Fix the three defects that cost a production restore

A live migration on 2026-08-24 stopped a production mail server and then
discovered the host's stalwart-cli was 0.13.4 - present, but from when the
CLI shipped with the server, with no `apply` command. The migration needs
v1.0.2+ from the separately-versioned stalwartlabs/cli repository.

Recovery was closed in both directions. v0.16's recovery-mode boot had
already bumped the store schema to v6, so the 0.15.5 binary refused to
reopen it ("expected 5 or below, found 6"). Going forward needed
export.json, which this tool's own failure path had deleted - and
regenerating it required a settings dump from a live v0.15 instance that
could no longer start. The operator restored a day-old snapshot and lost a
day of mail across nine domains.

Three fixes:

1. preflight.CheckExternalTools verifies stalwart-cli exists and is v1.0.2
   or later, and that python3 runs - before anything is touched. Every fact
   needed to prevent this was available in under a second from a stopped
   state. Skipped for a patch upgrade, which invokes neither tool.

2. A failed run no longer deletes its work directory. Cleaning up on every
   exit path was right for a sandboxed rehearsal and catastrophic here:
   once the service is stopped the settings dump cannot be regenerated, so
   deleting it removes the only way forward. The failure now prints the
   resume command instead.

3. `run --resume <id>` continues an interrupted run. The checkpoint
   machinery existed but never engaged, because run created a new run every
   invocation - so a retry re-ran preflight against a binary already moved
   aside, and failed. Completed steps are skipped from the checkpoint.

Proven against a VM built to match the failure: stalwart-cli 0.15.5,
accounts and mail seeded.

  * preflight refused, service still active, mail still accepted
  * a stub CLI passing --version and failing apply left the run stopped
    with all eight inputs intact and the resume command printed
  * --resume carried it to a clean finish: five seconds of downtime,
    listeners regenerated, admin role restored, quotas rebuilt

That failure-path test is the one that should have run before production.
Every earlier test had stalwart-cli installed from the start, and the one
failure I did exercise happened to leave its artifacts behind.
This commit is contained in:
2026-08-23 23:20:47 -07:00
parent 5a4c175042
commit 9faa21f4f1
10 changed files with 370 additions and 11 deletions
+8
View File
@@ -44,6 +44,10 @@ func countLines(t *testing.T, path string) int {
}
func TestCheckerRunEndToEndAndResume(t *testing.T) {
// preflight now verifies the external tools before anything else; give
// this environment ones that satisfy it.
fakeTool(t, "stalwart-cli", "stalwart-cli 1.0.12")
fakeTool(t, "python3", "Python 3.13.5")
// -- fixtures --------------------------------------------------------
counterPath := filepath.Join(t.TempDir(), "invocations")
binaryPath := writeFakeBinary(t, "0.15.5", counterPath)
@@ -200,6 +204,10 @@ func TestCheckerRunFlagsInsufficientDiskSpace(t *testing.T) {
}
func TestCheckerRunCapturesAccountSnapshotWhenAdminURLSet(t *testing.T) {
// preflight now verifies the external tools before anything else; give
// this environment ones that satisfy it.
fakeTool(t, "stalwart-cli", "stalwart-cli 1.0.12")
fakeTool(t, "python3", "Python 3.13.5")
counterPath := filepath.Join(t.TempDir(), "invocations")
binaryPath := writeFakeBinary(t, "0.15.5", counterPath)