Fix the three defects that cost a production restore
A live migration on 2026-08-24 stopped a production mail server and then
discovered the host's stalwart-cli was 0.13.4 - present, but from when the
CLI shipped with the server, with no `apply` command. The migration needs
v1.0.2+ from the separately-versioned stalwartlabs/cli repository.
Recovery was closed in both directions. v0.16's recovery-mode boot had
already bumped the store schema to v6, so the 0.15.5 binary refused to
reopen it ("expected 5 or below, found 6"). Going forward needed
export.json, which this tool's own failure path had deleted - and
regenerating it required a settings dump from a live v0.15 instance that
could no longer start. The operator restored a day-old snapshot and lost a
day of mail across nine domains.
Three fixes:
1. preflight.CheckExternalTools verifies stalwart-cli exists and is v1.0.2
or later, and that python3 runs - before anything is touched. Every fact
needed to prevent this was available in under a second from a stopped
state. Skipped for a patch upgrade, which invokes neither tool.
2. A failed run no longer deletes its work directory. Cleaning up on every
exit path was right for a sandboxed rehearsal and catastrophic here:
once the service is stopped the settings dump cannot be regenerated, so
deleting it removes the only way forward. The failure now prints the
resume command instead.
3. `run --resume <id>` continues an interrupted run. The checkpoint
machinery existed but never engaged, because run created a new run every
invocation - so a retry re-ran preflight against a binary already moved
aside, and failed. Completed steps are skipped from the checkpoint.
Proven against a VM built to match the failure: stalwart-cli 0.15.5,
accounts and mail seeded.
* preflight refused, service still active, mail still accepted
* a stub CLI passing --version and failing apply left the run stopped
with all eight inputs intact and the resume command printed
* --resume carried it to a clean finish: five seconds of downtime,
listeners regenerated, admin role restored, quotas rebuilt
That failure-path test is the one that should have run before production.
Every earlier test had stalwart-cli installed from the start, and the one
failure I did exercise happened to leave its artifacts behind.
This commit is contained in:
@@ -27,7 +27,12 @@ type Options struct {
|
||||
AdminPassword string
|
||||
TargetVersion string // e.g. "0.16.14" or "latest"
|
||||
MinFreeMultiple float64
|
||||
HTTPClient *http.Client
|
||||
// CLIPath and PythonPath are the external programs the migration
|
||||
// shells out to. Checked before anything is touched - see
|
||||
// CheckExternalTools.
|
||||
CLIPath string
|
||||
PythonPath string
|
||||
HTTPClient *http.Client
|
||||
}
|
||||
|
||||
// Checker runs the preflight checks described in ARCHITECTURE.md §4.1.
|
||||
@@ -107,7 +112,7 @@ func (c *Checker) Run(ctx context.Context, store *checkpoint.Store, rs *checkpoi
|
||||
return report, err
|
||||
}
|
||||
|
||||
if _, err := runCheck("upgrade-direction", func() (CheckResult, string) {
|
||||
boundaryOutcome, err := runCheck("upgrade-direction", func() (CheckResult, string) {
|
||||
curV, errCur := parseSemver(versionOutcome.Extra)
|
||||
tgtV, errTgt := parseSemver(targetOutcome.Extra)
|
||||
if errCur != nil || errTgt != nil {
|
||||
@@ -123,15 +128,27 @@ func (c *Checker) Run(ctx context.Context, store *checkpoint.Store, rs *checkpoi
|
||||
return CheckResult{
|
||||
Status: StatusOK,
|
||||
Detail: fmt.Sprintf("%s -> %s crosses the 0.15/0.16 major boundary: full recovery-mode migration plan required (ARCHITECTURE.md §4.4)", curV, tgtV),
|
||||
}, ""
|
||||
}, "crosses"
|
||||
}
|
||||
return CheckResult{
|
||||
Status: StatusOK,
|
||||
Detail: fmt.Sprintf("%s -> %s is a same-boundary patch upgrade: fast-path plan applies (ARCHITECTURE.md §4.6)", curV, tgtV),
|
||||
}, ""
|
||||
}); err != nil {
|
||||
}, "patch"
|
||||
})
|
||||
if err != nil {
|
||||
return report, err
|
||||
}
|
||||
crossesBoundary := boundaryOutcome.Extra != "patch"
|
||||
|
||||
// Before anything else that matters: are the tools this migration
|
||||
// depends on actually here? Discovering a missing stalwart-cli after
|
||||
// the service has been stopped is what this exists to prevent.
|
||||
for _, res := range CheckExternalTools(ctx, c.opts.CLIPath, c.opts.PythonPath, crossesBoundary) {
|
||||
result := res
|
||||
if _, err := runCheck(result.Name, func() (CheckResult, string) { return result, "" }); err != nil {
|
||||
return report, err
|
||||
}
|
||||
}
|
||||
|
||||
deploymentOutcome, err := runCheck("deployment-kind", func() (CheckResult, string) {
|
||||
kind := DetectDeploymentKind(ctx, c.opts.ContainerName)
|
||||
|
||||
Reference in New Issue
Block a user