# stalwart-migrator In-place upgrade tool for Stalwart Mail Server, 0.15.5 → latest: no data loss, a checkpoint at every step so an interrupted run resumes instead of restarting, and automated validation that the server still works afterwards. **Recovery from a failed migration is your own snapshot or backup — this tool does not undo a migration.** See [Recovery is your job](#recovery-is-your-job) before using it on anything you care about. Go, standard library only — no external dependencies. ## Status Partially implemented. Roughly 8,800 lines of tested code. Every phase except staging now exists as a package, including cutover, but nothing wires them into a production run yet, so `run` still refuses. | Command | State | |---|---| | `stalwart-migrate preflight` | **Works** — read-only checks and a migration plan | | `stalwart-migrate run --dry-run` | **Works** — preflight, real backup, sandboxed trial conversion | | `stalwart-migrate run` | **Refuses on purpose** — see below | | `stalwart-migrate status ` | **Works** | | `stalwart-migrate report ` | Not implemented | **`run` without `--dry-run` deliberately refuses to proceed.** Cutover (ARCHITECTURE.md §4.5) is implemented and tested, but nothing calls it: the staging phase (§4.3) and the production pipeline that would run preflight → backup → stage → recovery-mode → cutover → validate against real paths don't exist yet. `run` stops rather than going partway. That refusal is the correct behaviour today, not a bug. Package state: | Package | Lines | Tests | |---|---|---| | `internal/backup` | 1806 | yes | | `internal/preflight` | 1329 | yes | | `internal/stalwartapi` | 1276 | yes | | `internal/cutover` | 1186 | yes | | `internal/validate` | 792 | yes | | `internal/recovery` | 702 | yes | | `internal/checkpoint` | 556 | yes | | `internal/service` | 467 | yes | | `internal/plan` | 195 | yes | | `internal/config` | stub | — | ## Why not a shell script Stalwart's 0.15 → 0.16 boundary is not a drop-in binary swap: settings move, and the data directory has to be migrated rather than merely copied. The failure mode that matters is a half-migrated mail store with no way back — which is why backup verification and checkpointing are the design centre rather than conveniences bolted on afterwards — and why the tool refuses to cut over until you confirm you have a way back. [`ARCHITECTURE.md`](ARCHITECTURE.md) covers this in full: §1 on why a thin wrapper is insufficient, §4 on the migration phases, §5 on the checkpoint state machine, §6 on the CLI surface, and §8 on what is still open. ## Build and test ```sh go build ./... go test ./... ``` Requires Go 1.26 or newer. ## Trying it safely `preflight` is the sensible starting point — its checks against the Stalwart installation are read-only: ```sh sudo go run ./cmd/stalwart-migrate preflight ``` It still needs write access, because it records the run as a checkpoint before doing anything else: ``` create run: checkpoint: create run directory: mkdir /var/lib/stalwart-migrator: permission denied ``` That path is `checkpoint.DefaultBaseDir`, a compile-time constant with no flag or environment override — so preflight needs either root or a pre-created writable `/var/lib/stalwart-migrator`. (`run` takes `--work-dir` for its scratch space, but that is a different directory and does not move the checkpoint store.) `run --dry-run` performs a **real backup**, which touches the live data directory — read the caveat the command prints before using it on anything you care about. Where the plan crosses the 0.15/0.16 boundary it clones that verified backup into a disposable sandbox and converts the copy, leaving the original untouched. ## Recovery is your job **This tool does not undo a migration.** There is no `rollback` command. Recovery from a failed migration is your own snapshot or backup, taken by whatever method you already trust and know how to restore — a ZFS, LVM or btrfs snapshot, a VM or volume snapshot, or a restorable backup. Choosing that method, taking it, and verifying you can actually restore from it is out of scope for this tool: it does not take one, does not check that one exists, and cannot restore from one. Cutover refuses to start until you confirm a recovery point exists. That confirmation is an acknowledgement, not a check — nothing here can verify your snapshot. Its only purpose is that nobody migrates a production mail server having never been asked the question. **Take the snapshot with the service stopped** if you want a clean one. A snapshot of a running Stalwart is crash-consistent rather than clean; RocksDB will usually recover from its WAL, but "usually" is doing real work in that sentence. ### Restoring from a snapshot loses mail delivered since Reverting to any pre-migration recovery point discards mail delivered between taking it and restoring it. This is inherent to restoring a point in time and this tool cannot solve it — plan your migration window with that in mind, and consider holding inbound mail at a secondary MX for the duration if the gap matters to you. ### What the tool does to make a manual restore easier - **The old binary is preserved**, never deleted, next to the new one as `.v` — so putting things back doesn't depend on re-downloading a specific old release under pressure. - **The original service definition is preserved** as `.pre-` before cutover rewrites it, so you aren't reconstructing a unit file from memory. - **The settings and principals dumps** taken during backup stay on disk. - **Every artifact path and checksum is in the checkpoint**, and `stalwart-migrate status ` prints exactly which steps completed and which failed — which is the first thing you want when deciding what to restore. None of this is a substitute for the snapshot. It's what makes the twenty minutes after restoring one less unpleasant. ### Why it works this way An earlier version of this tool implemented rollback itself: it restored the filesystem backup, verified every restored file against a manifest, replayed SQL dumps, reinstalled the old binary, and re-validated the result. It was tested and it looked good. It was removed, because restoring bytes correctly is not the hard part — it copied contents and permissions but not *ownership*, so run as root it would have produced a byte-perfect, checksum-verified, root-owned data directory that Stalwart, running as its own user, could not open, and it would have reported success. A filesystem snapshot has no such failure mode, because it never lost the metadata to begin with. ARCHITECTURE.md §4.8 records the full reasoning.