Commit Graph
5 Commits
Author SHA1 Message Date
jcoffey-dev 56152ed7de Ask the container what it is running
Preflight's first check ran `--version` on --binary. A container-only host
has no such file, so the check failed, and because it is first, nothing
downstream ever ran — including every container check that exists to
decide whether that container can be migrated at all. The container path
was unreachable on exactly the hosts it is for.

A host that happens to have a binary is the worse case, not the better
one: a stray /usr/local/bin/stalwart from an older install answers
confidently with a version nothing is running, and the whole migration
plan is derived from that number.

The source version now comes from running the image the container is on,
by ID rather than by the tag it was started from, using the same command
and the same fallback stage already uses for the target image — the two
have to agree about what a Stalwart image reports or the source and
target could be read by different rules. Verified against a real
stalwartlabs/stalwart image, not only the fake.

Deployment kind is detected once and shared, rather than asked again by
the check that reports it. Two answers for one run is not a thing this
should be able to produce.

The same reasoning retires preserve-binary on a container: there is
nothing on this host to move aside, and the equivalent is already
guaranteed, since cutover renames the old container and never prunes the
old image. Renaming a stray binary would have preserved something nothing
was running.
2026-08-29 17:56:32 -07:00
jcoffey-dev e96e72bf79 Tell what a container inherits from what it overrides
Checked against a real stalwartlabs/stalwart image, `docker inspect` on an
ordinary container reports User "stalwart", Entrypoint
["/usr/local/bin/stalwart"] and Cmd ["--config",
"/etc/stalwart/config.json"] — all three inherited, none of them given.

Two things followed from reading those as the operator's.

A container user was listed as configuration a recreate would drop, so
every container off the official image was refused as unrecreatable. That
refusal lived in cutover, downstream of the stop, the settings conversion
and the store migration: it arrived with mail down and data already
moved, which is the failure issue #1 was filed for. Each of the three is
now compared against `docker image inspect` of the image the container is
on. Inherited values are left to the new image, whose own defaults are
the ones that go with it. Overrides are carried: --user, --entrypoint,
and the rest of an entrypoint as leading argv. Cmd and Entrypoint were
not being read at all, so an overridden one was silently dropped — the
exact loss the unsupported list exists to prevent.

The recreatability question also moved into preflight, while the server
is still running. Cutover asks it again, since the two are separated by
the whole migration, but only one of them can refuse without cost.

The other half: the recreated container is now started with `--config`
pointing at the migrated config in the data volume. Left to the image's
default command it came up on /etc/stalwart/config.json — a different
volume, holding whatever the old version left there — so cutover would
have produced a running server with nothing to do with the migration that
preceded it. An overridden command and that --config are the same argv
and cannot be merged honestly, so a container with one is refused and
told why.

The config is also chowned to whatever owns the data directory, before
the recovery cycle opens it. The image runs as uid 2000 and this tool
writes as root; §4.8 is the standing reminder that byte-perfect and
unreadable is a way to report success.

Found while checking @kaya-eu's field report in #1 against a real image.
Their three manual migrations are where the config step comes from.
2026-08-29 17:45:23 -07:00
jcoffey-dev 272439cf2a Wire the container path up, behind a flag that says what it is
Everything the container migration needs has landed a piece at a time and
nothing called any of it. `run` now does: stage pulls and verifies an image
instead of downloading a binary, the recovery cycle launches a throwaway
container against the live container's own mounts, and cutover recreates it.
Preflight's blanket refusal of docker goes with it -- what still refuses is
specific to a container rather than to containers, which is compose and
data that is not on a volume.

It refuses without --container-path-unproven, and that flag is the honest
part of this change. Every test drives a fake docker. That proves the right
commands are assembled and proves nothing about whether a real image reads
the config it is handed -- which is the exact limit ARCHITECTURE.md section
4.8 records about the rollback code that was deleted for being tested only
against fakes. A doc note seemed too quiet for a tool that stops a mail
server, so it is a flag nobody reaches without being told.

The converted config reaches the container through the data volume. It is
written under the host side of whichever mount covers --data-dir and named
on the container side, because cutover recreates a container with the mounts
it had and cannot invent a new one for a config file. --data-dir therefore
names the path inside the container, which preflight already says when it
matches no mount.

PatchPaths stays unused, deliberately. Its documented purpose is pointing a
rehearsal at a sandbox; a real container's dumped settings already carry
container-side paths, because they come from the live server rather than
from a file on this host.

The preflight test that asserted docker was refused outright now asserts
the replacement rather than being deleted -- "docker is allowed through
here" is the thing that would be wrong to regress. Its fixture had to make
--data-dir both a real host directory and one the fake container mounts,
since disk-space stats it and container-data-volume wants it covered.

README gains the container section and, at the top, the note that this is
ihasmail's companion.
2026-08-28 17:42:31 -07:00
jcoffey-dev ad7c2d5135 Cut over a container, or refuse to for a reason
A container cannot be edited in place the way a unit file can, so cutting
one over means rebuilding it. That makes silent loss the default failure:
a container recreated without its capabilities, its custom network or its
device mappings starts cleanly and is quietly not the server it was.

Section 4.5 already answers this for a unit file -- it rewrites in place
rather than regenerating, because a generated unit would drop hardening
options this tool has no business having an opinion about, and it refuses
to edit a line it only partly understands. The same rule applies here,
where the whole definition has to be rebuilt: the parts this understands
are carried across, and a container using anything else is refused by name
rather than rebuilt without it. The list of what it looks for is
conservative and not exhaustive, which is the safe direction: docker's
HostConfig has far more fields, and one this does not know about is a
reason not to be recreating that container at all.

The old container is renamed, not removed, and nothing here prunes the old
image. Together they are the container's manual restore path -- one command
starts the previous container again -- which is as close to section 4.2's
preserved binary as a container gets. The inspect output is preserved as an
artifact before anything is replaced, for the reason the unit file is: an
operator putting a machine back by hand should not also be reconstructing
the definition from memory.

Recovery-mode variables are stripped from the recreated container's
environment. Leaving STALWART_RECOVERY_MODE set would recovery-boot on
every restart, which is the same footgun the unit rewrite exists to
prevent.

Run now branches, with the health check and quota recalculation shared:
those ask the same question whatever started the server. The binary path
moved inside an else and is otherwise untouched -- no existing test needed
editing, which is the evidence for that.

The container path is opt-in through Options.Container, so a Docker
deployment without it is still refused exactly as before. Nothing calls it
yet; wiring `run` up and lifting preflight's refusal is what remains of #3,
and ARCHITECTURE.md says so in both places it previously said Docker was
refused outright.
2026-08-28 17:30:33 -07:00
jcoffey-dev 6e7a05e478 Tell a container operator what stands in the way
Preflight refuses a container, and until now that was all it said. The two
things that actually decide whether one could be migrated at all were never
looked at, so an operator was refused without being told what to fix or what
a manual migration would involve.

It now inspects the container and reports three things. What image it is
running, tag and digest kept apart because a moved tag makes them disagree
and only one of them says what is really there. Whether compose manages it,
which matters beyond this tool's current refusal: recreating a
compose-managed container out from under compose leaves the container and
the compose file disagreeing about what is deployed, and the next
`compose up` reverts the migration. And whether the data is on a volume at
all -- a container keeping it in its own writable layer loses it when the
container is replaced, and replacing the container is what migrating it
means, so that one is fatal in a way no later phase could recover from.

All three are advisory under rehearse, on the same reasoning that already
made the deployment check advisory there: rehearse never stops or recreates
anything, and an operator doing this by hand needs these facts more than an
automated run does.

Preflight stays read-only. Preserving the container definition as an
artifact belongs with the phase that replaces it, so it goes with cutover
rather than here.

One limitation this surfaced and does not fix: --data-dir naming a path
inside the container breaks the host-side disk-space check, which stats it
locally. Path translation is the next piece of work; the data-volume check
says so when the path matches no mount.
2026-08-28 16:50:09 -07:00