From 18a55ccd5acb2ad33b9cf02fe5374e472c748254 Mon Sep 17 00:00:00 2001 From: John Coffey Date: Fri, 4 Sep 2026 19:16:01 -0700 Subject: [PATCH] Realign the roadmap with what is already built Three of the roadmap's claims were contradicted by the repository itself. Fleet management was Phase 11, "Planned", and positioning.md said config still flowed to the agent from the host rather than from the platform. agent-management-design.md has recorded the opposite for some time: punch list complete, verified live, with central authoring, versioning, rollout on the next check-in, observation, and a restart command a real agent picks up and acts on. There is no Phase 11 now. Its remainder is either already named there -- stop/uninstall, per-host multi-row alerting, a rule-per-host generator -- or belongs to Phase 8, since distributing rules is the one genuinely new thing the mechanism has to carry, and a rule language nobody can push to a fleet is not worth having. That inverts the old ordering argument, which put fleet last on the grounds that it manages configuration the earlier phases define. Sound reasoning; the world went the other way and built the mechanism first. Recorded rather than quietly dropped, because the instinct behind it is a good one that happened not to apply. Retention was listed as a question Phase 10 would finally have to answer. Half of it is answered: api/logretention serves operator-driven preview and delete with an owner-only per-agent floor. What is missing is an automatic TTL, so Phase 10 owns tiering and automatic TTL rather than retention from nothing. The rule-language recommendation is now settled rather than proposed, and not on its own authority: DesiredOverride is already a closed typed shape that cannot carry arbitrary code, so no channel exists that would deliver JavaScript to an agent even if the language argument had gone the other way. That surfaced a requirement nothing had written down. Agent management rests on an invariant it states outright -- every editable field degrades behaviour without cutting off the agent's ability to receive the next correction. Processing rules break it: a rule that panics or loops strands the agent exactly the way a corrupted ingest endpoint would, across every host it reached first. Phase 8 now owes either total evaluation or apply-then-verify with rollback, chosen deliberately rather than discovered mid-rollout. Signed-off-by: John Coffey --- README.md | 2 +- docs/architecture.md | 13 +++--- docs/positioning.md | 100 +++++++++++++++++++++++++++++++++++-------- docs/status.md | 28 +++++++++--- 4 files changed, 114 insertions(+), 29 deletions(-) diff --git a/README.md b/README.md index 7c8c252..8d79bd9 100644 --- a/README.md +++ b/README.md @@ -217,7 +217,7 @@ to compute. So `enterprise/` stays in the tree, AGPLv3 and working, and is not on the roadmap. It remains the answer for serving other people's data — a question -this project is not asking. Nothing in Phases 8-11 depends on it. +this project is not asking. Nothing in Phases 8-10 depends on it. ### What would close the gap to production-ready diff --git a/docs/architecture.md b/docs/architecture.md index 84bd1d3..6ff6fb0 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -271,11 +271,14 @@ defense against application-layer bugs, not an operational control. ## Open questions for you to resolve -- Retention/TTL policy for the ClickHouse `logs` table — not specified yet. - This was deferred until storage sizing became a real concern; Phase 10's - archive/replay work is where it stops being deferrable, since tiering to - object storage and reading back from it is the same question asked from - the other side. See [positioning.md](positioning.md). +- **Automatic** retention/TTL policy for the ClickHouse `logs` table — not + specified yet. The operator-driven half exists: `api/logretention` serves + preview and delete, with an owner-only per-agent floor. What is missing is + a policy that expires data without somebody asking it to. Deferred until + storage sizing became a real concern; Phase 10's archive/replay work is + where it stops being deferrable, since tiering to object storage and + reading back from it is the same question asked from the other side. See + [positioning.md](positioning.md). - Exact OTel log schema field mapping (which OTel resource/log attributes map to which ClickHouse columns) — Phase 0 uses a minimal subset (timestamp, host, service, severity, message, attributes map); full diff --git a/docs/positioning.md b/docs/positioning.md index a611b35..62b0bf4 100644 --- a/docs/positioning.md +++ b/docs/positioning.md @@ -131,6 +131,39 @@ Central processing at the ingest tier is the complement: rules that need context the agent lacks, and a place to change behaviour without a fleet rollout. +**Treat that recommendation as settled, because the fleet design already +settled it.** `DesiredOverride` — the only channel that can deliver +anything to an agent — is deliberately a closed, typed shape that cannot +carry arbitrary code, and permanently excludes any field capable of +stranding an agent. Two documents reached the same conclusion from +opposite directions, one reasoning about expressiveness and one about +blast radius. There is no version of this where rules arrive as +JavaScript, because there is no channel that would carry it. + +**And one requirement that falls out of the same design, which nothing +has written down until now.** Agent management rests on an invariant it +states explicitly: every editable field "degrades the agent's behavior +without ever cutting off its ability to receive the next correction." A +bad batch size is survivable precisely because the agent still checks in +and can be corrected. + +Processing rules break that invariant. A rule that panics, loops +forever, or exhausts memory strands the agent exactly the way a corrupted +`ingest.endpoint` would — the one channel capable of fixing the mistake +is the thing the mistake killed, across however many hosts the rule +reached before anyone noticed. So Phase 8 owes one of two things, chosen +deliberately rather than discovered during a rollout: + +1. **Total evaluation** — a rule set that provably cannot panic, cannot + loop unboundedly, and cannot allocate without limit. A typed + declarative DSL can offer this; it is a third argument for one. +2. **Apply-then-verify** — the agent treats a new rule set as + provisional, and reverts to the last known-good set if it crash-loops + before the next successful check-in. + +The first is better if it can be had, because the second is a recovery +mechanism and the first is an absence of the failure. + ### 2. Routing and multiple destinations Conditional routing — this source, matching this rule, to these @@ -144,17 +177,36 @@ Elastic bulk, OTLP, Kafka, plain HTTP. An archive format on object storage, and the ability to read it back into the pipeline or into a destination later. This is the feature that makes aggressive reduction safe: you can drop something from the hot path -precisely because you can get it back. It also folds in the retention/TTL -question `/docs/architecture.md` currently lists as unresolved and -deferred — that question stops being deferrable here. +precisely because you can get it back. -### 4. Fleet management +Retention is half of a question rather than all of one. `api/logretention` +already ships operator-driven preview and delete, with an owner-only +per-agent floor. What does not exist is an *automatic* TTL policy, which +is the half `/docs/architecture.md` still lists as unresolved, and the +half that stops being deferrable here — tiering to object storage and +reading back from it is the same question asked from the other side. -Central agent configuration: author, version, roll out, and observe. Much -of the substrate exists — agents check in, report their own version and -source config, and there is an Agents page that already knows when one -goes stale. What is missing is the direction of travel: config currently -flows *to* the agent from the host, not from the platform. +### 4. Fleet management — already built + +This section used to say that config flowed to the agent from the host +rather than from the platform, and that the direction of travel was what +was missing. That has not been true for some time. +[`agent-management-design.md`](agent-management-design.md) records the +whole thing as complete and verified live: config authored centrally +(`PUT /agents/{host}/config`), versioned +(`desired_override_version`/`applied_override_version`), rolled out on the +agent's next check-in, observed on the Agents page, and a restart command +that a real agent picks up and acts on. + +What remains is named there and is small: `stop`/`uninstall` lifecycle +commands, true per-host multi-row alerting, and a rule-per-host generator. + +The consequence for the roadmap is bigger than the correction. Fleet was +going to be the last phase, on the argument that it manages configuration +the earlier phases define. The mechanism arrived first instead, so the +question is no longer "how do we distribute config" but "what new thing +does Phase 8 need it to carry" — which makes distribution part of Phase 8 +rather than a phase of its own. ### 5. Schema normalisation as a feature, not a detail @@ -227,21 +279,33 @@ continuation of the first, and it is worth numbering separately rather than appending forever to a list that was about analytics. - **Phase 8 — Processing.** The rule DSL, agent-side execution, - ingest-side execution, and the tests that prove a rule does the same - thing in both places. + ingest-side execution, the tests that prove a rule does the same thing + in both places, and distribution: carrying rule sets through + `DesiredOverride` without breaking the strand-safety invariant above. - **Phase 9 — Routing and sinks.** Multiple destinations, conditional routing, per-destination delivery guarantees, and the first three sinks: object storage, OTLP, Splunk HEC. -- **Phase 10 — Archive and replay.** The archive format, retention and - tiering, and replay back into the pipeline or out to a destination. -- **Phase 11 — Fleet.** Config authored centrally, versioned, rolled out - and observed. +- **Phase 10 — Archive and replay.** The archive format, tiering, + automatic TTL, and replay back into the pipeline or out to a + destination. + +**There is no Phase 11.** It was going to be fleet management, and fleet +management is built — see section 4 above. What is left of it either +belongs to Phase 8 (distributing rules) or is a small named remainder +already tracked in +[`agent-management-design.md`](agent-management-design.md). Ordering is deliberate. Processing without routing still pays for itself by shrinking what is stored; routing without processing forwards -everything and helps nobody. Archive depends on both. Fleet is last -because it manages configuration the earlier phases define — building it -first would mean managing settings that do not exist yet. +everything and helps nobody. Archive depends on both. + +The original ordering put fleet last, reasoning that it manages +configuration the earlier phases define. That reasoning was sound and the +world went the other way: the mechanism was built first, and Phase 8 now +inherits a distribution channel rather than needing one built for it +afterwards. Worth recording, because the instinct to schedule the +management layer after the thing it manages is a good one that happened +not to apply here. ## The names already argue the case diff --git a/docs/status.md b/docs/status.md index 41023ba..b27da6e 100644 --- a/docs/status.md +++ b/docs/status.md @@ -20,19 +20,37 @@ verification procedure and its results. | 5 | Frontend redesign and design system | Shipped | | 6 | License compliance audit and remediation | Shipped | | 7 | AI-assisted query authoring | Shipped | -| 8 | Processing: rule DSL, agent-side and ingest-side | Planned | +| 8 | Processing: rule DSL, agent-side and ingest-side, and its distribution | Planned | | 9 | Routing: multiple destinations and delivery guarantees | Planned | -| 10 | Archive and replay, retention and tiering | Planned | -| 11 | Fleet: central agent configuration | Planned | +| 10 | Archive and replay, tiering, automatic TTL | Planned | +| — | Fleet: central agent configuration | Mostly shipped, see below | | — | AI-assisted analysis and explanation, on a local model | End-state goal | -Phases 8-11 are a second axis rather than a continuation of the first: -0-7 built the destination, and those four build the road to it. The +Phases 8-10 are a second axis rather than a continuation of the first: +0-7 built the destination, and those three build the road to it. The argument for taking that on, including the part where cheap storage removes the usual reason to buy a pipeline at all, is [`positioning.md`](positioning.md). Nothing in them is started. None of them depends on Phase 4. +**There is no Phase 11.** It was going to be fleet management, and fleet +management already exists: central config authoring, versioning, pull-based +rollout, observation, and a working restart command, all recorded and +verified live in +[`agent-management-design.md`](agent-management-design.md). The roadmap +was describing it as future work long after it stopped being any. Its real +remainder is small and already disclosed there -- `stop`/`uninstall` +lifecycle commands, true per-host multi-row alerting, and a rule-per-host +generator -- and the one piece that is genuinely new, distributing +processing rules, belongs to Phase 8, because a rule language nobody can +push to a fleet is not worth having. + +Retention is half-built too, which changes what Phase 10 owns. +`api/logretention` ships operator-driven preview and delete with an +owner-only per-agent floor. What does not exist is an *automatic* TTL +policy, which is why Phase 10 is now "tiering, automatic TTL" rather than +"retention" from nothing. + **Decision, 2026-09-05: tenancy is for other people's data; installations are for environments.** Phase 4 stays shipped and stays in the tree, and comes off the roadmap. Separating two environments means running two