main
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7cc2fd8c78 |
Draft the Phase 8 processing design
Starts from the distribution channel and the safety invariant it protects, and derives the language from them, rather than designing a rule language and asking later how to ship it. Four decisions proposed. A rule is a matcher plus ordered typed actions, with no expressions and nothing resembling eval -- less expressive than Cribl on purpose, and the only shape that can be pushed to ten thousand hosts and audited by reading it. Total evaluation is the primary safety guarantee, with apply-then-verify as a backstop. One spec with two implementations means a language-neutral conformance suite is the specification and should be built first, not last. Distribution reuses DesiredOverride, following extra_file_paths as the precedent for a repeated field. It also corrects something #21 got wrong. That change said a fatal rule would strand an agent the way a corrupted ingest endpoint does. Reading apply_override's actual semantics, overrides live only in the running process's memory and are never written to disk, so a restarted agent boots clean and re-syncs -- a fatal rule set crash-loops rather than strands, and the agent keeps checking in, so it stays correctable. That is a much better failure mode, and it was acquired by accident: the "don't persist" choice was made to avoid filesystem writes on read-only images, not for safety. This design promotes it to a constraint, since persisting overrides later would silently convert every crash-loop into a strand. positioning.md is corrected to match rather than left disagreeing. Five open questions are left open rather than answered to look decisive, the sharpest being that there is no staged rollout today: an edit reaches every matching agent at once, which for executable rules is the difference between breaking one host and breaking all of them. The v1 acceptance test is real data, not a fixture: two processes on the maintainer's own workstation account for 308 of 325 journal entries in five minutes, and a suppress rule should remove about 60% of that host's volume. Signed-off-by: John Coffey <[email protected]> |
||
|
|
18a55ccd5a |
Realign the roadmap with what is already built
Three of the roadmap's claims were contradicted by the repository itself. Fleet management was Phase 11, "Planned", and positioning.md said config still flowed to the agent from the host rather than from the platform. agent-management-design.md has recorded the opposite for some time: punch list complete, verified live, with central authoring, versioning, rollout on the next check-in, observation, and a restart command a real agent picks up and acts on. There is no Phase 11 now. Its remainder is either already named there -- stop/uninstall, per-host multi-row alerting, a rule-per-host generator -- or belongs to Phase 8, since distributing rules is the one genuinely new thing the mechanism has to carry, and a rule language nobody can push to a fleet is not worth having. That inverts the old ordering argument, which put fleet last on the grounds that it manages configuration the earlier phases define. Sound reasoning; the world went the other way and built the mechanism first. Recorded rather than quietly dropped, because the instinct behind it is a good one that happened not to apply. Retention was listed as a question Phase 10 would finally have to answer. Half of it is answered: api/logretention serves operator-driven preview and delete with an owner-only per-agent floor. What is missing is an automatic TTL, so Phase 10 owns tiering and automatic TTL rather than retention from nothing. The rule-language recommendation is now settled rather than proposed, and not on its own authority: DesiredOverride is already a closed typed shape that cannot carry arbitrary code, so no channel exists that would deliver JavaScript to an agent even if the language argument had gone the other way. That surfaced a requirement nothing had written down. Agent management rests on an invariant it states outright -- every editable field degrades behaviour without cutting off the agent's ability to receive the next correction. Processing rules break it: a rule that panics or loops strands the agent exactly the way a corrupted ingest endpoint would, across every host it reached first. Phase 8 now owes either total evaluation or apply-then-verify with rollback, chosen deliberately rather than discovered mid-rollout. Signed-off-by: John Coffey <[email protected]> |
||
|
|
0ee2e9183b |
Take multi-tenancy off the roadmap
Cairn OBS is self-hosted, and the way to separate two environments is to run two installations rather than two tenants inside one. Tenancy is the wrong boundary for that, on three counts this repository demonstrates rather than assumes: chwriter.WriteBatch is all-or-nothing across tenants, so one tenant's failure stalls offset progress for every other; CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT puts every tenant's data behind a single superuser credential, as docker-compose.yml's own comment says; and one binary with one set of migrations moves every tenant together, which is the opposite of what separate environments are for. A whole installation idles at about 1.3 GB, so the sharing buys nothing. The project led with multi-tenant RBAC in the README banner and in PROJECT-SPEC's goal statement. Both now say what it is instead: self-hosted. "Open-core" goes with them -- it was already inaccurate, since CONTRIBUTING states there is no feature gate and no paid tier, and with enterprise/ off the roadmap there will not be one. A second identity provider comes off the list of things standing between this and production-ready. SSO belongs to enterprise/, and a self-hosted deployment is not waiting on it. Terraform's tenant/RBAC resources move from "disclosed future work" to not planned. Nothing is scrubbed from the record. Phase 4 stays shipped, its runbook stays, and its known gaps stay stated -- rewriting that history would contradict the candour the Status section is built on. enterprise/ stays in the tree, AGPLv3 and working, as the answer to a question this project is not asking. Signed-off-by: John Coffey <[email protected]> |
||
|
|
d49ddb943e |
Add the AI axis: plain English as an option, analysis as the end state
Cost is the argument against Splunk and control is the argument against Cribl. AI is the third, and the difference there is not a feature comparison, it is where the model runs. Plain-English querying shipped in Phase 7 and stays an option rather than a replacement for writing a query: every generated query compiles through the same IR and executor as a hand-written one, with the same tenant scoping, cost guardrails and audit logging. The model suggests, it does not get a private path to the data. AI-assisted analysis and explanation is the end state and is not built. Authoring answers "how do I ask this"; the valuable question is "what does this mean" -- what changed in a result set, why an alert fired and what preceded it, summarising an incident from the records around it. Recorded on the status page as an end-state goal rather than a numbered phase, because it is a property the product keeps rather than a thing to finish and tick off. Local is the non-negotiable part, and it is worth stating as position rather than as a bullet: the default runs qwen2.5-coder through Ollama on the customer's own hardware, Apache-2.0 weights chosen so Phase 6's licence work survives contact with the model, and the cloud adapter is opt-in and off by default. Logs are the most sensitive unstructured data most organisations hold -- credentials in stack traces, customer identifiers, internal topology -- so an assistant that reads them is either running where the data already is, or it is a data-egress decision wearing a helpful interface. The constraint it imposes is stated too, because it bounds what can be promised: a 7B model on a customer's hardware will not match a frontier model, and the honest claim is not that it is as clever but that it is good enough at a bounded task and runs somewhere you control. Analysis features have to be designed to that budget rather than assuming an API is one call away. |
||
|
|
e57486add5 |
Position against Cribl as well as Splunk, and say what that costs us
Splunk and Cribl are not the same competitor and the claim is not the same claim twice. Splunk is the destination and Cairn OBS replaces it, which is what phases 0-7 were for. Cribl is the road: routing, reduction, enrichment, redaction and replay on the way to wherever data is going. Cairn OBS is already a road in shape -- agent, Redpanda, ingest -- and exposes none of a pipeline's controls. The agent cannot filter, drop, sample, mask or re-route anything; ingest normalises a schema and writes it; there is exactly one destination and it is us. Reconciling the two turns up something a cost-led project has to face rather than paper over: most people buy Cribl because Splunk is expensive per gigabyte, so being genuinely cheap per gigabyte removes the main reason to buy Cribl in front of us. That makes the strongest pitch "one system where there were two" rather than "we are also a pipeline vendor" -- but that pitch only survives a buyer if we also do the four things people buy a pipeline for that are not about spend: routing to several destinations, redacting before data leaves the network, archive and replay, and not being locked to one analytics vendor. Those are about control, which is better ground anyway: cost advantages get matched and architectural ones do not. The consequence is uncomfortable and is written down as a decision rather than left to be discovered: competing with Cribl means being able to send data to S3, Splunk HEC, Elastic, OTLP and Kafka -- building features whose purpose is to help data leave this platform. A project that refuses lock-in in its licence and then builds it into its egress would be lying about itself. Four phases follow, ordered so each pays for itself: processing (8), routing (9), archive and replay (10), fleet (11). Processing without routing still shrinks what is stored; routing without processing forwards everything and helps nobody. One design decision is called out now because it collides with a non-negotiable constraint. Cribl's rule language is JavaScript, and embedding a JS engine in a statically-linked musl agent would end "no glibc runtime deps" as a claim. The recommendation is a declarative rule DSL -- matchers and typed actions, no arbitrary code -- deliberately less expressive, small enough to audit and safe to push to ten thousand hosts. The retention/TTL question in architecture.md is no longer deferrable and now says so: Phase 10 asks it from the other side. |
||
|
|
252c207ccf |
Say that Phase 4 shipped, and that the environment proving it is gone (#7)
The status file and the README both still said Phase 4 was not shipped because the environment had lost Docker and database access partway through, and that only the audit-logging guarantees had been confirmed against a live database. That stopped being true some time ago. phase-4-runbook.md records the opposite in detail: Docker access came back, a real docker-compose stack ran with real ClickHouse and Postgres and two provisioned tenants, a local kind cluster ran the Helm chart end to end, and both SSO protocols were verified against a real Auth0 tenant with full browser round trips. Eight real bugs came out of that, six from compose and two from the chart's first real install -- none of them findable without the infrastructure. PROJECT-SPEC.md sends readers to status.md and tells them to read it before assuming a capability works end to end, so the one file that is meant to be authoritative was the one understating the project by the widest margin. Correcting it matters more now than it would have last week, because the evidence cannot be regenerated: proto.cairnobs.org and the VPS under it were retired on 2026-09-04, taking the mTLS CA, the server certificate and six enrolled agents with them. The runbooks are what is left. Three gaps are now stated rather than implied: The prototype is gone, so none of this can be re-run today without building one. The DNS was kept for that; the certificates deliberately were not. demo.cairnobs.org is live and is not evidence for Phase 4. It runs COMPOSE_PROFILES=single-tenant, so it exercises the OSS path and says nothing about RBAC, tenant isolation or per-tenant ClickHouse. A healthy demo proving multi-tenancy is exactly the wrong inference to leave available. SSO has been tried against one IdP and one local kind cluster, not two IdPs and not a production-grade cluster. The Terraform entry now names its cause instead of pointing at another file: alerting exposes no PUT for rules or targets, and neither rulestore.Store nor notifystore.Store has an Update method to wire one to, so Terraform destroys and recreates -- which resets alert_state and delivery-log continuity. |
||
|
|
f756a9d4f6 |
Rename the project spec and update every reference to it
The charter file carried a tool-specific name while being the repository's own document: mission, non-negotiable constraints, the pinned stack, repo conventions and phase status, cited as authority by thirty files across the agent, api, deploy, docs, search and terraform trees. PROJECT-SPEC.md says what it is. All 42 references are updated in the same commit, including the relative link in docs/status.md, so nothing points at a filename that no longer exists. |
||
|
|
914c0af467 |
docs: split project status out of CLAUDE.md
CLAUDE.md was doing two jobs: durable repo conventions, and a ~500-line phase-by-phase status narrative that duplicates the per-phase runbooks and goes stale the moment a phase ships. Keep mission, constraints, pinned stack, conventions, and "when in doubt" in CLAUDE.md (572 -> 78 lines). Move the phase record verbatim to docs/status.md, prefaced with a summary table and the known verification gaps. Content is byte-identical; nothing was reworded or dropped. Link both directions, and point the README's status section at the new file. |