Commit Graph
163 Commits
Author SHA1 Message Date
dependabot[bot] 82cfe1fa0a Bump the go-minor-and-patch group across 5 directories with 9 updates
Bumps the go-minor-and-patch group with 1 update in the /alerting directory: [golang.org/x/sync](https://github.com/golang/sync).
Bumps the go-minor-and-patch group with 1 update in the /api directory: [golang.org/x/crypto](https://github.com/golang/crypto).
Bumps the go-minor-and-patch group with 3 updates in the /deploy/operator directory: [k8s.io/apimachinery](https://github.com/kubernetes/apimachinery), [k8s.io/client-go](https://github.com/kubernetes/client-go) and [sigs.k8s.io/controller-runtime](https://github.com/kubernetes-sigs/controller-runtime).
Bumps the go-minor-and-patch group with 6 updates in the /enterprise directory:

| Package | From | To |
| --- | --- | --- |
| [golang.org/x/sync](https://github.com/golang/sync) | `0.22.0` | `0.23.0` |
| [k8s.io/apimachinery](https://github.com/kubernetes/apimachinery) | `0.31.0` | `0.37.0` |
| [k8s.io/client-go](https://github.com/kubernetes/client-go) | `0.31.0` | `0.37.0` |
| [github.com/coreos/go-oidc/v3](https://github.com/coreos/go-oidc) | `3.20.0` | `3.21.0` |
| [github.com/go-jose/go-jose/v4](https://github.com/go-jose/go-jose) | `4.1.4` | `4.1.5` |
| [golang.org/x/oauth2](https://github.com/golang/oauth2) | `0.36.0` | `0.37.0` |

Bumps the go-minor-and-patch group with 1 update in the /ingest directory: [golang.org/x/sync](https://github.com/golang/sync).


Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/crypto` from 0.55.0 to 0.56.0
- [Commits](https://github.com/golang/crypto/compare/v0.55.0...v0.56.0)

Updates `k8s.io/apimachinery` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/apimachinery/compare/v0.31.0...v0.37.0)

Updates `k8s.io/client-go` from 0.31.0 to 0.37.0
- [Changelog](https://github.com/kubernetes/client-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/kubernetes/client-go/compare/v0.31.0...v0.37.0)

Updates `sigs.k8s.io/controller-runtime` from 0.19.3 to 0.25.0
- [Release notes](https://github.com/kubernetes-sigs/controller-runtime/releases)
- [Changelog](https://github.com/kubernetes-sigs/controller-runtime/blob/main/RELEASE.md)
- [Commits](https://github.com/kubernetes-sigs/controller-runtime/compare/v0.19.3...v0.25.0)

Updates `k8s.io/api` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/api/compare/v0.31.0...v0.37.0)

Updates `k8s.io/apimachinery` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/apimachinery/compare/v0.31.0...v0.37.0)

Updates `k8s.io/client-go` from 0.31.0 to 0.37.0
- [Changelog](https://github.com/kubernetes/client-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/kubernetes/client-go/compare/v0.31.0...v0.37.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `k8s.io/apimachinery` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/apimachinery/compare/v0.31.0...v0.37.0)

Updates `k8s.io/client-go` from 0.31.0 to 0.37.0
- [Changelog](https://github.com/kubernetes/client-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/kubernetes/client-go/compare/v0.31.0...v0.37.0)

Updates `github.com/coreos/go-oidc/v3` from 3.20.0 to 3.21.0
- [Release notes](https://github.com/coreos/go-oidc/releases)
- [Commits](https://github.com/coreos/go-oidc/compare/v3.20.0...v3.21.0)

Updates `github.com/go-jose/go-jose/v4` from 4.1.4 to 4.1.5
- [Release notes](https://github.com/go-jose/go-jose/releases)
- [Commits](https://github.com/go-jose/go-jose/compare/v4.1.4...v4.1.5)

Updates `golang.org/x/oauth2` from 0.36.0 to 0.37.0
- [Commits](https://github.com/golang/oauth2/compare/v0.36.0...v0.37.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `k8s.io/api` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/api/compare/v0.31.0...v0.37.0)

Updates `k8s.io/apimachinery` from 0.31.0 to 0.37.0
- [Commits](https://github.com/kubernetes/apimachinery/compare/v0.31.0...v0.37.0)

Updates `k8s.io/client-go` from 0.31.0 to 0.37.0
- [Changelog](https://github.com/kubernetes/client-go/blob/master/CHANGELOG.md)
- [Commits](https://github.com/kubernetes/client-go/compare/v0.31.0...v0.37.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

Updates `golang.org/x/sync` from 0.22.0 to 0.23.0
- [Commits](https://github.com/golang/sync/compare/v0.22.0...v0.23.0)

---
updated-dependencies:
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/crypto
  dependency-version: 0.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/apimachinery
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/client-go
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: sigs.k8s.io/controller-runtime
  dependency-version: 0.25.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/api
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/apimachinery
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/client-go
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/apimachinery
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/client-go
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: github.com/coreos/go-oidc/v3
  dependency-version: 3.21.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: github.com/go-jose/go-jose/v4
  dependency-version: 4.1.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/oauth2
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/api
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/apimachinery
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: k8s.io/client-go
  dependency-version: 0.37.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
- dependency-name: golang.org/x/sync
  dependency-version: 0.23.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: go-minor-and-patch
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 17:00:23 +00:00
jcoffey e071216378 Merge pull request #34 from Coffey-Labs/dependabot/cargo/agent/quick-xml-0.42.0
Bump quick-xml from 0.41.0 to 0.42.0 in /agent
2026-09-10 09:58:12 -07:00
jcoffey 7c9aeb7406 Merge pull request #41 from Coffey-Labs/go-1.26-pins
Move the Go toolchain pins to 1.26, in CI and in every image
2026-09-10 09:58:08 -07:00
jcoffey-dev c60028aad1 Move the Go toolchain pins to 1.26, in CI and in every image
Two Dependabot PRs are stuck behind the same number.

#35 raises the go directive to 1.26.0 in six modules, because
golang.org/x/crypto v0.56.0 requires it -- x/crypto tracks the two most
recent Go releases and 0.56 dropped 1.25. A module that says 1.26 cannot
be built by the 1.25 this repository pins in two places, so that PR
fails every Go job.

#29 raises actions/setup-go to v7, which sets GOTOOLCHAIN=local. With
that set, `go install golang.org/x/vuln/cmd/govulncheck@latest` cannot
quietly fetch a newer toolchain, and stops with

  golang.org/x/[email protected] requires go >= 1.26.0 (running go 1.25.14)

Under setup-go v5 the same install succeeded by downloading 1.26 behind
our backs, which is its own reason to be on 1.26 deliberately instead.

So: security-scan's go-version and all eight Dockerfiles move together,
1.25 -> 1.26. Nothing else needs to. A newer toolchain builds an older
directive happily, so this stands on its own before #35 lands, and the
go.mod files stay where they are here.

Checked by building rather than by reading: the api and ingest images
both build on golang:1.26-alpine, and api, ingest and enterprise still
`go build ./...` clean against their existing 1.25 directives.
2026-09-10 09:54:45 -07:00
dependabot[bot] 585305126b Bump quick-xml from 0.41.0 to 0.42.0 in /agent
Bumps [quick-xml](https://github.com/tafia/quick-xml) from 0.41.0 to 0.42.0.
- [Release notes](https://github.com/tafia/quick-xml/releases)
- [Changelog](https://github.com/tafia/quick-xml/blob/master/Changelog.md)
- [Commits](https://github.com/tafia/quick-xml/compare/v0.41.0...v0.42.0)

---
updated-dependencies:
- dependency-name: quick-xml
  dependency-version: 0.42.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 16:50:34 +00:00
jcoffey 907bf52541 Merge pull request #37 from Coffey-Labs/dependabot/npm_and_yarn/web/npm-minor-and-patch-3c6815572c
Bump the npm-minor-and-patch group in /web with 6 updates
2026-09-10 09:48:59 -07:00
jcoffey 0b37199615 Merge pull request #33 from Coffey-Labs/dependabot/cargo/agent/windows-service-0.8.1
Bump windows-service from 0.7.0 to 0.8.1 in /agent
2026-09-10 09:48:54 -07:00
jcoffey 75303930a9 Merge pull request #32 from Coffey-Labs/dependabot/cargo/agent/toml-1.1.5spec-1.1.0
Bump toml from 0.8.23 to 1.1.5+spec-1.1.0 in /agent
2026-09-10 09:48:51 -07:00
jcoffey 0670ff55b8 Merge pull request #30 from Coffey-Labs/dependabot/cargo/agent/windows-0.62.2
Bump windows from 0.58.0 to 0.62.2 in /agent
2026-09-10 09:48:47 -07:00
jcoffey be1840a94f Merge pull request #40 from Coffey-Labs/funding-username-jcoffey-dev
Point the Sponsor button at the current GitHub username
2026-09-10 09:19:32 -07:00
jcoffey-dev df6f9d049f Point the Sponsor button at the current GitHub username
The account behind it was renamed from LINUXexpert-org to jcoffey-dev,
and GitHub does not redirect the old name: github.com/sponsors/
LINUXexpert-org answers 404 while the new one answers 200. So the
Sponsor button on this repository has been leading nowhere.

Worth fixing rather than leaving to redirect, because a released
username can be registered by anyone -- a stale link stops being a dead
end and starts being someone else's page.
2026-09-10 09:17:02 -07:00
dependabot[bot] c0f0d51887 Bump the npm-minor-and-patch group in /web with 6 updates
Bumps the npm-minor-and-patch group in /web with 6 updates:

| Package | From | To |
| --- | --- | --- |
| [@codemirror/commands](https://github.com/codemirror/commands) | `6.10.4` | `6.11.0` |
| [@codemirror/state](https://github.com/codemirror/state) | `6.7.1` | `6.7.4` |
| [@codemirror/view](https://github.com/codemirror/view) | `6.43.8` | `6.43.11` |
| [@sveltejs/kit](https://github.com/sveltejs/kit/tree/HEAD/packages/kit) | `2.70.2` | `2.70.3` |
| [svelte](https://github.com/sveltejs/svelte/tree/HEAD/packages/svelte) | `5.56.9` | `5.57.0` |
| [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) | `8.2.1` | `8.2.2` |


Updates `@codemirror/commands` from 6.10.4 to 6.11.0
- [Changelog](https://github.com/codemirror/commands/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/commands/commits)

Updates `@codemirror/state` from 6.7.1 to 6.7.4
- [Changelog](https://github.com/codemirror/state/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/state/commits)

Updates `@codemirror/view` from 6.43.8 to 6.43.11
- [Changelog](https://github.com/codemirror/view/blob/main/CHANGELOG.md)
- [Commits](https://github.com/codemirror/view/commits)

Updates `@sveltejs/kit` from 2.70.2 to 2.70.3
- [Release notes](https://github.com/sveltejs/kit/releases)
- [Changelog](https://github.com/sveltejs/kit/blob/version-3/packages/kit/CHANGELOG.md)
- [Commits](https://github.com/sveltejs/kit/commits/@sveltejs/[email protected]/packages/kit)

Updates `svelte` from 5.56.9 to 5.57.0
- [Release notes](https://github.com/sveltejs/svelte/releases)
- [Changelog](https://github.com/sveltejs/svelte/blob/main/packages/svelte/CHANGELOG.md)
- [Commits](https://github.com/sveltejs/svelte/commits/[email protected]/packages/svelte)

Updates `vite` from 8.2.1 to 8.2.2
- [Release notes](https://github.com/vitejs/vite/releases)
- [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md)
- [Commits](https://github.com/vitejs/vite/commits/v8.2.2/packages/vite)

---
updated-dependencies:
- dependency-name: "@codemirror/commands"
  dependency-version: 6.11.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: npm-minor-and-patch
- dependency-name: "@codemirror/state"
  dependency-version: 6.7.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: npm-minor-and-patch
- dependency-name: "@codemirror/view"
  dependency-version: 6.43.11
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: npm-minor-and-patch
- dependency-name: "@sveltejs/kit"
  dependency-version: 2.70.3
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: npm-minor-and-patch
- dependency-name: svelte
  dependency-version: 5.57.0
  dependency-type: direct:development
  update-type: version-update:semver-minor
  dependency-group: npm-minor-and-patch
- dependency-name: vite
  dependency-version: 8.2.2
  dependency-type: direct:development
  update-type: version-update:semver-patch
  dependency-group: npm-minor-and-patch
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 16:07:03 +00:00
dependabot[bot] eb48ee0aa5 Bump windows-service from 0.7.0 to 0.8.1 in /agent
Bumps [windows-service](https://github.com/mullvad/windows-service-rs) from 0.7.0 to 0.8.1.
- [Release notes](https://github.com/mullvad/windows-service-rs/releases)
- [Changelog](https://github.com/mullvad/windows-service-rs/blob/main/CHANGELOG.md)
- [Commits](https://github.com/mullvad/windows-service-rs/compare/v0.7.0...v0.8.1)

---
updated-dependencies:
- dependency-name: windows-service
  dependency-version: 0.8.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 16:06:35 +00:00
dependabot[bot] 3db8850598 Bump toml from 0.8.23 to 1.1.5+spec-1.1.0 in /agent
Bumps [toml](https://github.com/toml-rs/toml) from 0.8.23 to 1.1.5+spec-1.1.0.
- [Commits](https://github.com/toml-rs/toml/compare/toml-v0.8.23...toml-v1.1.5)

---
updated-dependencies:
- dependency-name: toml
  dependency-version: 1.1.5+spec-1.1.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 16:06:31 +00:00
dependabot[bot] df90b85ec4 Bump windows from 0.58.0 to 0.62.2 in /agent
Bumps [windows](https://github.com/microsoft/windows-rs) from 0.58.0 to 0.62.2.
- [Release notes](https://github.com/microsoft/windows-rs/releases)
- [Commits](https://github.com/microsoft/windows-rs/commits)

---
updated-dependencies:
- dependency-name: windows
  dependency-version: 0.62.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <[email protected]>
2026-09-10 16:06:20 +00:00
Coffey Labs bb42a07392 Merge pull request #28 from Coffey-Labs/grpc-1.83.2
Take grpc to 1.83.2 across the nine modules that carry it
2026-09-10 09:04:46 -07:00
jcoffey-dev b7f99b49f2 Take grpc to 1.83.2 across the nine modules that carry it
GHSA-2v4p-qf9q-27wj is a panic in gRPC-Go's xDS routing interceptor: a
request arriving with neither `:authority` nor `Host` indexes an empty
slice, the per-RPC goroutine does not recover, and the process dies.
High, and nine alerts, because nine go.mod files pin the same version --
eight directly, terraform indirectly.

Nothing here was reachable. The interceptor is installed by
`xds.NewGRPCServer`, which this repo never calls: the one production
server is `grpc.NewServer(grpc.Creds(...))` in ingest/internal/grpcserver
and the only other is a plain one in a searchclient test. That is also
why security-scan has been green throughout -- govulncheck reports on
reachability and found nothing on 1.83.1, while Dependabot reports on
version ranges and found nine. Both were right.

Taken anyway: it is a patch release, and the next advisory in this
dependency may well land somewhere we do reach.

`go mod tidy` carried the indirect requirements grpc 1.83.2 asks for --
x/net, x/text, x/sys and friends. No CI job builds or tests Go here, so
all nine modules were built locally and api, ingest and enterprise
tested with -count=1, since a cached pass would not have exercised the
new version.

The dependabot.yml is the other half. There was no config, so nothing
opened a PR against any of this. Go majors stay out of the group, being
import path changes rather than bumps.
2026-09-10 09:01:27 -07:00
Coffey Labs 6a17fa1562 Update GitHub Sponsors username in FUNDING.yml
Signed-off-by: Coffey Labs <[email protected]>
2026-09-05 00:51:06 -07:00
Coffey Labs 91dff9cc46 Merge pull request #27 from Coffey-Labs/chore/upgrade-tantivy
Upgrade tantivy to 0.26 and clear the lru advisory
2026-09-04 22:19:22 -07:00
jcoffey-dev a53c309bad Upgrade tantivy to 0.26 and clear the lru advisory
Dependabot #10: lru's IterMut violates Stacked Borrows, fixed in 0.16.3.
lru was transitive through tantivy 0.22.1, which pins lru ^0.12.0, so
there was no in-range fix -- cargo update -p lru locks nothing. The
advisory was also not reachable: tantivy calls only get, put, len,
peek_lru and new on its LruCache, never iter_mut. Upgrading rather than
dismissing because it is early enough that carrying four versions of
drift costs more than paying it now, and the alert then closes on its
own evidence rather than on an argument.

lru resolves to 0.16.4, past the patch line.

One API change across the four releases. TopDocs no longer implements
Collector on its own -- an ordering has to be chosen rather than
defaulted into. order_by_score() is exactly what bare TopDocs did in
0.22, so result order is preserved rather than quietly changed, which
matters for a search endpoint whose contract is "most relevant first".

Index compatibility checked rather than assumed, since a format change
would have meant a reindex for every existing deployment. Against a live
index of 1757 documents written by 0.22: the service opened it without
error, a document indexed hours earlier by 0.22 is still findable, new
documents written by 0.26 are findable, and a phrase query spans both.
No migration needed.

The compliance inventory is regenerated for the new graph: 17 crates
added, 5 gone, 20 bumped, and three duplicate-version entries collapsed
where the graph no longer needs two. Nothing newly flagged -- every
addition is permissive -- and cargo-deny check licenses, which is the
gate CI actually runs, passes. Rows for crates whose name and version
are unchanged are left byte-for-byte alone, so the diff shows the real
change rather than 200 rows of SPDX term reordering.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 22:17:28 -07:00
Coffey Labs bd49fb7685 Merge pull request #26 from Coffey-Labs/docs/phase-8-aggregate-count
Specify what aggregate_count emits
2026-09-04 22:01:21 -07:00
jcoffey-dev ec4b860ba8 Specify what aggregate_count emits
The last unanswered action, and the only one whose output is not the
input with edits -- it emits a record that never existed, which is why
it was deferred twice.

It emits the window's first record unchanged, tagged with
cairnobs.aggregated, cairnobs.count, and the observed window bounds.
That follows the convention the agent already uses for heartbeat and
host-metrics records rather than inventing a second synthetic-record
mechanism, and keeping the first record intact means a reader sees a
real example of what was collapsed instead of an invented summary.

It tags even when the count is one. Emitting a bare record there would
be tidier and would make cairnobs.count present only sometimes, so
summing it silently breaks on quiet windows. window_last is the last
record that actually contributed, never window_start + window_ms,
because a window flushed early must not claim an end that never
happened.

Specifying it surfaced a problem the other nine actions do not have.
Windows are measured on record time, so a window can only be closed by a
later record arriving. suppress_duplicates never has anything pending;
aggregate_count holds state, so a matching stream that goes quiet leaves
its aggregate unemitted indefinitely -- data loss dressed as latency.
Emission therefore has a second trigger, end of stream, which the corpus
defines as an implicit flush after the last input and which production
gets from the batch flush. The cost is stated rather than hidden:
window_ms becomes a maximum, not a guarantee, and one burst can produce
more than one aggregate.

And it has a consequence nobody should meet in production first: stats
count undercounts aggregated data silently, so every panel and alert
counting rows changes meaning the moment a rule aggregates the data
behind it. Nothing here fixes that. The correct idiom is summing
cairnobs.count; teaching the query layer to do it automatically is a
Phase 2 change to the IR, recorded as the open question this decision
leaves in its place rather than quietly inherited.

Six cases added, corpus at 45. The validator's unspecified-action guard
stays in place with an empty set, still rejecting anything added to it.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 21:58:40 -07:00
Coffey Labs 17256634db Merge pull request #25 from Coffey-Labs/docs/phase-8-apply-then-verify
Put apply-then-verify in v1, and specify it
2026-09-04 21:49:29 -07:00
jcoffey-dev c2f74b4e1a Put apply-then-verify in v1, and specify it
It was a defensible cut while a canary might have caught a bad rollout
partway through. With no canary it is the only thing that recovers a
host without an operator noticing, and "total evaluation should make it
unreachable" is what every crash-loop was before it happened.

Specified rather than named. A two-field file is written and fsynced
before rules are applied: the version being attempted, and whether it is
trying, good or quarantined. The rule set itself is never written, which
is what keeps the crash-loop-not-strand property -- an agent that loses
the file still boots clean and re-syncs. A version found still "trying"
at boot is quarantined, the agent starts with no rules, and it reports
the quarantined version so the failure is visible rather than merely
survived. A different version clears the quarantine, because pushing new
rules is the correction.

Three failure modes decided instead of discovered. An agent that cannot
write the file logs once and runs with total evaluation alone: degrading
to "no rules" would punish every read-only deployment for a failure that
has not happened, and the backstop is best-effort by construction rather
than by accident. An agent killed for an unrelated reason quarantines a
blameless rule set, which is a deliberate false positive -- the
alternative is claiming to distinguish "died because of the rules" from
"died while they happened to be loaded", which it cannot do honestly.
And a rule set fatal on only some hosts quarantines per host, which is
the closest thing to a canary this design has: the first host to hit it
reports while the rest carry on.

Also states that the conformance corpus cannot test any of this, so
nobody tries. The corpus pins rule semantics -- records in, records out.
Process death and file state across restarts are not expressible that
way and need agent-side tests driving a real process through crash and
restart.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 21:02:35 -07:00
Coffey Labs 244bfbd540 Merge pull request #24 from Coffey-Labs/docs/phase-8-decisions-regex-and-rollout
Settle Phase 8's regex and rollout questions
2026-09-04 20:58:55 -07:00
jcoffey-dev 02ebbc86a1 Settle Phase 8's regex and rollout questions
regex-lite on the agent, full regex at ingest, conformance corpus
limited to the syntax both accept. The agent had no regex dependency at
all and the full crate is a megabyte-plus against a binary whose pitch
is that it is small and static. The rejected alternative, regex
ingest-side only, would have given up redacting PII before it leaves the
host -- the one capability that most needs to be on the agent, and most
of the reason on-host processing exists at all.

Checked rather than assumed: every pattern the corpus uses compiles and
behaves under regex-lite, including named captures and replace_all
replacing every occurrence, which is what mask needs and what a
redaction stopping at the first hit would get wrong. Both engines also
resolve alternation leftmost-first, confirmed by running the same
pattern through regex-lite and Go's regexp and getting the same output.
A case now pins that, since it is the kind of semantic two independent
implementations can differ on silently.

No canary gate. An edit reaches every matching agent at once, and this
design had suggested that might have to block the feature. It does not,
but the risk is accepted rather than waved away: total evaluation is the
real mitigation, a fatal rule set crash-loops rather than strands
because overrides are never persisted, apply-then-verify makes that
self-healing, and applied_override_version already shows the blast
radius. What is genuinely given up is the ability to stop a bad rollout
partway through -- everything else shortens the outage without
preventing the rule reaching every host first.

That raises the stakes on apply-then-verify, which is still open.
Without a canary, total evaluation stops being a nice property of a
well-designed DSL and becomes the only thing between a bad rule and
every host at once, so cutting the backstop from v1 is a harder call
than it looked when it was written down.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 20:57:02 -07:00
Coffey Labs bd0d71c45b Merge pull request #23 from Coffey-Labs/feat/processing-conformance-suite
Build the Phase 8 conformance corpus
2026-09-04 20:53:43 -07:00
jcoffey-dev 7fabc6a067 Build the Phase 8 conformance corpus
The design argues the conformance suite is the specification and should
be built before either implementation, since two hand-written
implementations of one language diverge unless something shared pins
them. This is that suite: 38 cases in /processing, a language-neutral
top-level directory for the same reason /proto is one -- the Rust agent
and the Go ingest tier both consume the definition and neither owns it.

Nothing executes the cases, because neither implementation exists. A
stdlib-only validator checks the corpus stays well-formed and runs in
CI: known actions, addressable fields, compilable patterns, names
matching filenames, and no case depending on record_id, which is
withheld so the question of whether an ingest-side rule can see one
stays open. The validator was checked against seven deliberately broken
cases before being trusted, since "38/38 valid" means nothing from a
validator that cannot fail.

Writing the cases first has already paid for itself twice.

It forced two determinism decisions the prose had left vague, both of
which a conformance suite cannot avoid answering. Sampling is
counter-based rather than random: random is statistically nicer and
impossible to assert on. Windows are measured on record timestamps
rather than wall-clock, which makes replay deterministic and, not
incidentally, makes backfill behave correctly where a wall-clock window
would not.

And it made the missing aggregate_count answer concrete. The design
does not say what that action emits, or what a query not expecting a
synthetic record sees. Rather than invent one by writing cases, the
validator rejects any case using it, so the design question has to be
answered before the behaviour can be frozen by accident.

The corpus includes the acceptance case from real measured data: the
two processes that account for roughly 60% of a real workstation's
journal volume, and the one kernel message worth keeping.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 20:19:51 -07:00
Coffey Labs 8787c1d087 Merge pull request #22 from Coffey-Labs/docs/phase-8-processing-design
Draft the Phase 8 processing design
2026-09-04 20:11:53 -07:00
jcoffey-dev 7cc2fd8c78 Draft the Phase 8 processing design
Starts from the distribution channel and the safety invariant it
protects, and derives the language from them, rather than designing a
rule language and asking later how to ship it.

Four decisions proposed. A rule is a matcher plus ordered typed actions,
with no expressions and nothing resembling eval -- less expressive than
Cribl on purpose, and the only shape that can be pushed to ten thousand
hosts and audited by reading it. Total evaluation is the primary safety
guarantee, with apply-then-verify as a backstop. One spec with two
implementations means a language-neutral conformance suite is the
specification and should be built first, not last. Distribution reuses
DesiredOverride, following extra_file_paths as the precedent for a
repeated field.

It also corrects something #21 got wrong. That change said a fatal rule
would strand an agent the way a corrupted ingest endpoint does. Reading
apply_override's actual semantics, overrides live only in the running
process's memory and are never written to disk, so a restarted agent
boots clean and re-syncs -- a fatal rule set crash-loops rather than
strands, and the agent keeps checking in, so it stays correctable. That
is a much better failure mode, and it was acquired by accident: the
"don't persist" choice was made to avoid filesystem writes on read-only
images, not for safety. This design promotes it to a constraint, since
persisting overrides later would silently convert every crash-loop into
a strand. positioning.md is corrected to match rather than left
disagreeing.

Five open questions are left open rather than answered to look decisive,
the sharpest being that there is no staged rollout today: an edit
reaches every matching agent at once, which for executable rules is the
difference between breaking one host and breaking all of them.

The v1 acceptance test is real data, not a fixture: two processes on the
maintainer's own workstation account for 308 of 325 journal entries in
five minutes, and a suppress rule should remove about 60% of that host's
volume.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 20:09:00 -07:00
Coffey Labs b5a3ff2b6d Merge pull request #21 from Coffey-Labs/docs/roadmap-realignment
Realign the roadmap with what is already built
2026-09-04 19:20:25 -07:00
jcoffey-dev 18a55ccd5a Realign the roadmap with what is already built
Three of the roadmap's claims were contradicted by the repository
itself.

Fleet management was Phase 11, "Planned", and positioning.md said config
still flowed to the agent from the host rather than from the platform.
agent-management-design.md has recorded the opposite for some time:
punch list complete, verified live, with central authoring, versioning,
rollout on the next check-in, observation, and a restart command a real
agent picks up and acts on. There is no Phase 11 now. Its remainder is
either already named there -- stop/uninstall, per-host multi-row
alerting, a rule-per-host generator -- or belongs to Phase 8, since
distributing rules is the one genuinely new thing the mechanism has to
carry, and a rule language nobody can push to a fleet is not worth
having.

That inverts the old ordering argument, which put fleet last on the
grounds that it manages configuration the earlier phases define. Sound
reasoning; the world went the other way and built the mechanism first.
Recorded rather than quietly dropped, because the instinct behind it is
a good one that happened not to apply.

Retention was listed as a question Phase 10 would finally have to
answer. Half of it is answered: api/logretention serves operator-driven
preview and delete with an owner-only per-agent floor. What is missing
is an automatic TTL, so Phase 10 owns tiering and automatic TTL rather
than retention from nothing.

The rule-language recommendation is now settled rather than proposed,
and not on its own authority: DesiredOverride is already a closed typed
shape that cannot carry arbitrary code, so no channel exists that would
deliver JavaScript to an agent even if the language argument had gone
the other way.

That surfaced a requirement nothing had written down. Agent management
rests on an invariant it states outright -- every editable field
degrades behaviour without cutting off the agent's ability to receive
the next correction. Processing rules break it: a rule that panics or
loops strands the agent exactly the way a corrupted ingest endpoint
would, across every host it reached first. Phase 8 now owes either total
evaluation or apply-then-verify with rollback, chosen deliberately
rather than discovered mid-rollout.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 19:16:01 -07:00
Coffey Labs 9e1efba38f Merge pull request #19 from Coffey-Labs/docs/tenancy-off-the-roadmap
Take multi-tenancy off the roadmap
2026-09-04 19:01:57 -07:00
jcoffey-dev 0ee2e9183b Take multi-tenancy off the roadmap
Cairn OBS is self-hosted, and the way to separate two environments is to
run two installations rather than two tenants inside one. Tenancy is the
wrong boundary for that, on three counts this repository demonstrates
rather than assumes: chwriter.WriteBatch is all-or-nothing across
tenants, so one tenant's failure stalls offset progress for every other;
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT puts every tenant's data behind a
single superuser credential, as docker-compose.yml's own comment says;
and one binary with one set of migrations moves every tenant together,
which is the opposite of what separate environments are for. A whole
installation idles at about 1.3 GB, so the sharing buys nothing.

The project led with multi-tenant RBAC in the README banner and in
PROJECT-SPEC's goal statement. Both now say what it is instead:
self-hosted. "Open-core" goes with them -- it was already inaccurate,
since CONTRIBUTING states there is no feature gate and no paid tier, and
with enterprise/ off the roadmap there will not be one.

A second identity provider comes off the list of things standing between
this and production-ready. SSO belongs to enterprise/, and a self-hosted
deployment is not waiting on it. Terraform's tenant/RBAC resources move
from "disclosed future work" to not planned.

Nothing is scrubbed from the record. Phase 4 stays shipped, its runbook
stays, and its known gaps stay stated -- rewriting that history would
contradict the candour the Status section is built on. enterprise/ stays
in the tree, AGPLv3 and working, as the answer to a question this
project is not asking.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 18:31:23 -07:00
Coffey Labs f644825692 Merge pull request #17 from Coffey-Labs/feat/local-login-compose-and-docs
Make local login reachable, and write down how it works
2026-09-04 18:02:50 -07:00
Coffey Labs 10a7c2fee6 Merge pull request #16 from Coffey-Labs/docs/query-api-request-shape
Catch the runbooks up with the query API they describe
2026-09-04 18:02:41 -07:00
Coffey Labs 6a45832d80 Merge pull request #15 from Coffey-Labs/fix/nav-local-auth-hidden
Show the account controls the deployment actually has
2026-09-04 18:02:31 -07:00
Coffey Labs 81ec542452 Merge pull request #14 from Coffey-Labs/fix/self-delete-lockout
Stop an owner deleting the account they are signed in as
2026-09-04 18:02:13 -07:00
jcoffey-dev a24860ac2d Make local login reachable, and write down how it works
Local login is implemented, wired through api, alerting and web, and
undiscoverable. No compose file turns it on, the Helm chart sets none
of its variables, and no markdown in the repository mentions
-seed-admin, LOCAL_AUTH_ENABLED or local login at all. The only way to
find it is to read cmd/api/main.go's authorizer switch.

Enabling it in docker-compose.yml is not the answer: a plain
`docker compose up` has no authentication, and every Phase 0-3 runbook
verifies the pipeline with bare curl against /query. Turning login on
by default would break the project's own documented verification.
So it's an opt-in overlay instead.

Four settings have to agree, and only one of them is obviously about
login. Each fails differently and none of the failures name the cause:
the route 404s, or the browser refuses the request before sending it,
or login returns 200 and every later request is anonymous because the
cookie was never stored, or the same symptom again from the opposite
end because the bundle never attaches it. That is what the new document
is mostly for.

The Helm chart still has no local-login support. Recorded in the
document as a gap rather than papered over.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 17:48:20 -07:00
jcoffey-dev 9bbd802a91 Catch the runbooks up with the query API they describe
Phase 2 unified the two query languages behind one endpoint and renamed
the request field, and the runbooks were never updated. Following them
today does not work:

  {"sql": ...}                 -> 400 query must not be empty
  POST /api :8080/search       -> 404, the route no longer exists

Both appear in the Phase 0 and Phase 1 runbooks and in the
windows-fixture README. That matters more than a normal doc typo,
because status.md cites the Phase 0 runbook as the record of how Phase
0 was verified -- so the documented verification procedure is one
nobody can re-run as written.

The Phase 1 step is rewritten rather than search-and-replaced: it
checked the SQL and full-text paths against two different endpoints,
and its exit criterion (the same record_id from both) now has to be
expressed against /query twice, once with SQL and once with a bare
word.

Phase 0's expected output for SELECT 1 also gained a warnings field
since it was written.

Every command here was run against a live stack before being written
down, including confirming both paths return the same record_id.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 17:44:54 -07:00
jcoffey-dev dd5d5a77c3 Show the account controls the deployment actually has
The sidebar decided which auth mode was live from enterpriseAuthBase,
so any deployment with VITE_ENTERPRISE_AUTH_BASE_URL set rendered the
enterprise block -- and compose sets it unconditionally, so the tenant
picker can exist. On a single-tenant stack with local login on, that
meant the local block could never render: no username, no role, no
Change password, no Log out, and in their place a "Sign in" link
pointing at enterprise-auth's OIDC route, which is disabled unless
OIDC_ISSUER_URL is configured. A dead link where the account controls
should be.

The build-time flag was never the right thing to ask. api registers
/auth/* only when LOCAL_AUTH_ENABLED is set and ENTERPRISE_AUTH_URL is
not, so the frontend cannot know the mode from its own build args --
the two can disagree, and here they did. getLocalSession already
distinguishes 'disabled' (a 404 from /auth/session) from null (a 401,
logged out); the sidebar collapsed both to null and threw the answer
away. It now keeps that distinction and branches on it, so the mode
comes from what the server actually serves.

Logged out under local auth, the sidebar previously rendered no auth
block at all -- no way back to the login page from the nav. It now
offers Sign in, pointing at /login.

Neither block renders until the probe lands, so nothing flashes the
wrong mode on load.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 17:42:26 -07:00
jcoffey-dev 7becb7344d Stop an owner deleting the account they are signed in as
Deleting your own user succeeded, and logged you out doing it:
local_sessions.user_id is ON DELETE CASCADE, so the delete took the
caller's own live session with it. Nothing refused this. The
last-owner guard is the only thing in the path, and it passes cleanly
as soon as a second owner exists -- which is exactly the state you are
in just after creating one.

The way back in was then whatever other account happened to exist, and
-seed-admin could not help: it skipped whenever *any* local user was
present, so the command documented as the way to create an
administrator refused precisely when there was no usable one, because
some other account still existed. It now asks whether the admin
account itself is missing, which is what its own help text always
claimed, and what makes it useful as recovery rather than only as
first-run bootstrap.

TestCanDeleteAnOwnerWhenAnotherRemains signed in as admin1 and deleted
admin1, asserting 204 -- it encoded the lockout as intended behaviour.
It now deletes the other owner, which is what it meant to cover, and a
new test holds the refusal in place.

runSeedAdmin takes a small interface so the bootstrap path is tested
without a Postgres pool; it had no tests before.

Signed-off-by: John Coffey <[email protected]>
2026-09-04 17:39:22 -07:00
Coffey Labs 5374c1946a Merge pull request #13 from Coffey-Labs/demo-currency-usd
Price the demo in dollars
2026-09-04 17:02:28 -07:00
jcoffey-dev a2baf2af74 Price the demo in dollars
The commerce generators shipped with currency=GBP on every order, refund,
authorisation and chargeback -- twelve places, all of them a default
nobody chose. Dollars is the convention everywhere else.

Only the label changes. The SKU prices stay as they are: a $489 task
chair and a $629 standing desk are as plausible as the pound figures
were, and moving them would have shifted average order value on the
demo's dashboards for no reason other than tidiness.

Nothing else referenced the currency -- no dashboard panel and no alert
rule filters or groups on it -- so this is the whole change.
2026-09-04 16:59:42 -07:00
Coffey Labs e22b99d11e Merge pull request #12 from Coffey-Labs/positioning-splunk-and-cribl
Position against Cribl and Splunk, and name the local-AI axis
2026-09-04 16:45:27 -07:00
jcoffey-dev 50dc3f7ee3 Write down the naming contrast, because it argues the position
Splunk is from spelunking: caving, a lamp, feeling your way along in the
dark. An honest description of search-driven investigation -- powerful
with expertise, unforgiving without it.

Cribl is from cribble, to sift, from Latin cribrum, a sieve; the same root
gives engraving its maniere criblee, the dotted ground punched into a
plate. Both senses land together: the data is a medium to be worked and
thinned on the way through. A name about the material, not the
destination.

A cairn is a stack of stones on open ground, where the path is not
obvious, doing one job -- somebody came this way, and this is the way.

Three of its properties map onto things this project already does rather
than things it claims. It is left by whoever went first for whoever comes
next, which is the runbook culture and the reason every phase records
what was actually run including the failures. Anyone passing adds to it,
which is AGPLv3 throughout and an egress path that helps data leave.
And you can see it from a distance in daylight, which is a legible query
language, an AI that explains rather than divines, and a plan that
publishes what has not been proven.

Written into positioning.md rather than kept as a marketing note because
it is a reason the position coheres, not decoration on top of it. Also
noted there that it should not be turned into a slogan.
2026-09-04 16:43:01 -07:00
jcoffey-dev 57ffb6697e Frame the Status section so its candour reads as rigour
The Status section went from a table of "Shipped" straight into three
paragraphs of caveats, with nothing in between telling a reader what
standard was being applied. Read cold, that is a project confessing. Read
with the standard stated first, it is a project that refuses to call
something done because the tests pass.

So the section now opens by saying what stage this is -- pre-1.0, no
production workload -- and what "shipped" means here: verified against
real infrastructure with a runbook recording how, including what the
verification could not reach. And it says plainly why the caveats are
long, which is that they are disclosed rather than discovered. A shorter
Status section would not mean a more finished product, only a less
careful one.

It closes with what would actually close the gap: a second IdP, a real
cluster, the Windows agent on Windows, sustained load, and somebody
else's data. None of that is research, it is time on real infrastructure
-- which makes it the list a pilot works through, and makes a 1.0 tag the
wrong next milestone to reach for.

Written because the risk of a public repo at this stage is not a
competitor reading the roadmap, it is a prospect reading unusual honesty
as immaturity. The fix for that is context, not privacy.
2026-09-04 16:39:54 -07:00
jcoffey-dev d49ddb943e Add the AI axis: plain English as an option, analysis as the end state
Cost is the argument against Splunk and control is the argument against
Cribl. AI is the third, and the difference there is not a feature
comparison, it is where the model runs.

Plain-English querying shipped in Phase 7 and stays an option rather than
a replacement for writing a query: every generated query compiles through
the same IR and executor as a hand-written one, with the same tenant
scoping, cost guardrails and audit logging. The model suggests, it does
not get a private path to the data.

AI-assisted analysis and explanation is the end state and is not built.
Authoring answers "how do I ask this"; the valuable question is "what
does this mean" -- what changed in a result set, why an alert fired and
what preceded it, summarising an incident from the records around it.
Recorded on the status page as an end-state goal rather than a numbered
phase, because it is a property the product keeps rather than a thing to
finish and tick off.

Local is the non-negotiable part, and it is worth stating as position
rather than as a bullet: the default runs qwen2.5-coder through Ollama on
the customer's own hardware, Apache-2.0 weights chosen so Phase 6's
licence work survives contact with the model, and the cloud adapter is
opt-in and off by default. Logs are the most sensitive unstructured data
most organisations hold -- credentials in stack traces, customer
identifiers, internal topology -- so an assistant that reads them is
either running where the data already is, or it is a data-egress decision
wearing a helpful interface.

The constraint it imposes is stated too, because it bounds what can be
promised: a 7B model on a customer's hardware will not match a frontier
model, and the honest claim is not that it is as clever but that it is
good enough at a bounded task and runs somewhere you control. Analysis
features have to be designed to that budget rather than assuming an API
is one call away.
2026-09-04 16:35:44 -07:00
jcoffey-dev e57486add5 Position against Cribl as well as Splunk, and say what that costs us
Splunk and Cribl are not the same competitor and the claim is not the
same claim twice. Splunk is the destination and Cairn OBS replaces it,
which is what phases 0-7 were for. Cribl is the road: routing, reduction,
enrichment, redaction and replay on the way to wherever data is going.
Cairn OBS is already a road in shape -- agent, Redpanda, ingest -- and
exposes none of a pipeline's controls. The agent cannot filter, drop,
sample, mask or re-route anything; ingest normalises a schema and writes
it; there is exactly one destination and it is us.

Reconciling the two turns up something a cost-led project has to face
rather than paper over: most people buy Cribl because Splunk is expensive
per gigabyte, so being genuinely cheap per gigabyte removes the main
reason to buy Cribl in front of us. That makes the strongest pitch "one
system where there were two" rather than "we are also a pipeline vendor"
-- but that pitch only survives a buyer if we also do the four things
people buy a pipeline for that are not about spend: routing to several
destinations, redacting before data leaves the network, archive and
replay, and not being locked to one analytics vendor. Those are about
control, which is better ground anyway: cost advantages get matched and
architectural ones do not.

The consequence is uncomfortable and is written down as a decision rather
than left to be discovered: competing with Cribl means being able to send
data to S3, Splunk HEC, Elastic, OTLP and Kafka -- building features whose
purpose is to help data leave this platform. A project that refuses
lock-in in its licence and then builds it into its egress would be lying
about itself.

Four phases follow, ordered so each pays for itself: processing (8),
routing (9), archive and replay (10), fleet (11). Processing without
routing still shrinks what is stored; routing without processing forwards
everything and helps nobody.

One design decision is called out now because it collides with a
non-negotiable constraint. Cribl's rule language is JavaScript, and
embedding a JS engine in a statically-linked musl agent would end "no
glibc runtime deps" as a claim. The recommendation is a declarative rule
DSL -- matchers and typed actions, no arbitrary code -- deliberately less
expressive, small enough to audit and safe to push to ten thousand hosts.

The retention/TTL question in architecture.md is no longer deferrable and
now says so: Phase 10 asks it from the other side.
2026-09-04 16:29:29 -07:00
Coffey Labs 8660576eb1 Merge pull request #11 from Coffey-Labs/demo-alerts-and-dashboards
Add two storefronts, a payment gateway, and alerts worth waking up for
2026-09-04 16:16:03 -07:00