Files
cairnobs/alerting
jcoffey-dev 7a86008062 Complete the low-risk half of the Sentry -> Cairn OBS rebrand
Sweeps the references that carry no runtime coupling, and fixes one that
turned out to be a real bug rather than stale branding.

Docker network: sentry_default -> cairnobs_default across 23 runbook and
test-header `docker run` commands. Compose derives the network from the
directory name, so this lands together with renaming the working copy to
cairnobs/ -- the two are only correct as one change.

Stale references corrected: four Dockerfile "repo root (sentry/)"
headers; .env pointing at the long-renamed deploy/helm/sentry/ chart;
five Helm comments describing the topic as sentry.logs.raw when all four
code paths have defaulted to cairnobs.logs.raw for some time; an
absolute /home/john/Projects/sentry/ path in the operator's package doc,
now repo-relative; the hand-written Tenant CRD description in both of
its identical copies, whose Go source already said Cairn OBS.

Migration 0043 repoints the default tenant's data source. 0026 seeded it
with ('sentry', '/var/lib/sentry-search') to match what
api/internal/config then defaulted to; the rebrand later moved those
defaults to "cairnobs" and /var/lib/cairnobs-search without moving the
already-applied row, leaving the default tenant naming a ClickHouse
database nothing writes to. Scoped to the exact stale values so it is a
no-op on any deployment that set them deliberately. 0026's comment is
annotated as superseded; its applied SQL is untouched.

Deliberately not included: the gRPC wire packages (sentry.logs.v1,
sentry.agent.v1) and proto/sentry/ import paths, which cannot change
without a lockstep agent/server upgrade; the Helm chart's
sentry_metadata database and sentry role, which need a real Postgres
migration on existing deployments; and the compliance audit records in
docs/compliance/, which are a dated historical record.

go build, go vet, and go test pass for ingest and deploy/operator.
2026-08-22 18:39:25 -07:00
..
2026-08-21 20:53:32 -07:00
2026-08-21 20:53:32 -07:00

alerting

Alert rule CRUD, the ticker-driven evaluator, and webhook/Slack/PagerDuty delivery. See /docs/phase-3-alerting-design.md for the full design (data model, the ok/pending/firing state machine, and the four correctness properties this implementation follows exactly).

Running

POSTGRES_PASSWORD=cairnobs-dev-only API_QUERY_URL=http://localhost:8080 go run ./cmd/alerting

Talks to the same cairnobs_metadata Postgres database as /api (different tables — see /metadata/README.md), and to /api's POST /query over plain HTTP for rule evaluation. Never connects to ClickHouse or Tantivy directly.

HTTP API

POST   /rules                    create a rule
GET    /rules                    list rules (with current state)
GET    /rules/{id}               get a rule (with current state)
DELETE /rules/{id}
GET    /rules/{id}/deliveries    delivery log for a rule, most recent first

POST   /targets                  create a notification target
GET    /targets
GET    /targets/{id}
DELETE /targets/{id}

GET    /healthz

A rule's condition_type is "threshold" (requires comparator + threshold_value, and the query must resolve to exactly one row) or "absence" (fires when the query returns zero rows in its own earliest=/latest= window — no separate window field). A notification target's kind is "webhook", "slack", or "pagerduty" — all three deliver via the same HTTP POST + retry/backoff mechanism (internal/delivery/webhook.go); slack/pagerduty are payload formatters only, not separate delivery paths.

Environment variables

Var Default
HTTP_LISTEN_ADDR :8081
POSTGRES_ADDR localhost:5432
POSTGRES_DATABASE cairnobs_metadata
POSTGRES_USERNAME cairnobs
POSTGRES_PASSWORD (empty — must be set)
API_QUERY_URL http://localhost:8080
CORS_ALLOWED_ORIGIN *
EVALUATOR_TICK_SECONDS 5 — how often the scheduler checks for due rules
EVALUATOR_CLAIM_BATCH_SIZE 1000 — how many due rules one tick can pull off the queue
EVALUATOR_WORKER_POOL_SIZE 20 — bounded concurrency for /query calls within a claimed batch
EVALUATOR_QUERY_TIMEOUT_SECONDS 30 — per-evaluation POST /query timeout

EVALUATOR_CLAIM_BATCH_SIZE and EVALUATOR_WORKER_POOL_SIZE are deliberately separate knobs, not the same number — see internal/config/config.go's doc comment for the real bug this separation fixes (found by hack/alert-load-test, see /docs/phase-3-runbook.md): with both capped at 20, 500 rules due at once took 125s to cycle through instead of the configured 60s.

Package layout

cmd/alerting/          wires config, Postgres pool, api client; runs the
                        HTTP server + evaluator + delivery worker concurrently (errgroup)
internal/httpapi/      REST handlers -- Handler/RegisterRoutes, same shape as api/internal/dashboards
internal/rulestore/    pgx CRUD for alert_rules + alert_state; ClaimDueRules
                        (fix 1's atomic claim) and ApplyTransition (fix 2's transactional outbox)
internal/notifystore/  pgx CRUD for notification_targets
internal/queryclient/  thin HTTP client to api's POST /query -- no querylang import here
internal/evaluator/    the ticker + worker pool; transitions.go is the pure,
                        exhaustively-tested ok/pending/firing state machine;
                        condition.go implements fixes 3/4 (errors never
                        coerced to "condition false"; threshold zero-rows
                        is an error, not a 0)
internal/delivery/     webhook.go is the claim-and-send worker (all three
                        kinds go through it); slack.go/pagerduty.go are
                        payload formatters only

Building & testing

go build ./...
go vet ./...
go test ./...
docker build -f Dockerfile -t cairnobs-alerting .   # context is alerting/, not the repo root -- no /proto needed