Sweeps the references that carry no runtime coupling, and fixes one that
turned out to be a real bug rather than stale branding.
Docker network: sentry_default -> cairnobs_default across 23 runbook and
test-header `docker run` commands. Compose derives the network from the
directory name, so this lands together with renaming the working copy to
cairnobs/ -- the two are only correct as one change.
Stale references corrected: four Dockerfile "repo root (sentry/)"
headers; .env pointing at the long-renamed deploy/helm/sentry/ chart;
five Helm comments describing the topic as sentry.logs.raw when all four
code paths have defaulted to cairnobs.logs.raw for some time; an
absolute /home/john/Projects/sentry/ path in the operator's package doc,
now repo-relative; the hand-written Tenant CRD description in both of
its identical copies, whose Go source already said Cairn OBS.
Migration 0043 repoints the default tenant's data source. 0026 seeded it
with ('sentry', '/var/lib/sentry-search') to match what
api/internal/config then defaulted to; the rebrand later moved those
defaults to "cairnobs" and /var/lib/cairnobs-search without moving the
already-applied row, leaving the default tenant naming a ClickHouse
database nothing writes to. Scoped to the exact stale values so it is a
no-op on any deployment that set them deliberately. 0026's comment is
annotated as superseded; its applied SQL is untouched.
Deliberately not included: the gRPC wire packages (sentry.logs.v1,
sentry.agent.v1) and proto/sentry/ import paths, which cannot change
without a lockstep agent/server upgrade; the Helm chart's
sentry_metadata database and sentry role, which need a real Postgres
migration on existing deployments; and the compliance audit records in
docs/compliance/, which are a dated historical record.
go build, go vet, and go test pass for ingest and deploy/operator.
43 lines
1.9 KiB
Markdown
43 lines
1.9 KiB
Markdown
# alert-load-test
|
|
|
|
Seeds a realistic number of concurrent alert rules via `/alerting`'s real
|
|
create API (not a direct DB insert) and measures whether the evaluator's
|
|
claim scheduling keeps up under load. See
|
|
`/docs/phase-3-alerting-design.md`'s "Load-testing plan" and
|
|
`/docs/phase-3-runbook.md` for the methodology and real measured results.
|
|
|
|
```sh
|
|
# 1. Push real data so rule queries have real work to do (reuses
|
|
# hack/benchmark-fixture):
|
|
cd ../benchmark-fixture
|
|
go run . --count 500000
|
|
|
|
# 2. Run a webhook-sink so the (never-firing, by design) rules have a
|
|
# valid notification target to point at:
|
|
docker run -d --name cairnobs-webhook-sink --network cairnobs_default \
|
|
-p 9099:9099 -v $(pwd)/../webhook-sink:/src -w /src golang:1.25-alpine go run .
|
|
|
|
# 3. Run the load test:
|
|
cd ../alert-load-test
|
|
go run . --rule-count 500 --eval-interval-seconds 60 --duration 3m30s
|
|
```
|
|
|
|
Each rule queries a different host's count over the last minute
|
|
(`earliest=-1m host="host-01" | stats count`) against real ClickHouse
|
|
data, with `threshold_value` set unreachably high so rules stay `ok` --
|
|
this isolates evaluator/ClickHouse scheduling throughput from
|
|
delivery-worker load (a query that never fires still exercises the exact
|
|
same claim → `/query` → evaluate → `ApplyTransition` path every tick).
|
|
|
|
The report shows, per rule, the observed intervals between consecutive
|
|
`last_evaluated_at` changes (polled at `--poll-interval`), compared
|
|
against the configured `eval_interval_seconds`. A real, significant
|
|
finding from actually running this: the evaluator's claim batch size and
|
|
worker-pool concurrency limit defaulted to the same number (20), so 500
|
|
rules all due at once took 125s to cycle through instead of the
|
|
configured 60s. Fixed by separating `EVALUATOR_CLAIM_BATCH_SIZE` from
|
|
`EVALUATOR_WORKER_POOL_SIZE` (see `alerting/internal/config/config.go`).
|
|
|
|
Cleans up the seeded rules and notification target on exit unless
|
|
`--no-cleanup` is passed.
|