Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat panels via gridstack + uPlot, global + per-panel time range, JSON export/import) and threshold/absence alert rules with an ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery. - New /metadata component: Postgres control-plane store for dashboards, panels, notification targets, alert rules/state, and delivery log -- see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree family isn't a fit for this access pattern (needs real row-level locking and read-your-writes consistency). - api/internal/dashboards: dashboard/panel CRUD, pure -- panel query execution stays client-side, reusing the existing /query endpoint. - New /alerting service: rule/target CRUD, a ticker-driven evaluator (claim-then-evaluate concurrency control, transactional-outbox delivery, query errors and threshold zero-rows never coerced into a false transition) and webhook/Slack/PagerDuty delivery with retry/backoff. See docs/phase-3-alerting-design.md for the full state-machine design and the four correctness properties it implements. - web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts list/get/apply, seeding a future Terraform provider's JSON contract. - hack/alert-load-test: 500 rules against real ClickHouse data, real measured results in docs/phase-3-runbook.md. Five real bugs found by actually running this against a live stack (documented in the runbook, not just fixed silently): a latent Phase 2 bug where ClickHouse rejected the timestamp format used for earliest=/latest= queries; a "now" literal token injected into query text; a GridStack/uPlot layout-timing race; JS's Date.parse being too lenient to use as a timestamp-detection heuristic; a rule's "enabled" field silently defaulting to false when omitted; and the evaluator's claim-batch-size and worker-pool-concurrency defaulting to the same value, causing 500 concurrently-due rules to take 125s to cycle through instead of the configured 60s.
This commit is contained in:
@@ -95,6 +95,42 @@ ClickHouse/Tantivy routing strategy, and
|
||||
`/docs/query-language-reference.md` for the user-facing syntax reference
|
||||
once built.
|
||||
|
||||
## What "done" looks like for Phase 3
|
||||
|
||||
A user can build a multi-panel dashboard from saved Phase 2 queries (at
|
||||
least a line chart panel and a table panel, working end-to-end against
|
||||
live data), save an alert rule that fires a Slack webhook when a
|
||||
condition is met (threshold comparison, or "absence" — the query returned
|
||||
zero rows in its own time window), and see the delivery attempt logged —
|
||||
all from the web UI, without touching the API directly. See
|
||||
`/docs/phase-3-dashboard-design.md` and `/docs/phase-3-alerting-design.md`
|
||||
for the data models and the alerting evaluator's firing/resolved state
|
||||
machine, and `/docs/phase-3-runbook.md` for the live-stack verification,
|
||||
including a load test of the alert evaluator against ~500 concurrent
|
||||
rules.
|
||||
|
||||
This phase adds PostgreSQL as a new pinned-stack component (see the
|
||||
dashboard design doc for why ClickHouse can't do this job — dashboards
|
||||
and alert state need real row-level locking and transactional
|
||||
read-modify-write, which ClickHouse's MergeTree family doesn't provide),
|
||||
scoped strictly to control-plane config: dashboards, panels, notification
|
||||
targets, alert rules, alert state, delivery log. Log data itself stays on
|
||||
ClickHouse/Tantivy only, unchanged.
|
||||
|
||||
Non-goals for this phase (same discipline as every phase so far):
|
||||
- No multi-tenancy enforcement and no `enterprise/` module work — single
|
||||
tenant/org assumed. New tables carry a `tenant_id` column so Phase 4's
|
||||
retrofit doesn't require a schema migration + backfill, but nothing
|
||||
reads or enforces it yet.
|
||||
- No raw-SQL dashboard panels (time-range injection isn't reliable
|
||||
against arbitrary SQL) — pipe-syntax queries only.
|
||||
- No per-group/multi-row threshold alerting (e.g. "alert separately per
|
||||
host") — a threshold rule's query must resolve to a single row.
|
||||
- No debounce on the way down — a firing alert resolves on the first
|
||||
false evaluation, no symmetric "stay firing for N more minutes" hold.
|
||||
- No Kubernetes Operator/Helm deployment work — still docker-compose,
|
||||
`/deploy` remains stubbed.
|
||||
|
||||
## When in doubt
|
||||
Ask before: changing the pinned stack, adding a new external dependency
|
||||
that pulls in a large transitive tree, or making an architectural decision
|
||||
|
||||
Reference in New Issue
Block a user