Files
cairnobs/web
jcoffey-dev 9435115ab7 Phase 3: dashboards and alerting
Saved, shareable multi-panel dashboards (table/line/bar/single-stat
panels via gridstack + uPlot, global + per-panel time range, JSON
export/import) and threshold/absence alert rules with an
ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery.

- New /metadata component: Postgres control-plane store for dashboards,
  panels, notification targets, alert rules/state, and delivery log --
  see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree
  family isn't a fit for this access pattern (needs real row-level
  locking and read-your-writes consistency).
- api/internal/dashboards: dashboard/panel CRUD, pure -- panel query
  execution stays client-side, reusing the existing /query endpoint.
- New /alerting service: rule/target CRUD, a ticker-driven evaluator
  (claim-then-evaluate concurrency control, transactional-outbox
  delivery, query errors and threshold zero-rows never coerced into a
  false transition) and webhook/Slack/PagerDuty delivery with
  retry/backoff. See docs/phase-3-alerting-design.md for the full
  state-machine design and the four correctness properties it
  implements.
- web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts
  list/get/apply, seeding a future Terraform provider's JSON contract.
- hack/alert-load-test: 500 rules against real ClickHouse data, real
  measured results in docs/phase-3-runbook.md.

Five real bugs found by actually running this against a live stack
(documented in the runbook, not just fixed silently): a latent Phase 2
bug where ClickHouse rejected the timestamp format used for
earliest=/latest= queries; a "now" literal token injected into query
text; a GridStack/uPlot layout-timing race; JS's Date.parse being too
lenient to use as a timestamp-detection heuristic; a rule's "enabled"
field silently defaulting to false when omitted; and the evaluator's
claim-batch-size and worker-pool-concurrency defaulting to the same
value, causing 500 concurrently-due rules to take 125s to cycle through
instead of the configured 60s.
2026-08-13 17:29:38 -07:00
..
2026-08-13 17:29:38 -07:00
2026-08-13 17:29:38 -07:00
2026-08-13 17:29:38 -07:00

web

SvelteKit frontend. Phase 0: one page, one query box, one table. No auth, no styling polish, no routing beyond /.

What it does

Textarea for a raw SQL string → POST {VITE_API_BASE_URL}/query on /api → renders {columns, rows} as an HTML table, or shows {error} from a rejected/failed query. That's the whole app — see src/routes/+page.svelte.

Why a static build, not a Node server

Scaffolded with @sveltejs/adapter-static: this page has no server-side data loading (all data comes from a client-side fetch triggered by the submit button), so there's nothing here that needs a running SvelteKit server. A prerendered static site is simpler to build, deploy, and reason about than running Node in production for a page that's this thin.

Because it's static, VITE_API_BASE_URL is baked in at build time, not read at container start. Set it before npm run build (or pass --build-arg VITE_API_BASE_URL=... to docker build) — changing it later means rebuilding, not just restarting the container.

Building & running

npm install
cp .env.example .env   # adjust VITE_API_BASE_URL if /api isn't on localhost:8080
npm run dev             # local dev server with hot reload
npm run check            # svelte-check, type errors
npm run build             # static output to build/
npm run preview            # serve the static build locally to sanity-check it
docker build -f Dockerfile -t sentry-web .   # context is web/, not the repo root
docker run -p 3000:3000 sentry-web

Why nginx, not distroless

The repo convention prefers distroless/scratch base images. Serving a static SPA still needs some HTTP server, though, and nginx:alpine is the boring, standard choice for that job — writing a custom static-file binary just to stay distroless would be more engineering than a Phase 0 placeholder page justifies. nginx.conf here is minimal: serve build/, fall back to index.html for client-side routing (only one route exists today, but this is what you want the moment a second one is added).