This is a large squashed commit covering two batches of prior uncommitted
work plus a full security-audit remediation pass, kept together because
go.mod/go.sum and several shared files (main.go, handler.go) were touched
by both and splitting risked non-building intermediate commits.
Features (built earlier, previously uncommitted):
- Local username/password login for single-tenant deployments with no
SSO configured (api/localauth, alerting/internal/sessioncheck,
sentryctl users, web/src/routes/login, metadata migrations 0040/0041).
- Remotely-editable additional log file paths for agents, on top of
their existing primary source (api/agents, agent/sentry-agent
extra-file-path diffing, web agent config UI).
- IPv4/IPv6 addresses reported alongside other host system metrics.
Security audit remediation (this pass, all live-verified in production):
- Critical: block ClickHouse SSRF table functions (url/remote/file/s3/...)
in the raw-SQL query escape hatch.
- High: deny sensitive paths and require Admin to add agent
extra_file_paths (Editor could previously point an agent at /etc/shadow
or an SSH key); alerting webhook targets now validate against
internal/metadata/loopback addresses, both at creation and send time;
alerting's session middleware now enforces an Editor+ floor on
mutating requests instead of "any authenticated session"; bumped
goxmldsig to close a SAML signature-verification bypass (GO-2026-4753).
- Medium: per-IP login rate limiting; security response headers
(HSTS/CSP/nosniff/X-Frame-Options/Referrer-Policy/Permissions-Policy)
on web/nginx.conf; a DevCredentialWarnings check in every Go service's
config loader, logging loudly at startup if a deployment is still on
docker-compose.yml's literal dev-only credentials; dependency bumps
(golang.org/x/text, grpc, x/net, quick-xml, h2) across every affected
Go module and both Rust crates, including a previously-uncovered x/net
vulnerability in deploy/operator; a new security-scan.yml CI workflow
running cargo-deny/govulncheck/npm-audit, mirroring the existing
license-compliance.yml matrix shape.
- Low: removed sentryctl's plaintext --password flag (shell
history/`ps` exposure) in favor of stdin and a --password-stdin flag
for reset-password's optional specific-password path; a dummy bcrypt
comparison closes a login response-time username-enumeration
side-channel.
Saved, shareable multi-panel dashboards (table/line/bar/single-stat
panels via gridstack + uPlot, global + per-panel time range, JSON
export/import) and threshold/absence alert rules with an
ok/pending/firing evaluator and webhook/Slack/PagerDuty delivery.
- New /metadata component: Postgres control-plane store for dashboards,
panels, notification targets, alert rules/state, and delivery log --
see docs/phase-3-dashboard-design.md for why ClickHouse's MergeTree
family isn't a fit for this access pattern (needs real row-level
locking and read-your-writes consistency).
- api/internal/dashboards: dashboard/panel CRUD, pure -- panel query
execution stays client-side, reusing the existing /query endpoint.
- New /alerting service: rule/target CRUD, a ticker-driven evaluator
(claim-then-evaluate concurrency control, transactional-outbox
delivery, query errors and threshold zero-rows never coerced into a
false transition) and webhook/Slack/PagerDuty delivery with
retry/backoff. See docs/phase-3-alerting-design.md for the full
state-machine design and the four correctness properties it
implements.
- web: /dashboards and /alerts UIs; cli: sentryctl dashboards/alerts
list/get/apply, seeding a future Terraform provider's JSON contract.
- hack/alert-load-test: 500 rules against real ClickHouse data, real
measured results in docs/phase-3-runbook.md.
Five real bugs found by actually running this against a live stack
(documented in the runbook, not just fixed silently): a latent Phase 2
bug where ClickHouse rejected the timestamp format used for
earliest=/latest= queries; a "now" literal token injected into query
text; a GridStack/uPlot layout-timing race; JS's Date.parse being too
lenient to use as a timestamp-detection heuristic; a rule's "enabled"
field silently defaulting to false when omitted; and the evaluator's
claim-batch-size and worker-pool-concurrency defaulting to the same
value, causing 500 concurrently-due rules to take 125s to cycle through
instead of the configured 60s.
End-to-end log pipeline for Linux hosts, per /docs/architecture.md:
- proto: shared gRPC contract (agent <-> ingest), Go bindings checked in
- agent: Rust, musl-targeted, journald/file sourcing, RFC5424 parser,
mTLS gRPC client, no required config for the common case
- ingest: Go, single binary with --mode server|consumer|all; gRPC front
end forwards to Redpanda unchanged, consumer normalizes and
batch-writes to ClickHouse with at-least-once delivery
- storage: ClickHouse schema + a plain SQL-file migration runner
- api: minimal SELECT-only query endpoint, plain REST (not gRPC+gateway
yet -- see api/README.md)
- web: SvelteKit static SPA, one query page
- transport: Redpanda compose + topic provisioning
- cli: sentryctl ping stub
- hack/dev-certs: throwaway CA + cert generation for local mTLS
- root docker-compose.yml + docs/phase-0-runbook.md tie it together
Not yet run end-to-end against real Docker/ClickHouse/Redpanda -- see the
runbook's caveats section before relying on this working as-is.