templates/enterprise-auth.yaml never set POSTGRES_ADDR/DATABASE/
USERNAME/PASSWORD at all -- enterprise-auth silently fell back to its
localhost:5432 default and could never actually reach Postgres,
crash-looping forever. Fixed to match api.yaml's existing pattern
(Service DNS name + Secret-sourced password), plus a wait-for-postgres
initContainer for the same startup-ordering reason api.yaml has one.
templates/clickhouse.yaml was missing CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT
-- same real bug docker-compose.yml had, now fixed there too: the
official image's default user lacks CREATE USER privilege without it,
so tenantprovision's -provision-tenant could never actually provision a
tenant through this chart.
Neither of these had ever been caught before because this chart had
never been installed against a real cluster -- both surfaced and were
fixed running the full "Trying the two-tenant example" walkthrough
against a real kind cluster, ending with both tenants reaching
status.phase: Active and real generated ClickHouse credentials in their
Secrets, closing /docs/phase-4-runbook.md's last remaining gap.
RBAC (api/internal/authz) is live on /query and /dashboards, backed by a
new enterprise/ module (session issuance, audit logging, RBAC storage,
OIDC/SAML protocol wiring) that core never imports -- only calls over
HTTP. Found and fixed a real cross-tenant vulnerability in dashboards
(no tenant_id filtering at all) while writing the threat model doc.
Two things are explicitly NOT done, documented rather than hidden:
tenant isolation for log data itself (/query still shares one ClickHouse
connection and Tantivy index across every tenant -- RBAC controls who
can query, not what a query can see), and human SSO login (protocol
wiring exists, no HTTP handler calls it yet). See
docs/security/threat-model.md and docs/phase-4-runbook.md.
Also adds deploy/ (Go Operator + Helm chart, validated offline only --
no cluster was reachable in this environment).