Files
cairnobs/deploy/helm/sentry/README.md
T
jcoffey-dev 3037b31b0f Phase 4: Helm chart enforces api vs enterprise-api, closing the deployment-topology gap
deploy/helm/sentry/templates/api.yaml and the new enterprise-api.yaml
are mutually exclusive, gated on opposite sides of the same
enterprise.enabled flag -- exactly one renders, both as a Deployment+
Service named {{ .Release.Name }}-api on port 8080, so every consumer
(alerting's API_QUERY_URL, web's build args) needs zero conditional
logic of its own. This is the concrete fix for what the threat model
named as the single largest remaining gap once both storage engines'
isolation mechanisms were built: previously nothing forced or flagged
whether a deployment ran the tenant-isolated binary. Now the same flag
that turns on RBAC/audit/SSO also chooses the query binary.

Verified by parsing (not eyeballing) helm template's rendered output
under both value sets: exactly one sentry-api Deployment/Service either
way, with the right image, and kubeconform -strict clean against the
real Kubernetes 1.31 schema. Not applied to a live cluster (still no
cluster in this environment) -- docker-compose.yml also still runs
plain api unconditionally, so this enforcement is Helm-only for now.

Updated the threat model, architecture doc, CLAUDE.md, and deploy/
READMEs to reflect this and to name what's left: ingest has no tenant
concept for either storage engine (undesigned), and the Tenant CRD
(deploy/operator) and enterprise-api -provision-tenant are still two
separate, unreconciled provisioning mechanisms.
2026-08-14 06:20:20 -07:00

112 lines
5.4 KiB
Markdown

# deploy/helm/sentry
A Helm chart covering every `docker-compose.yml` service (Redpanda,
ClickHouse, Postgres, ingest, search, alerting, web) plus, when
`enterprise.enabled: true`: enterprise-auth, the `deploy/operator`
tenant-operator, and `Tenant` CRs from `values.tenants`. See
`/deploy/README.md` for what "multi-tenant-aware" does and doesn't mean
at this layer, and its verification-status section before trusting this
against a real cluster.
This chart never builds images -- push every image its `values.yaml`
references to a registry the cluster can pull from first, same division
of labor as `docker compose build` vs. `docker compose up`.
## `api` vs `enterprise-api`: one Deployment, chosen by `enterprise.enabled`
`templates/api.yaml` and `templates/enterprise-api.yaml` are mutually
exclusive, gated on opposite sides of the same `enterprise.enabled` flag
-- exactly one of them ever renders, both under the same
`{{ .Release.Name }}-api` Service name and port 8080. This is the fix
for what `/docs/security/threat-model.md` named as Phase 4's single
largest remaining gap once both storage engines' isolation mechanisms
were built: previously nothing forced or even flagged whether a
deployment ran the tenant-isolated binary. Now it's not a second knob to
remember -- the same flag that turns on RBAC/audit/SSO also swaps which
query binary actually serves `/query` and `/dashboards` traffic. Every
consumer (`alerting`'s `API_QUERY_URL`, `web`'s build args) needs zero
conditional logic of its own, since both variants answer on the same
name/port.
`enterprise-api` starts with an empty tenant set until
`-provision-tenant` has been run for at least one tenant (see
`/enterprise/README.md`) -- until then it's up and healthy, but every
`/query` request correctly fails closed with no tenant to route to.
## Startup ordering
`docker-compose.yml` uses `depends_on: condition: service_healthy` /
`service_completed_successfully` to sequence startup (e.g. `api` waits
for `clickhouse-migrate` to actually finish, not just for `clickhouse` to
be reachable). This chart approximates that more loosely:
- Migration Jobs (`clickhouse-migrate`, `metadata-migrate`,
`redpanda-provision`) are plain `Job` resources (not Helm hooks --
making the StatefulSets they depend on into hooks too, to get
ordering, would break `helm upgrade`/`helm uninstall`'s normal
ownership tracking of stateful resources, a worse tradeoff), with
`backoffLimit: 6` so they retry a few times if their dependency isn't
up yet.
- App Deployments get an `initContainer` that busy-waits for their
dependency's **TCP port**, not for a specific Job's completion (see
`templates/_helpers.tpl`'s `sentry.waitForTCP`) -- this covers "is
ClickHouse/Postgres/Redpanda up" but not "has the migration Job
actually finished."
- The gap that leaves (a pod starts before its migration has completed)
is covered by every Go service here already calling `os.Exit(1)` on a
failed startup DB ping (see e.g. `api/cmd/api/main.go`) --
Kubernetes' pod restart policy retries with backoff until the schema
is ready. This is a real, working, but *looser* guarantee than
docker-compose's explicit ordering -- documented here rather than
implied to be equivalent.
## Trying the two-tenant example
```sh
# Quote each --set value -- zsh globs an unquoted tenants[0] as a
# pattern and fails with "no matches found."
helm install sentry . --include-crds \
--set enterprise.enabled=true \
--set tenantOperator.enabled=true \
--set 'tenants[0].name=acme' --set 'tenants[0].displayName=Acme Corp' \
--set 'tenants[1].name=globex' --set 'tenants[1].displayName=Globex Corporation'
kubectl get tenants
kubectl get secret sentry-tenant-acme-clickhouse sentry-tenant-globex-clickhouse
```
This proves the K8s-side half of Phase 4's "two tenants... with their
own users, roles, dashboards" exit criteria (`/CLAUDE.md`) -- a real
per-tenant credential Secret exists for each, generated by
`deploy/operator`'s `Tenant` controller. It does **not** by itself give
either tenant a working ClickHouse database or Tantivy index -- that
needs `enterprise-api -provision-tenant=<id>` (a separate, deliberately
manual operator action; the `Tenant` CRD and `-provision-tenant` are two
independent mechanisms today, not yet unified -- see
`/enterprise/README.md`), and OIDC login (built, but still needs a
manual `tenant_memberships` row -- see `/docs/phase-4-runbook.md` §3a)
before a human can actually query as that tenant.
## `web`'s image needs rebuilding per environment
`web` is a static SvelteKit build (`adapter-static`) -- its three API
base URLs (`VITE_API_BASE_URL`/`VITE_ALERTING_API_BASE_URL`/
`VITE_ENTERPRISE_AUTH_BASE_URL`) are baked in at **image build time**
(`web/Dockerfile`'s build args), not read from the container's
environment at runtime. `values.yaml`'s `web.builtWithApiBaseURL` etc.
document what the image you point `web.image` at needs to have been
built with (an Ingress hostname, a LoadBalancer IP, etc.) -- this chart
has no Ingress resources and can't itself act on those values; rebuild
`web`'s image with the right build args for wherever this release is
actually reachable from a browser before pointing real users at it.
## Validating without a cluster
```sh
helm lint .
helm template sentry . --include-crds > /tmp/rendered.yaml
```
See `/deploy/README.md`'s verification section for what was actually
checked this way (and what wasn't -- no live cluster was available).