Full rebrand across cosmetic branding, code identifiers, and infrastructure/data-plane naming, using the supplied Cairn OBS logo package. Cosmetic: favicon/logo swap (also closes a stale license-audit finding -- the old favicon was SvelteKit's unreplaced scaffold logo), new centered welcome landing page, larger/legible sidebar logo, page titles, CLAUDE.md/README/docs prose. Code identifiers: Go module path github.com/sentry/sentry -> github.com/cairnobs/cairnobs across all 13 modules and ~91 files (protoc regenerated); Rust crates sentry-agent/sentry-parser/sentry-search -> cairnobs-*; CLI sentryctl -> cairnobsctl; Terraform provider fully renamed (sentry_dashboard etc. -> cairnobs_dashboard, provider type, env vars); every session/auth cookie name; agent config paths and Windows service identity. Deliberately preserved: the gRPC wire protocol's protobuf packages (sentry.logs.v1, sentry.agent.v1) and their Go import directory (proto/sentry/...) -- renaming the wire-level package would break every currently-deployed agent binary (confirmed two real hosts, including mail.inbuxa.com, are actively streaming through this exact contract) until rebuilt and redeployed in lockstep with an ingest cutover. Only the Go module path wrapping the generated code changes. Infrastructure: every docker-compose container name (root and three component-level compose files); the Helm chart (directory, Chart.yaml, named-template helpers, all templates, values.yaml image repos); Kubernetes Operator (CRD group sentry.io -> cairnobs.io, both CRD YAML files, Go identifiers, RBAC markers); the coupled enterprise/tenantcrd package. Caught and fixed real path-coupling bugs along the way: the Helm chart's search/ingest volume mounts and the dev-only-credential detection constant vs. docker-compose.yml's literal values had to move together or a security warning would have silently stopped firing. Data plane: Postgres database sentry_metadata -> cairnobs_metadata and role sentry -> cairnobs; ClickHouse database sentry -> cairnobs; Kafka topic sentry.logs.raw -> cairnobs.logs.raw and its consumer groups. Source-level defaults, docker-compose.yml, and every migrate.sh/ provision script default updated together; already-applied migration files left untouched per this repo's immutable-migration convention. Verified at every layer: all 13 Go modules build/vet/test clean, both Rust workspaces (agent, search) build/clippy/test clean, npm run check/ build clean, docker compose config validates on all four compose files. Live-verified against a real docker stack multiple times through this work, including a final fresh-volume run confirming the actual renamed Postgres database/role, ClickHouse database, and Kafka topic all work end to end with a real login and query, zero console errors.
7.3 KiB
deploy
Kubernetes deployment for Cairn OBS, added in Phase 4 (/deploy was
deliberately stubbed through Phase 3 -- see /CLAUDE.md's Phase 3
non-goals). Two pieces:
operator/-- a small Go controller-runtime Operator managing one CRD (Tenant). Seeoperator/README.md.helm/cairnobs/-- a Helm chart covering everydocker-compose.ymlservice, plus the operator andTenantCRs whenenterprise.enabled=true. Seehelm/cairnobs/README.md.
What "multi-tenant-aware" means here, precisely
Per /docs/phase-4-isolation-design.md, tenant isolation itself lives at
the application layer inside enterprise/ (enterprise-api holds a
map of per-tenant ClickHouse connection pools via internal/chrunner;
search holds a map of per-tenant Tantivy indices via
src/registry.rs) -- not at the Kubernetes layer. This directory is
not "one Deployment per tenant" or a general multi-cluster system;
that's an explicit Phase 4 non-goal (see /CLAUDE.md). What it does
add:
- A
TenantCRD + controller (operator/internal/controller) that reflects real provisioning state ontostatus.phase/aReadycondition, derived from whetherenterprise-api -provision-tenanthas reported real ClickHouse provisioning. - A Helm chart that can install zero-or-more
TenantCRs (values.tenants) alongside the rest of the stack, and swapsapi's Deployment forenterprise-api's wheneverenterprise.enabledis true, so which query binary actually serves traffic is no longer a separately-forgettable decision (seehelm/cairnobs/README.md's "apivsenterprise-api" section).
Now unified, in a deliberately lightweight way: enterprise-api -provision-tenant=<id> stays the sole real actor — it's the only thing
that calls ClickHouse (CREATE DATABASE/CREATE USER/GRANT, via
enterprise/internal/tenantprovision) and writes rbacstore. What
changed: once it succeeds, it also syncs the result into the Tenant
CRD (enterprise/internal/tenantcrd) — creating the Secret with real
credentials (the controller no longer generates a placeholder one that
authenticated against nothing) and setting the status fields the
controller reads to compute Phase/Ready. The controller itself
gained no new credentials and still never touches ClickHouse/Postgres --
it's a pure function of spec.suspended and whatever
-provision-tenant has reported, never an independent second guess at
"is this tenant really provisioned." A Tenant reaching
status.phase: Active now means the same thing rbacstore.tenants. status='active' does, not two different claims — see
enterprise/internal/tenantcrd's and operator/internal/controller/ tenant_controller.go's doc comments for the full split, and
enterprise-api -provision-tenant's TENANT_CRD_NAMESPACE env var
(set automatically by the Helm chart when tenantOperator.enabled) to
turn this on. Deliberately not built: the operator's reconcile loop
itself calling ClickHouse/rbacstore directly (a "full unification"
option considered and set aside — it would give the operator two new
credential sets and require real reconcile-loop idempotency design for
an inherently one-shot external side effect, a bigger and riskier
change than this repo's provisioning story needed to close the actual
gap, which was two disconnected sources of truth, not two actors).
Verification status -- read before trusting this against a real cluster
Now verified against a real live Kubernetes cluster. kind/
kubectl/helm were installed without root (static binaries into
~/.local/bin), a real local cluster was created, every image this
chart references was built and loaded into it, and the full two-tenant
walkthrough (helm/cairnobs/README.md) was run end to end -- both tenants
reached Tenant.status.phase: Active with real generated ClickHouse
credentials in their Secrets. See /docs/phase-4-runbook.md §7 for the
exact commands and the two real chart bugs this run found and fixed
(enterprise-auth missing its Postgres connection env vars entirely,
ClickHouse missing the env var that grants CREATE USER privilege) --
neither helm lint, helm template, nor the kubeconform schema check
below could have caught either, since both only manifest once real pods
actually try to start and connect to each other.
What was verified offline, before real cluster access existed (still true, kept as additional evidence, not superseded by the above):
deploy/operator:go build/go vet/go test ./...all pass, including reconciler tests against controller-runtime's fake client (internal/controller/tenant_controller_test.go) -- real reconcile logic exercised, but not against a real apiserver (noenvtestbinaries available; see that test file's doc comment).enterprise/internal/tenantcrd(the "lightweight unification" half-provision-tenantruns):go testpasses againstk8s.io/client-go's fake dynamic and typed clientsets -- real client library, fake transport, same shape asenterprise/internal/ searchclient's in-process gRPC tests. What this doesn't prove: thatcairnobs.io/v1alpha1.Tenant's real CRD schema (a real apiserver's OpenAPI validation) accepts exactly what this package writes -- thehelm template/kubeconform check below covers the schema shape, not a live write against it.deploy/operator/config/crd/cairnobs.io_tenants.yaml: parsed withsigs.k8s.io/yaml+ strict-unmarshaled into the realk8s.io/apiextensions-apiserverCustomResourceDefinitionGo type -- catches YAML syntax errors and structural mistakes, not a live-cluster admission check.deploy/helm/cairnobs:helm lintpasses;helm templaterenders cleanly under both default values and aenterprise.enabled: true+ two-tenant override; the rendered output was checked withkubeconform -strictagainst the real Kubernetes 1.31 OpenAPI schema for every built-in resource kind (22-31 resources depending on values, 0 invalid) -- this catches schema mistakes (wrong field names, wrong types) but not whether the resources actually reconcile correctly together on a live cluster (Job/StatefulSet startup ordering, PVC provisioning, actual pod scheduling). Specifically confirmed by parsing the rendered YAML (not just eyeballing it): exactly oneDeployment/Servicenamedcairnobs-apirenders in each mode, with theenterprise.enabled: truerender using theenterprise-apiimage and the default render using plainapi's. Also confirmed for thetenantOperator.enabled: truecase:enterprise-apigets its own ServiceAccount/Role/RoleBinding (exactlytenants/tenants/status/secrets, no more),tenant-operator's own ClusterRole no longer grantssecretsat all, andTENANT_CRD_NAMESPACEis set onenterprise-api's container only whentenantOperator.enabledis true.- Docker image builds (
operator/Dockerfileand every otherDockerfilethis chart references) are now confirmed working too -- all twelve images this chart needs were built and loaded into the testkindcluster above.
Before relying on this in production: ingest needs a real cert-manager
(or equivalent) issued Secret, not hack/dev-certs's throwaway dev
certs; ClickHouse/Postgres data isn't backed by anything durable beyond
the cluster's own PVC provisioner in this chart; and only Auth0 has been
tried as a real external IdP so far (see
/docs/phase-4-runbook.md §3a/§3b).