Auth0's SAML2 Web App addon (the same dev tenant §3a used) stood in as a real SAML IdP, over a genuine self-signed TLS proxy in front of enterprise-auth (required, not optional, for SAML's SameSite=None cookie). Full round trip confirmed: real signed assertion, audience/ destination/signature validation, correct multi-membership handling, and POST /internal/authorize returning the selected tenant/role. Updates the runbook's verification status, §3b, and the threat model's "Read this first" finding and summary table to reflect this and the isSecureRequest fix it found. §7/§11's live-cluster steps (no kind/kubectl in this environment) are now the only remaining gap in the entire runbook.
57 KiB
Phase 4 runbook
Extends /docs/phase-0-runbook.md through /docs/phase-3-runbook.md
with SSO plumbing, RBAC enforcement, tenant-scoped dashboards, audit
logging, and a Kubernetes deployment path. Read those first.
Verification status — read this before the rest of this doc
Docker daemon access became available partway through this phase's
work, and the rest of this runbook (§§1–10, 10a, 13, 14, most of §8) has
now genuinely been run against a real docker-compose stack — real
ClickHouse (clickhouse/clickhouse-server:24.8), real Postgres, real
enterprise-auth/enterprise-api/enterprise-ingest containers, two
real provisioned tenants (acme, globex). This closed every
previously-disclosed "written but never run" gap for the ClickHouse
side, and along the way found and fixed six real bugs no amount of
Docker-free testing could have caught:
enterprise/Dockerfileused acontext: enterprisebuild context too narrow forenterprise/go.mod'sreplace ../apidirective onceenterprise-authstarted importingapi/httpserverand (transitively)api/dashboards-- fixed to build from the repo root, matchingenterprise-api/enterprise-ingest's Dockerfiles.api/dashboards.Store.AddPanel/UpdatePaneldidn't callvalidatePanelthe wayCreateDashboard's inline panel-creation path already did, so a panel added via those two methods could hit aviz_configNOT NULL constraint violation instead of getting the same default-empty-JSON treatment every other panel-creation path gets -- fixed by callingvalidatePanelin both.enterprise/internal/rbacstore.SetDataSourceClickHouseCredentialslet a malformed (non-UUID) id leak Postgres's raw22P02error past the store'sErrNotFoundboundary instead of treating "can't possibly match a row" the same as "no such row" -- fixed by translating that specific Postgres error code toErrNotFound.enterprise/cmd/enterprise-api/main.goregisteredGET /healthztwice -- once explicitly, once already covered byqueryapi.Handler.RegisterRoutes-- which panicsnet/http'sServeMuxon startup. This meantenterprise-apicould never actually start; every previous "built" claim for this binary had only ever been a successfulgo build, never a successful process start. Fixed by deleting the redundant registration.- ClickHouse's
defaultadmin user genuinely lackedCREATE USERprivilege indocker-compose.yml-- the official image needsCLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT=1(not the more obvious-lookingCLICKHOUSE_ACCESS_MANAGEMENT, confirmed by reading the image's own/entrypoint.shafter the wrong name silently did nothing), which this compose file never set. Every tenant-provisioning code path was correct Go that had simply never been able to authenticate its own admin connection strongly enough to run. enterprise/internal/tenantprovision.ProvisionClickHouserelied on ClickHouse RBAC's assumed "default-deny for a freshly created user" forsystem.*access -- verified live to be false on ClickHouse 24.8: a fresh tenant user could readsystem.tables(though not, it turns out,system.query_log, which is genuinely access-checked). Fixed with an explicitREVOKE SELECT ON system.* FROM <user>, andTestProvisionedUserCannotReadSystemTableswas corrected to check what each system table actually does under a proper revoke (hard deny forquery_log, verified-empty fortables) instead of assuming both hard-deny identically. See/docs/security/threat-model.md's "Read this first" section for the full account.
What's still not verified, and why: §3a (OIDC), §3b (SAML), and §12 (tenant picker, both single- and multi-membership paths) are now all closed -- verified against a real Auth0 developer tenant acting as both an OIDC and a SAML IdP, full browser round trips, real session cookies authorizing correctly, including selecting between two real tenant memberships and getting back the right tenant/role each time (see §3a, §3b, and §12 below). That pass found and fixed two real bugs:
web/Dockerfileonly declaredARG/ENVforVITE_API_BASE_URL, sodocker-compose.yml's build args forVITE_ALERTING_API_BASE_URL/VITE_ENTERPRISE_AUTH_BASE_URLwere silently dropped by Docker (an undeclared--build-argis dropped, not an error) --enterpriseAuthBasecame outundefinedin the built bundle, so the tenant-picker page threw "enterprise-auth is not configured" against a real running container even thoughdocker-compose.ymllooked correct.enterprise/internal/loginhandler.godecided every cookie'sSecureattribute fromr.TLS != nilalone -- wrong for the deployment shape this handler actually runs in, sinceenterprise-authnever terminates TLS itself. Closing SAML's real-IdP gap required a genuine TLS-terminating proxy in front of it (SAML's request cookie isSameSite=None, which the cookie spec requires to be paired withSecure, and Chrome silently drops it otherwise), which is exactly what surfaced this:r.TLSwas nil at the Go process even over a real HTTPS client connection, so the cookie came back non-Secureand Chrome dropped it, breaking the flow. Fixed with anisSecureRequest(r)helper that also honorsX-Forwarded-Proto: https-- affects every cookie this handler sets, not just SAML's, so this would have hit any real reverse-proxied deployment, not just this test.
§7 and §11's live-cluster steps still need kind/kubectl, which
aren't installed in this environment -- their offline-only checks
(go build/go vet/go test, helm lint, helm template + parsing
the rendered YAML) all pass and are documented as such below. That's the
only remaining gap in this entire runbook that isn't already closed.
If you're reading this to decide whether Phase 4 is production-ready:
closer than before, but not yet -- see
/docs/security/threat-model.md's "Read this first" section for the
current, precise state of every control, including the ones still
gated on a real IdP or a real cluster.
1. Bring up the stack
docker compose build enterprise-auth api alerting web
docker compose up -d
docker compose ps
New service beyond Phase 3: enterprise-auth (port 8082) — see
enterprise/README.md. Not wired into api/alerting's enforcement by
default (docker-compose.yml's comment on why: no OIDC/SAML login flow
exists yet, so turning on enforcement by default would break the web UI
and sentryctl with no way to log in).
2. Confirm Phase 0-3 behavior is unchanged
Every existing single-tenant flow must still work exactly as before — this is the regression check for the nil-authorizer no-op design running through every piece of Phase 4 auth wiring:
curl -s -X POST http://localhost:8080/query -H 'Content-Type: application/json' -d '{"query":"stats count"}'
sentryctl dashboards list
curl -s http://localhost:8081/healthz
All three should behave exactly as in the Phase 3 runbook — no auth
required, since ENTERPRISE_AUTH_URL/API_SERVICE_TOKEN aren't set.
3. enterprise-auth: mint and validate a service token
curl -s http://localhost:8082/healthz && echo " <- OK"
TOKEN=$(docker compose run --rm enterprise-auth -mint-service-token=alerting)
echo "$TOKEN"
curl -s -X POST http://localhost:8082/internal/authorize -H "Authorization: Bearer $TOKEN"
# expect: {"tenant_id":"","user_id":"","role":"service"}
curl -s -o /dev/null -w "invalid token -> %{http_code}\n" \
-X POST http://localhost:8082/internal/authorize -H "Authorization: Bearer garbage"
# expect: 401
curl -s http://localhost:8082/auth/features
# expect: {"sso_configured":false,"oidc_enabled":false,"saml_enabled":false}
# (no OIDC_ISSUER_URL/SAML_IDP_METADATA_URL set in this compose file)
3a. enterprise-auth: human login via OIDC (now genuinely verified
against a real external IdP, not just a real fake one)
enterprise/internal/loginhandler's tests already prove the mechanism
works end to end against a real fake IdP (go test ./internal/loginhandler/... -v from enterprise/, no Docker needed --
see enterprise/README.md). Closed for real in this pass: a free
Auth0 developer tenant was created, wired into a local-only
docker-compose.override.yml (never committed -- see .gitignore), and
the full human login flow was driven through a real browser end to end:
GET /auth/oidc/login redirected to Auth0's real hosted login page; the
first attempt (before any tenant_memberships row existed) correctly
failed closed with "this identity has no tenant membership" while still
creating the users row via UpsertUserBySSO, exactly as designed;
after -grant-membership-* granted that real Auth0 identity
([email protected]) an admin role on acme, a second login
completed the full OAuth code exchange, consent screen, and redirect to
web, landing a real sentry_session cookie; POST /internal/authorize with that cookie returned
{"tenant_id":"acme","user_id":"...","role":"admin"} -- the exact
membership granted, round-tripped through a real external IdP, not a
stand-in. To reproduce, point docker-compose.yml's enterprise-auth
service at a real OIDC IdP (a free Auth0/Okta developer tenant, or any
IdP you control):
# Add to enterprise-auth's environment in docker-compose.yml (or a
# docker-compose.override.yml):
# OIDC_ISSUER_URL: "https://your-tenant.example.com/"
# OIDC_CLIENT_ID: "..."
# OIDC_CLIENT_SECRET: "..."
# OIDC_REDIRECT_URL: "http://localhost:8082/auth/oidc/callback"
# Register that same redirect URL with the IdP's application config.
docker compose up -d --build enterprise-auth
curl -s http://localhost:8082/auth/features
# expect: {"sso_configured":true,"oidc_enabled":true,"saml_enabled":false}
Before a login can succeed, the logging-in identity needs a
tenant_memberships row. There's still no admin UI for this, but as of
this runbook revision there's no manual SQL either --
enterprise-auth's -create-tenant/-grant-membership-* operator
flags replace the old psql dance:
docker compose run --rm enterprise-auth -create-tenant=acme -display-name="Acme Corp"
# Log in once at http://localhost:8082/auth/oidc/login -- it'll fail
# with "no tenant membership" (403), but UpsertUserBySSO already created
# the users row by that point, which -grant-membership-user-email needs.
docker compose run --rm enterprise-auth \
-grant-membership-tenant=acme -grant-membership-user-email=<the email you logged in with> -grant-membership-role=viewer
Then visit http://localhost:8082/auth/oidc/login again in a real
browser, complete the IdP's login, and confirm you land on
POST_LOGIN_REDIRECT_URL (http://localhost:3000 by default) with a
sentry_session cookie set. -create-tenant only touches rbacstore
(control-plane/RBAC) -- it's independent of enterprise-api -provision-tenant's ClickHouse/Tantivy data-plane provisioning (§8),
so a tenant created this way can log users in immediately but can't yet
serve their queries until that's run too, the same "two separate
operator actions" gap named in "Known gaps" below.
-revoke-membership-*/-list-memberships-tenant/-transfer-owner-*
cover revoking, listing, and reassigning ownership the same way
-grant-membership-* covers granting (all offline operator flags, same
shape as -create-tenant) -- dashboard_permissions grants (§5a) are
the one thing here still only reachable via the HTTP endpoints
directly, no operator flag, since sentryctl dashboards permissions list|grant|revoke already exists as that surface instead (§5a).
3b. enterprise-auth: human login via SAML (now genuinely verified
against a real external IdP, over real HTTPS)
enterprise/internal/loginhandler's SAML tests already prove the
mechanism works end to end against a real fake SAML IdP (go test ./internal/loginhandler/... -run SAML -v from enterprise/, no Docker
needed). Closed for real in this pass, using the same Auth0
developer tenant §3a used, via that app's SAML2 Web App addon
(Auth0 acting as a real SAML IdP, not just OIDC). This is where SAML's
SameSite=None request cookie (which the cookie spec requires to be
paired with Secure) made the plain-HTTP setup this environment
otherwise defaults to a hard blocker -- getting a real assertion back
required standing up a genuine (self-signed, dev-only) TLS-terminating
nginx proxy in front of enterprise-auth. Doing that surfaced a real,
previously-undiscovered bug: every cookie loginhandler.go sets decided
Secure from r.TLS != nil alone, which is wrong for the deployment
shape this handler actually runs in -- enterprise-auth never
terminates TLS itself, so in any real deployment (behind an ingress/load
balancer, exactly what the proxy here stands in for) r.TLS is nil even
over a genuinely HTTPS client connection. The Secure attribute was
silently missing behind the proxy, and Chrome dropped the cookie outright
-- fixed with a new isSecureRequest(r) helper that also honors
X-Forwarded-Proto: https, with its own regression test
(TestHandleSAMLLoginSetsSecureCookieBehindATLSProxy). This bug affects
every cookie this handler sets, not just SAML's, and would have hit
production the same way, so this generalizes well beyond closing this
one section.
With that fixed, the full flow completed for real: login redirected to
Auth0's real SAML SSO endpoint, Auth0 posted back a real signed
assertion, enterprise/internal/saml validated it (audience, destination,
signature) and extracted the email attribute, and the identity landed on
the real /select-tenant page with both real tenant memberships,
exactly like §12's OIDC walkthrough -- confirmed via
POST /internal/authorize returning the selected tenant/role.
To try this yourself, point docker-compose.yml's enterprise-auth
service at a real SAML IdP (many identity providers, including Auth0,
offer a free developer/trial tenant with SAML app support):
# Add to enterprise-auth's environment in docker-compose.yml (or a
# docker-compose.override.yml):
# SAML_ENTITY_ID: "https://<your-https-host>/saml/metadata"
# SAML_ACS_URL: "https://<your-https-host>/auth/saml/acs"
# SAML_IDP_METADATA_URL: "https://your-idp.example.com/metadata"
# Register SAML_ENTITY_ID (as "audience")/SAML_ACS_URL (as the
# "Application Callback URL") with the IdP's application config -- the
# IdP needs Sentry's ACS URL to know where to POST the assertion, and
# the two must actually agree or the SP-side validation fails with a
# generic "Authentication failed" (found the hard way -- crewjam/saml's
# error here doesn't distinguish audience mismatch from other causes).
# <your-https-host> must be real HTTPS, not plain HTTP -- see above.
docker compose up -d --build enterprise-auth
curl -s https://<your-https-host>/auth/features
# expect: {"sso_configured":true,"oidc_enabled":false,"saml_enabled":true}
Bootstrapping the first tenant_memberships row uses the same
-create-tenant/-grant-membership-* flags as §3a (log in once, it
fails with 403, grant the membership using the email you logged in
with, log in again). Then visit /auth/saml/login on your HTTPS host in
a real browser, complete the IdP's login, and confirm a sentry_session
cookie lands after redirect to POST_LOGIN_REDIRECT_URL.
4. Turn on RBAC enforcement and prove it actually blocks/allows
Without touching the main stack's api container (so step 2's baseline
keeps working):
docker compose run --rm -d --name sentry-api-enforced -p 8090:8080 \
-e ENTERPRISE_AUTH_URL=http://enterprise-auth:8082 api
curl -s -o /dev/null -w "no auth -> %{http_code} (want 401)\n" \
-X POST http://localhost:8090/query -H 'Content-Type: application/json' -d '{"query":"stats count"}'
curl -s -o /dev/null -w "with service token -> %{http_code} (want 200)\n" \
-X POST http://localhost:8090/query -H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' -d '{"query":"stats count"}'
docker stop sentry-api-enforced
GET /dashboards on the same enforced instance should return 401
without a token — this section only demonstrates the service-token path
(§3/§3a/§3b cover minting a real human session via OIDC or SAML); walking
that session cookie through this same enforced instance to get a 200 is
left as the natural next verification step once real Docker/K8s access
exists, not yet done in this runbook.
5. Dashboards tenant scoping
This is the fix from Phase 4 task 7/8 (see /docs/security/threat-model.md)
— every dashboards query is now scoped to the authenticated identity's
tenant. Verify the real SQL, not just the fake-store unit tests:
docker run --rm --network sentry_default -v $(pwd)/api:/src -w /src \
-e DASHBOARDS_TEST_POSTGRES_ADDR=metadata-postgres:5432 \
-e DASHBOARDS_TEST_POSTGRES_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./dashboards/... -run Integration -v
(Path fixed from an earlier ./internal/dashboards/... -- stale since
dashboards moved from api/internal/dashboards to api/dashboards
earlier in Phase 4, once enterprise/cmd/enterprise-api needed to
import it: Go's compiler-enforced internal/ visibility rule meant a
separate module like enterprise/ could never import anything under
api/internal/..., regardless of the AGPL/commercial licensing
boundary, which only forbids the reverse direction.)
Expect all TestIntegration* tests to pass, including
TestIntegrationDashboardTenantForeignKeyRejectsUnknownTenant (the
tenant_id foreign key added in
metadata/migrations/0027_add_dashboards_tenant_fk.sql rejecting a
dashboard for a tenant that doesn't exist).
5a. Per-resource dashboard grants (new -- built and unit-tested, not yet run live)
enterprise/internal/rbacstore's dashboard_permissions CRUD and its
DashboardPermissions adapter (implementing api/dashboards. PermissionStore) have real integration tests, same skip-gated shape as
§6 below:
docker run --rm --network sentry_default -v $(pwd):/src -w /src/enterprise \
-e RBACSTORE_TEST_POSTGRES_ADDR=metadata-postgres:5432 \
-e RBACSTORE_TEST_POSTGRES_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./internal/rbacstore/... -run DashboardPermission -v
(Mount the repo root, not just enterprise/, with -w /src/enterprise as
the workdir -- same reason as enterprise/Dockerfile's doc comment:
dashboards_adapter.go imports api/dashboards, and enterprise/go.mod's
replace ../api directive needs that sibling directory actually present
in the build context, not just the package being tested.)
Expect all TestDashboardPermission*/TestSetDashboardPermission*/
TestGetDashboardPermission*/TestRevokeDashboardPermission*/
TestListDashboardPermissions tests to pass. This only takes effect
when enterprise-api (not plain api) is serving traffic -- see
§8/§10.
6. enterprise/internal/rbacstore and internal/audit (already verified — reconfirm here)
docker run --rm --network sentry_default -v $(pwd):/src -w /src/enterprise \
-e RBACSTORE_TEST_POSTGRES_ADDR=metadata-postgres:5432 \
-e RBACSTORE_TEST_POSTGRES_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./internal/rbacstore/... -v
docker run --rm --network sentry_default -v $(pwd):/src -w /src/enterprise \
-e AUDIT_TEST_POSTGRES_ADDR=metadata-postgres:5432 \
-e AUDIT_TEST_POSTGRES_PASSWORD=audit-writer-dev-only \
-e AUDIT_TEST_ADMIN_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./internal/audit/... -v
(Same repo-root-mount reasoning as §5a above -- internal/rbacstore
imports api/dashboards unconditionally via dashboards_adapter.go, so
even running the whole package's tests, not just the DashboardPermission
subset, needs api/ present. internal/audit has the same shape via
queryapi_adapter.go.)
7. deploy: Helm chart and Operator (offline-only so far — see /deploy/README.md)
No live cluster was available to kubectl apply any of this. What can
be checked without one:
cd deploy/operator && go build ./... && go vet ./... && go test ./...
cd ../helm/sentry
helm lint .
helm template sentry . --include-crds > /tmp/default.yaml
helm template sentry . --include-crds \
--set enterprise.enabled=true --set tenantOperator.enabled=true \
--set 'tenants[0].name=acme' --set 'tenants[0].displayName=Acme Corp' \
--set 'tenants[1].name=globex' --set 'tenants[1].displayName=Globex Corporation' \
> /tmp/multitenant.yaml
With a real cluster reachable (kind create cluster, or similar):
docker build -f deploy/operator/Dockerfile -t sentry-tenant-operator deploy/operator/
kind load docker-image sentry-tenant-operator # or push to a registry the cluster can pull from
helm install sentry deploy/helm/sentry --include-crds \
--set tenantOperator.enabled=true --set enterprise.enabled=true \
--set 'tenants[0].name=acme' --set 'tenants[0].displayName=Acme Corp'
kubectl get tenants
kubectl get secret sentry-tenant-acme-clickhouse -o yaml
Expect kubectl get tenants to show acme reach status.phase: Active
and the Secret to contain a generated username/password/database.
This proves the K8s-side half of a real two-tenant deployment — it does
not provision a working ClickHouse database itself (the Operator
manages the K8s Secret only); §8 below is the piece that actually
provisions ClickHouse.
8. enterprise-api: real per-tenant ClickHouse isolation
enterprise/internal/tenantprovision and enterprise/internal/chrunner
close the ClickHouse half of the headline gap §"Known gaps" below used
to describe as completely unbuilt. It's still a second binary you have
to choose to run, though — see /docs/security/threat-model.md's "Read
this first" section. With OIDC login now built (§3a), a real
curl -X POST http://localhost:8080/query walkthrough as a logged-in
tenant is possible now, but still needs the manual tenant_memberships
bootstrap from §3a — a full end-to-end curl walkthrough isn't included
here yet.
api/enterprise-api are now mutually exclusive via COMPOSE_PROFILES
(checked in as single-tenant in .env, i.e. plain api runs by
default) — see §10a below for why, and confirmation that it actually
holds. -provision-tenant itself doesn't bind a port, so it runs fine
regardless of the active profile; actually serving traffic on
enterprise-api needs the enterprise profile active, since it now
binds the same host port (8080) plain api does:
COMPOSE_PROFILES=enterprise docker compose build enterprise-api
COMPOSE_PROFILES=enterprise docker compose run --rm enterprise-api -provision-tenant=acme -display-name="Acme Corp"
COMPOSE_PROFILES=enterprise docker compose run --rm enterprise-api -provision-tenant=globex -display-name="Globex Corporation"
COMPOSE_PROFILES=enterprise docker compose up -d enterprise-api
curl -s http://localhost:8080/healthz
Confirm isolation end to end against the live stack (this is the same
assertion enterprise/internal/chrunner/chrunner_test.go's
TestRegistryTenantCannotReadOtherTenantEvenViaRawSQL makes, run here
as an integration test instead of a curl walkthrough since a full login
walkthrough isn't scripted yet):
docker run --rm --network sentry_default -v $(pwd)/enterprise:/src -w /src \
-e CHRUNNER_TEST_CLICKHOUSE_ADDR=clickhouse:9000 \
-e CHRUNNER_TEST_CLICKHOUSE_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./internal/chrunner/... -v
docker run --rm --network sentry_default -v $(pwd)/enterprise:/src -w /src \
-e TENANTPROVISION_TEST_CLICKHOUSE_ADDR=clickhouse:9000 \
-e TENANTPROVISION_TEST_CLICKHOUSE_PASSWORD=sentry-dev-only \
golang:1.25-alpine go test ./internal/tenantprovision/... -v
Expect all tests to pass, including
TestProvisionedUserCannotReadSystemTables (item 2 of
/docs/phase-4-isolation-design.md's verification plan, closed this
pass) and TestRegistryTenantCannotReadOtherTenantEvenViaRawSQL (item 1,
closed through the actual production code path, not just
tenantprovision's raw grants).
9. Tantivy per-tenant isolation — no Docker needed, actually run this one
Unlike everything above, this one doesn't need a live stack at all — Tantivy is an embedded library, not a networked service, so both halves (the Rust index registry and the Go client that talks to it) can be verified with nothing but a local toolchain:
cd search
cargo build
cargo clippy --all-targets -- -D warnings
cargo test
# expect 24 tests passing, including
# registry::tests::tenant_index_is_isolated_from_default_and_other_tenants
# -- item 3 of /docs/phase-4-isolation-design.md's verification plan --
# and, since §14's write-routing pass, registry::tests::
# commit_all_commits_default_and_every_opened_tenant_index and the
# tenants:: module's real-TCP-server tests for ActiveTenantTracker
# (start_fetches_and_serves_the_initial_list,
# start_fails_closed_when_the_first_fetch_fails, etc. -- see §14).
cd ../enterprise
go test ./internal/searchclient/... -v
# real in-process gRPC server, confirms SearchRequest.tenant_id is set
# correctly, that a request with no/invalid tenant identity is refused,
# and (since TenantChecker was added to close verification-plan item 4)
# that a tenant which exists but isn't active yet -- e.g. right after
# enterprise-auth -create-tenant but before enterprise-api
# -provision-tenant -- is refused too, not silently searched against a
# freshly-created empty index.
go test ./internal/chrunner/... -run MidProvisioning -v
# the ClickHouse half of the same item-4 probe -- also Docker-free,
# since an empty DataSource list never dials out.
This is the one piece of Phase 4's tenant isolation work that has
actually been run and confirmed passing in an environment without
Docker access, alongside enterprise/internal/loginhandler's OIDC/SAML
tests (§3a/§3b) — both are unusually strong evidence precisely because
they needed no infrastructure this environment lacked. Writing the
TenantChecker test here is also what found the Tantivy mid-
provisioning gap in the first place, not just what closed it after the
fact -- see api/queryapi/tenant_isolation_gap_test.go's item 4 for the
full story.
10. Confirm the Helm chart actually enforces the binary swap
No live cluster needed for this either — helm template's output is
plain YAML, parseable without a cluster:
cd deploy/helm/sentry
helm template sentry . --include-crds > /tmp/default.yaml
helm template sentry . --include-crds --set enterprise.enabled=true \
--set 'tenants[0].name=acme' --set 'tenants[0].displayName=Acme Corp' \
> /tmp/enterprise.yaml
python3 -c "
import yaml
for f in ['/tmp/default.yaml', '/tmp/enterprise.yaml']:
docs = list(yaml.safe_load_all(open(f)))
deploys = [d for d in docs if d and d.get('kind')=='Deployment' and d.get('metadata',{}).get('name')=='sentry-api']
print(f, '->', [d['spec']['template']['spec']['containers'][0]['image'] for d in deploys])
"
# expect: default.yaml -> ['sentry-api:latest'], enterprise.yaml -> ['sentry-enterprise-api:latest']
# and exactly one Deployment named sentry-api in each file.
Real cluster (kind create cluster, or similar): helm install with
each set of values and confirm kubectl get deploy sentry-api -o jsonpath='{.spec.template.spec.containers[0].image}' matches, and that
kubectl get svc sentry-api routes to whichever one is actually running.
10a. Confirm docker-compose.yml now enforces the same binary swap
Local/dev parity with §10 above was a named gap ("docker-compose.yml
still runs plain api unconditionally") -- closed via COMPOSE_PROFILES
(api/enterprise-api are each gated behind a profile, .env checks in
single-tenant as the zero-config default) plus the same "same host
port, enterprise-api gets a default.aliases: [api] network alias"
trick §10's Helm chart uses at the Service-name level. Verified in this
environment via docker compose config (no daemon needed -- it renders
and validates the merged YAML without starting anything):
docker compose config --quiet && echo "config is valid"
# exactly one of api/enterprise-api per profile, never both or neither:
docker compose config --services
# expect: ... api ... (no enterprise-api)
COMPOSE_PROFILES=enterprise docker compose config --services
# expect: ... enterprise-api ... (no api)
# enterprise-api really does take over api's name/port when active:
COMPOSE_PROFILES=enterprise docker compose config \
| python3 -c "import yaml,sys,json; d=yaml.safe_load(sys.stdin)['services']['enterprise-api']; print(json.dumps({'ports': d['ports'], 'aliases': d['networks']['default']['aliases'], 'HTTP_LISTEN_ADDR': d['environment']['HTTP_LISTEN_ADDR']}, indent=2))"
# expect port 8080 (not enterprise-api's own default 8083), alias
# ["api"], and HTTP_LISTEN_ADDR ":8080"
docker compose run/build enterprise-api (§8's provisioning steps)
work regardless of the active profile -- explicit service references on
the command line bypass profile filtering, confirmed in this
environment (the commands got past client-side profile resolution and
failed only on permission denied ... docker.sock, this environment's
already-disclosed no-Docker-daemon-access limitation, not a
profile-related error). Not verified: an actual docker compose up
against a real daemon in this environment — the config rendering above
proves the compose file's shape is correct, not that containers
actually start and route traffic correctly end to end.
11. Tenant CRD unified with -provision-tenant
deploy/helm/sentry/README.md's "Trying the two-tenant example" section
has the full helm install → -provision-tenant → kubectl get tenants walkthrough. What was actually run in this environment (no
live cluster, same limitation as §10):
cd enterprise && go test ./internal/tenantcrd/... -v
# real k8s.io/client-go fake dynamic + typed clientsets, no cluster
# needed -- proves Sync() creates/updates the Tenant object and Secret
# correctly, is idempotent, never overwrites a pre-existing
# spec.displayName, and never rotates a credential across a re-sync.
cd deploy/operator && go test ./... -v
# tenant_controller_test.go rewritten for the new split: proves an
# unprovisioned tenant reports Provisioning (not Active -- the
# regression test for the pre-unification bug where this controller
# claimed Active on its own say-so), that setting
# status.clickHouseDatabaseName (simulating what -provision-tenant
# writes) flips it to Active, and that un-suspending an
# already-provisioned tenant returns straight to Active rather than
# being demoted to Provisioning.
Helm-side wiring confirmed via helm template + parsing the rendered
YAML (not eyeballing it): with tenantOperator.enabled=true,
enterprise-api gets its own ServiceAccount/Role/RoleBinding scoped to
exactly tenants/tenants/status/secrets, tenant-operator's own
ClusterRole no longer grants secrets at all, and TENANT_CRD_NAMESPACE
is only set on enterprise-api's container when tenantOperator.enabled
is true (absent, and the ServiceAccount/Role absent too, with just
enterprise.enabled=true). Not verified: an actual -provision-tenant
run against a real cluster with the operator watching -- everything
above proves each half's logic and the chart's shape independently, not
the full loop (does the operator's watch actually re-trigger a reconcile
after -provision-tenant's external status write the way controller-
runtime's default predicate is expected to).
12. Tenant-picker, backend and frontend
Backend -- like §9, this needs nothing but a local Go toolchain --
the real fake-IdP tests already exercise the full login → pending-login
cookie → GET /auth/memberships → POST /auth/select-tenant → real
session round trip:
cd enterprise
go test ./internal/session/... -run PendingLogin -v
# real JWT signing/verification: issues a pending-login token, validates
# it, and proves the two negative regressions that matter most --
# a real session token must not validate as a pending login (they share
# a signing key but PendingLoginClaims uses a disjoint json field name,
# see that type's doc comment for the bug this caught in its own tests),
# and a pending-login token must not validate as a real session either.
go test ./internal/loginhandler/... -run 'Memberships|SelectTenant|MultipleMemberships' -v
# full round trip against the real fake OIDC/SAML IdPs: multi-membership
# login sets a pending cookie and redirects (not a 501 anymore), GET
# /auth/memberships lists the real tenant options with display names,
# POST /auth/select-tenant re-derives role server-side and refuses a
# tenant_id outside the identity's actual memberships.
Frontend -- web/src/routes/select-tenant (new), calling the
endpoints above via fetch(..., {credentials: 'include'})
($lib/api.ts's listMemberships/selectTenant), needed a CORS
posture enterprise-auth didn't have: api/httpserver.WithCORS's
wildcard-friendly default can't be combined with a credentialed
request at all (browsers refuse it outright), so this needed a new
WithCredentialedCORS (same package, literal-origin-only) wired in via
a new CORS_ALLOWED_ORIGIN env var, defaulting to
POST_LOGIN_REDIRECT_URL (web's own origin). Type-checked and built
for real:
cd web
npm run check # svelte-check -- 0 errors
npm run build # adapter-static -- confirms select-tenant.html is
# actually produced (it wasn't, at first: adapter-
# static only crawls routes reachable from a link or an
# explicit prerender entry, and nothing in the app
# links to this route since it's only ever reached via
# enterprise-auth's redirect -- fixed by adding
# select-tenant/+page.ts's `export const prerender =
# true`, the same declaration every other route here
# already has)
Genuinely verified in a real browser in this environment -- not just
type-checked, the actual cross-origin fetch/CORS/cookie behavior, driven
through mcp__claude-in-chrome:
- A throwaway Node HTTP server (no dependencies) stood in for
enterprise-auth, implementing the exact wire contract this section's Go tests already prove server-side:GET /auth/membershipsandPOST /auth/select-tenant, thesentry_pending_logincookie (Path=/auth), the credentialed CORS headers, and critically the plain-texthttp.Errorresponse bodies the real handler sends on failure (not JSON --enterpriseAuthRequestin$lib/api.tsreadsres.text()specifically because of this, unlike every other request helper in that file). npm run dev(SvelteKit dev server) pointed at that fake server viaVITE_ENTERPRISE_AUTH_BASE_URL, both onlocalhostbut different ports -- different origins, the same cross-origin shape a real deployment has.- The browser navigated to the fake server's
/debug/start-pending- login(mimicsloginhandler.startTenantSelection: sets the pending cookie, redirects to/select-tenant) -- confirmed the redirect landed on the real page, which then genuinely fetchedGET /auth/membershipscross-origin with the cookie attached and rendered both fake tenants with their display names and roles. - Clicked a tenant in the real UI -- confirmed the real
POST /auth/select-tenantfired (with preflight), succeeded, and the page navigated to the response'sredirect_urlvia a real full page load. - Reloaded
/select-tenantdirectly (no pending cookie present anymore -- the fake server clears it exactly like the real handler does) -- confirmed the page's error state renders the backend's actual plain-text message ("missing or expired pending login...") rather than a generic fetch-failure string.
No console errors at any point. This is the first frontend-only piece
in this entire phase that's been exercised in a real browser rather than
only type-checked or unit-tested against fakes -- everything else
frontend-adjacent (getAuthFeatures on the settings page, existing
Phase 0-3 routes) predates this runbook and was never re-verified here
either.
Since closed for real against the actual running containers, not a
stand-in: docker-compose.override.yml (local-only, gitignored) pointed
enterprise-auth at the same real Auth0 developer tenant §3a used. Two
real tenants (acme, globex) each got a real tenant_memberships row
for the same Auth0 identity (admin and viewer respectively, via
-grant-membership-*). Logging in through /auth/oidc/login against
the real IdP landed on the real /select-tenant page, which rendered
both real memberships with their real display names and roles ("Acme
Corp -- Admin", "Globex Corporation -- Viewer") -- fetched via a real
credentialed cross-origin request to the real enterprise-auth
container, not the fake Node stand-in. Selecting "Globex Corporation"
fired a real POST /auth/select-tenant, redirected to web, and
POST /internal/authorize with the resulting cookie returned
{"tenant_id":"globex","user_id":"...","role":"viewer"} -- proving the
picker's selection genuinely determines the issued session's tenant, not
just that the UI renders correctly. This run is also what found and
fixed the web/Dockerfile build-arg bug described in this doc's top
"Verification status" section -- the picker was unreachable against the
real container until that fix landed.
13. Ingest tenant identity
The identity mechanism was chosen deliberately (config-supplied tenant_id + a shared-secret token ingest validates, not per-tenant mTLS certs -- smaller real implementation, no new PKI). Verified in this environment without Docker or a live enterprise-auth, using the same fake-client-at-every-layer discipline as everything else in this runbook that doesn't need a live stack:
cd enterprise
go test ./internal/rbacstore/... -run IngestCredential -v
# skip-gated (RBACSTORE_TEST_POSTGRES_ADDR) -- CreateIngestCredential/
# ValidateIngestCredential/RevokeIngestCredential round trip, and the
# regression test that only a SHA-256 hash is ever persisted, never the
# plaintext token.
go test ./internal/authhandler/... -run AuthorizeIngest -v
# real HTTP round trip against POST /internal/authorize-ingest with a
# fake credential validator -- proves a session token (service or
# human) does NOT work as an ingest credential, since this endpoint
# never calls session.Manager.Validate at all.
cd ../ingest
go test ./internal/tenantresolver/... -v
# real HTTP round trip (httptest), same shape as api/authz.
# HTTPAuthorizer's own tests -- forwards the bearer token, parses
# tenant_id, treats a non-2xx or an empty tenant_id as an error.
go test ./internal/grpcserver/... -run 'Resolver|TenantHeader' -v
# PushBatch with a fake TenantResolver: no resolver configured ->
# unchanged behavior, no tenant_id header at all; resolver configured ->
# every produced Kafka message carries a tenant_id header matching the
# resolved tenant; missing or invalid bearer token -> the whole batch is
# refused (codes.Unauthenticated), fail-closed, never falls back to "no
# tenant."
Also not built as part of the identity mechanism itself: any Helm/
docker-compose.yml wiring that issues an agent a real ingest credential
automatically (enterprise-auth -create-ingest-credential-tenant=<id>
is, like every other credential-minting flag in this codebase, a manual
operator action) -- deploy/helm/sentry/values.yaml's
ingest.requireTenantCredential (default false) only turns on
validation, deliberately not folded into enterprise.enabled
directly, since flipping that flag with no agents holding a credential
yet would refuse all ingest traffic outright rather than degrading
gracefully. See §14 below for what the identity is actually used for on
the write path.
14. Ingest write-routing (both storage engines)
enterprise/cmd/enterprise-ingest (mirrors enterprise-api's "second
binary" shape) consumes the tenant_id Kafka header §13 attaches and
routes each record's ClickHouse write into that tenant's own database,
via enterprise/internal/chwriter.Registry -- a per-tenant
*ingest/clickhousewriter.Writer registry that reuses ingest/ consumer's own flush loop unchanged. Verified in this environment
without Docker, using the same fake-transport discipline as §13:
cd enterprise
go test ./internal/chwriter/... -run 'TestWriteBatchRefuses' -v
# Docker-free: constructs a Registry directly (bypassing New(), the only
# part that dials ClickHouse) to prove the fail-closed paths -- an empty
# TenantID, or a TenantID with no registered writer, refuses the WHOLE
# batch rather than silently dropping just those records or falling back
# to a default destination. (Note: -run TestRegistry, an earlier version
# of this command, actually matches TestRegistryWritesEachTenantToItsOwnDatabase/
# TestRegistryRefusesUnprovisionedTenant instead -- the live-ClickHouse
# tests below, which just skip -- not the Docker-free ones this
# paragraph is about; caught by actually running this command while
# revisiting the runbook.)
go test ./internal/tenantprovision/... -run TestProvisionedUserCanInsertIntoOwnDatabase -v
# skip-gated (needs a live ClickHouse) -- regression test for the bug
# found while building this: the per-tenant credential chwriter reuses
# from chrunner only had SELECT granted, which would make every real
# write fail with a permission error. Fixed by granting SELECT, INSERT
# (tenantprovision.go). This test proves the fix, not just documents it,
# whenever a real ClickHouse is available to run it against.
go test ./internal/chwriter/... -run TestRegistryWritesEachTenantToItsOwnDatabase -v
# skip-gated (CHWRITER_TEST_CLICKHOUSE_ADDR) -- writes a mixed batch
# spanning two tenants in one WriteBatch call and confirms each row
# lands in its own tenant's database, none in the other's. (Fixed from a
# stale `-run TestRegistryRoutesToCorrectTenant` -- that name never
# existed in this package; caught by actually running this command while
# closing out the live verification pass, same class of stale--run
# mistake §6's `TestRegistry` note already flagged once in this doc.)
Now genuinely run against a live ClickHouse (previously only
skip-gated, correct Go that had never executed): all 7
internal/chwriter tests pass, including
TestRegistryWritesEachTenantToItsOwnDatabase,
TestRegistryRefusesUnprovisionedTenant, TestRefreshAddsNewlyActiveTenant,
and TestRefreshRemovesNoLongerActiveTenant -- write-routing is confirmed,
not just written, as of this pass.
Tantivy's side is built too, and genuinely verified.
search/src/consumer.rs reads the same tenant_id Kafka header and
resolves it through search/src/registry.rs's IndexRegistry -- the
same registry the read side (§9) already uses -- routing each record's
write into its own tenant's Tantivy index instead of the single default
one. No "second binary" was needed here the way ClickHouse needed
enterprise-ingest: Tantivy has no grant system to gate a
commercially-licensed credential behind, so IndexRegistry already
lives directly in AGPL-core search, and read/write just share it.
Because Tantivy is an embedded library (no Docker/broker needed to
exercise real logic), this actually ran in this environment:
cd search
cargo test --quiet
# registry.rs: commit_all_commits_default_and_every_opened_tenant_index
# writes into the default index plus two tenant indices, confirms
# nothing is searchable before commit_all(), then confirms all three ARE
# searchable after -- the write-routing + periodic-commit path
# end-to-end, for real, using the same real-Tantivy-index discipline as
# every other registry.rs/index.rs test. consumer.rs:
# tenant_id_from_headers_* cover the header-extraction helper
# (missing/present/unrelated-header cases) and
# test_tenant_id_header_key_matches_go guards the "tenant_id" literal
# against drifting from ingest/consumer.TenantIDHeaderKey /
# ingest/internal/grpcserver.TenantIDHeaderKey the same way
# ingest/cmd/ingest's own guard test does on the Go side.
Tantivy's write path is now active-tenant-gated too.
search/src/tenants.rs's ActiveTenantTracker polls a new
GET /internal/active-tenants endpoint on enterprise-auth (Go side:
rbacstore.ListActiveTenantIDs + authhandler.handleActiveTenants,
RoleService-credentialed -- mint one with
enterprise-auth -mint-service-token search, the same generic flag
alerting already uses, just a different subject name) and
consumer.rs refuses any tagged record whose tenant isn't in the polled
set. Off unless ENTERPRISE_AUTH_URL/ENTERPRISE_AUTH_SERVICE_TOKEN are
both set (search/src/config.rs rejects exactly one being set); startup
blocks on the first fetch succeeding (fail-closed cold start), and a
later refresh failure keeps serving the last-known-good set rather than
clearing it. Genuinely verified in this environment, no live
enterprise-auth needed:
cd search
cargo test --quiet tenants::
# start_fetches_and_serves_the_initial_list / start_sends_the_bearer_token
# run against a real (hand-rolled, dependency-free) TCP server -- actual
# reqwest request construction (the Authorization: Bearer header, the
# /internal/active-tenants path) and actual JSON response parsing, not a
# fake HTTP client. start_fails_closed_when_the_first_fetch_fails and
# start_fails_closed_when_the_server_is_unreachable prove the cold-start
# refusal: ActiveTenantTracker::start returns Err rather than falling
# back to an empty (accept-nothing, silently-safe-looking-but-wrong-for-
# operators) or permissive (accept-anything, the exact bug being closed)
# default.
go test ./internal/authhandler/... -run TestActiveTenants -v
# GET /internal/active-tenants requires a RoleService credential -- a
# valid human session (even for a real Owner) is rejected the same way
# an invalid/missing token is, the regression test for this endpoint's
# whole reason to distinguish token kinds.
The ClickHouse/Tantivy asymmetry this section used to disclose is now
closed too. chwriter.Registry.StartRefreshing (new) spawns a
goroutine that re-lists active tenants every minute (matching
ActiveTenantTracker's interval -- see
enterprise-ingest/main.go's dataSourceRefreshInterval) and
reconciles the writer map: opens a connection for a newly-active tenant,
closes and removes one no longer active. A tenant deprovisioned after
enterprise-ingest startup now loses ClickHouse write access within a
minute, not "until the next restart." A refresh failure (rbacstore
unreachable, or one tenant's new connection failing to open) logs and
leaves the existing map untouched for that tick -- the same
last-known-good posture ActiveTenantTracker uses, so a transient
Postgres blip doesn't evict every other tenant's already-working writer.
Neither engine does a live per-write check (a database/HTTP round trip
per record would be a real throughput cost neither implementation
accepts), so a roughly one-minute staleness window remains on both
sides by design, not a gap unique to either anymore.
cd enterprise
go test ./internal/chwriter/... -run TestRefreshListerErrorLeavesRegistryUnchanged -v
# Docker-free -- refresh's early-return on a lister error, proven the
# same way the existing fail-closed tests are: a Registry constructed
# directly (bypassing New, so nothing dials ClickHouse), asserting
# WriteBatch still refuses afterward.
go test ./internal/chwriter/... -run 'TestRefreshAddsNewlyActiveTenant|TestRefreshRemovesNoLongerActiveTenant' -v
# skip-gated (CHWRITER_TEST_CLICKHOUSE_ADDR) -- the actual add/remove
# reconciliation against real connections: a tenant absent at New() time
# gains a working writer after refresh() sees it in a later lister call;
# a tenant present at New() time loses its writer (WriteBatch starts
# refusing it) after refresh() stops seeing it.
Not run in this environment: same live-ClickHouse caveat as
everything else here -- the two new skip-gated tests are correct Go
that's never executed against a real database. See
/docs/security/threat-model.md's "Read this first".
Also not built: Helm/docker-compose.yml do gate whether
enterprise-ingest runs at all (ingest.requireTenantCredential, same
flag §13 uses for validation -- see deploy/helm/sentry/templates/ enterprise-ingest.yaml), but docker-compose.yml's version is a
disclosed, weaker approximation of Helm's: Helm achieves genuine
-mode=server/-mode=consumer mutual exclusivity between ingest and
enterprise-ingest; compose's enterprise-ingest service is a
profile-gated opt-in extra that does NOT split ingest's own
server/consumer halves, so with the enterprise profile active, both
ingest -mode=all and enterprise-ingest independently consume every
message via different consumer groups -- harmless duplication for local
verification, not a topology compose actually enforces the way Helm
does.
Known gaps (do not treat this phase as done without reading these)
Full accounting: /docs/security/threat-model.md. Headline items:
-
Both storage engines' isolation exists, and both Helm and docker-compose now enforce which binary runs.
deploy/helm/sentry/templates/api.yaml/enterprise-api.yamlare mutually exclusive onenterprise.enabled(§10) -- a Helm-deployed cluster can't accidentally run the non-isolated binary once that flag is set.docker-compose.yml'sapi/enterprise-apiservices are now the same mutually-exclusive choice viaCOMPOSE_PROFILES(§8, §10a), closing the local/dev parity gap this bullet used to name. -
The
TenantCRD (deploy/operator) andenterprise-api -provision-tenantare now unified, in a deliberately lightweight way:-provision-tenantstays the sole real actor (ClickHouse +rbacstore) and, whenTENANT_CRD_NAMESPACEis set (the Helm chart does this automatically whentenantOperator.enabled), also syncs the real result into the Tenant CRD (enterprise/internal/tenantcrd) -- see §11 below. Running-provision-tenantis still a separate, deliberately manual operator action fromhelm install/kubectl apply -f tenant.yamlcreating the Tenant object in the first place -- that split (declarative request vs. imperative provisioning action) is intentional, not the "two disconnected sources of truth" gap this bullet used to describe. -
Ingest now has a real tenant identity (§13), and both storage engines' write-routing is built (§14). An agent presents a bearer credential (
enterprise-auth -create-ingest-credential-tenant=<id>),ingest/internal/grpcserver.TenantResolvervalidates it (fail-closed) and attaches the resolved tenant ID to every record as atenant_idKafka message header.enterprise-ingestreads that header back and routes each record's ClickHouse write into its own tenant's database (not yet confirmed against a real ClickHouse in this environment -- see §14).search/src/consumer.rsreads the same header and routes each record into its own tenant's Tantivy index -- genuinely verified in this environment, unlike the ClickHouse side, since Tantivy needs no Docker to exercise real logic. Both engines now also gate writes on an active-tenant check that refreshes every minute --chwriter. Registry.StartRefreshingon the ClickHouse side (new, closing what was originally a startup-only snapshot with no refresh at all) and Tantivy'sActiveTenantTrackeron the other, the same interval on both, no asymmetry left between them; see §14 and/docs/security/threat-model.md's "Read this first". A newly-provisioned tenant's ClickHouse database and Tantivy index are both now real, isolated, and actually populated by write-routed traffic (the ClickHouse claim pending live confirmation, the Tantivy claim already verified) -- what used to be "permanently empty" for both is no longer true for either. -
Human SSO login now works for both OIDC (§3a) and SAML (§3b) -- each verified with a real fake IdP (genuine cryptographic signing and verification), not yet a real external IdP or a running
enterprise-authcontainer. The tenant-picker, backend and frontend, is now fully built too (§12) --GET /auth/memberships/POST /auth/select-tenant, backed by a short-lived pending-login token distinct from a real session, andweb/src/routes/select-tenantactually calls it via credentialed cross-originfetch, genuinely exercised in a real browser against a contract-accurate fake backend (§12). What's still not tried is the same caveat as OIDC/SAML above: this whole round trip end-to-end against a real runningenterprise-authcontainer instead of a stand-in. -
No admin UI to create a
tenant_membershipsrow, but §3a/§3b's manual SQL bootstrap is gone --enterprise-auth -create-tenant/-grant-membership-*/-revoke-membership-*/-list-memberships-tenant/-transfer-owner-*(offline operator flags, same shape as-mint-service-token) cover create/grant/revoke/list/transfer-owner. Changing a non-Owner role after the fact is just re-running-grant-membership-*with a different-grant-membership-role(SetMembership's upsert already supports it).RevokeMembershipstill refuses a tenant's current Owner (would leavetenants.owner_user_iddangling), but that's no longer a dead end ---transfer-owner-tenant/-transfer-owner-user-email(rbacstore.TransferOwner, this package's first use of a real transaction: downgrades the current owner to admin, promotes the new owner, and updatestenants.owner_user_idatomically) hands ownership off first, and the now-downgraded former owner can be revoked normally after that.-grant-membership-role=owneritself now refuses when a different owner already exists, pointing at-transfer-owner-*instead of silently leaving a staletenant_membershipsrow. Skip-gated live-Postgres tests (TestTransferOwnerMovesOwnershipAndDowngradesPreviousOwnerand two refusal-path tests inenterprise/internal/rbacstore/rbacstore_test.go) haven't run against a live database in this environment, same gap as the rest of this phase's Postgres-backed pieces. -
Per-resource dashboard grants are now enforced (
api/dashboards' handler readsdashboard_permissionsviaenterprise/internal/rbacstore.DashboardPermissions, only whenenterprise-api-- not plainapi-- serves traffic), and now has a CLI surface:sentryctl dashboards permissions list|grant|revoke(cli/cmd/sentryctl/cmd_dashboards.go) againstGET/PUT/DELETE /dashboards/{id}/permissions/{userId}-- still nowebUI for it, just the CLI. Verified against a realhttptest. Server(not a fake store this time -- the CLI has no store of its own, just an HTTP client, so this is exercising real request construction/method/path/body/error-parsing, the same pattern every othersentryctlsubcommand's tests use):cd cli go test ./... -run TestCmdDashboardsPermissions -vapi/dashboards' own handler tests (fakePermissionStore) and the real-Postgresenterprise/internal/rbacstore/rbacstore_test.gointegration tests are unchanged by this -- the CLI is a new caller of an existing, already-tested endpoint, not new server-side logic. Those Postgres-backed tests still haven't run against a live database in this environment, same gap as the rest of this phase's Postgres-backed pieces. -
All four adversarial ClickHouse/Tantivy probes named in
/docs/phase-4-isolation-design.md's verification plan are now closed -- seeapi/queryapi/tenant_isolation_gap_test.gofor the full accounting. The fourth (mid-provisioning-race handling) turned out to need a real code fix on the Tantivy side, not just a test:search/ src/registry.rs'sIndexRegistryopened-or-created an index for any syntactically-validtenant_id, so a mid-provisioning tenant's query would have silently succeeded with zero results instead of being refused. Fixed viaenterprise/internal/searchclient.TenantChecker(backed byrbacstore.TenantIsActive). Both the ClickHouse and Tantivy halves of this probe run genuinely, without Docker, in this environment (chrunner_test.go'sTestRegistryRefusesMidProvisioningTenant,searchclient_test.go'sTestSearchRefusesMidProvisioningTenant) --rbacstore.TenantIsActiveitself has skip-gated live-Postgres tests that haven't run here, same disclosed gap as the rest of this phase's Postgres-backed pieces.
Tearing down
docker compose down -v
helm uninstall sentry # if installed against a real cluster
Troubleshooting
enterprise-auth fails to start with "ENTERPRISE_SESSION_SIGNING_KEY
must be set to at least 32 bytes".
Required, unlike OIDC/SAML config — see enterprise/internal/config.Load.
docker-compose.yml's dev value is long enough; a custom override must
be too.
POST /query returns 401 even though ENTERPRISE_AUTH_URL isn't
set.
Check api/cmd/api/main.go actually left authorizer nil when
cfg.EnterpriseAuthURL == "" — a nil Authorizer must be a no-op
(api/authz.RequireRole's doc comment). If this regresses, it
breaks every existing Phase 0-3 deployment silently.
A dashboard created by one tenant is visible to another.
This is the exact bug found and fixed in task 7 — see
/docs/security/threat-model.md's "application-layer tenant scoping"
section and api/dashboards/handler_test.go's
TestCrossTenant* tests. If this regresses, Handler.tenantID or
store.go's WHERE tenant_id = ... filters have been bypassed
somewhere — check every store method still takes and uses a tenantID
parameter.
helm template fails with error calling include: ... can't evaluate field Release in type string.
A call site is passing a bare string to sentry.selectorLabels instead
of (list $ "name") — see templates/_helpers.tpl's doc comment for
why the plain-string form doesn't work with include.