Give ingest a real tenant identity (write-routing deferred, disclosed)

Ingest tenant-awareness was named "undesigned, not just unbuilt" across
CLAUDE.md/threat-model.md/the runbook since early Phase 4 -- the last
major standing gap. Scoping was agreed via AskUserQuestion: a
config-supplied tenant_id + shared-secret token ingest validates
(smaller real implementation, no new PKI), over per-tenant mTLS
certs. This change builds that identity mechanism end to end and
attaches it to every record at the point it enters the system; it
deliberately does NOT build per-tenant write-routing for ClickHouse or
Tantivy -- that's real, separately-scoped follow-up work, disclosed
explicitly everywhere this was previously called undesigned, not
silently left half-done.

New pieces:

- metadata/migrations/0034 + enterprise/internal/rbacstore/
  ingest_credentials.go: a per-tenant bearer credential, only its
  SHA-256 hash ever persisted (same reasoning a password gets hashed,
  not stored raw) -- CreateIngestCredential returns the plaintext
  exactly once, ValidateIngestCredential/RevokeIngestCredential/
  ListIngestCredentialsForTenant round it out.
- enterprise-auth gains -create-ingest-credential-tenant/
  -list-ingest-credentials-tenant/-revoke-ingest-credential (same
  offline-operator-flag shape as every other credential-minting flag in
  this binary) and a new POST /internal/authorize-ingest endpoint
  (internal/authhandler) validating a presented token and resolving its
  tenant -- a genuinely different credential type from session-backed
  /internal/authorize, so it doesn't touch session.Manager at all.
- ingest (AGPL core) gains an optional TenantResolver
  (internal/grpcserver, nil by default) and its HTTP client
  implementation (internal/tenantresolver.HTTPResolver) -- a plain HTTP
  call to enterprise-auth's new endpoint, never an enterprise/ import,
  same "network boundary, not import boundary" shape
  api/authz.HTTPAuthorizer already uses for the query path.
  PushBatch now requires an `authorization: Bearer <token>` gRPC
  metadata entry once a resolver is configured, fails the whole batch
  closed on a missing/invalid credential (never falls back to "no
  tenant"), and attaches the resolved tenant ID to every record as a
  `tenant_id` Kafka message header before producing it.

Verified with real round trips at every layer, no Docker needed:
rbacstore's credential CRUD (skip-gated on live Postgres, same as every
other rbacstore integration test this phase), authhandler's new
endpoint (real HTTP via httptest, including the regression test that a
session token must not validate as an ingest credential), tenantresolver
(real HTTP client against httptest, same pattern as
authz.HTTPAuthorizer's own tests), and grpcserver's PushBatch (fake
resolver/producer -- no resolver leaves messages unchanged, a configured
resolver attaches the right header or fails closed on a bad/missing
token).

Helm: ingest.requireTenantCredential (default false) is a deliberate,
separate opt-in from enterprise.enabled -- turning ENTERPRISE_AUTH_URL
on for ingest requires every agent to already hold a credential or be
refused outright, so it must not default on just because
enterprise.enabled does (same reasoning api.yaml's ENTERPRISE_AUTH_URL
isn't tied to enterprise.enabled directly either). docker-compose.yml
leaves it unset, same as ever.

Docs updated everywhere this was called "undesigned": CLAUDE.md,
docs/architecture.md, docs/security/threat-model.md (including its
summary table, now split into "identity: built" vs "write-routing: not
yet"), docs/phase-4-runbook.md (new §13), enterprise/README.md.
This commit is contained in:
2026-08-14 15:21:55 -07:00
parent d2c76aa3a4
commit 17fdc212c2
21 changed files with 1071 additions and 68 deletions
+24 -8
View File
@@ -228,14 +228,30 @@ than one `tenant_memberships` row now gets a real `GET
pending-login token, distinct from a real session by both Go type and
JWT claim name — a real token-confusion bug this design's own tests
caught before it shipped) instead of the flat refusal Phase 4 shipped
with earlier. What still keeps this phase from being done: the actual
tenant-picker *page* doesn't exist (`web` has no session/cookie-handling
code at all yet, and `enterprise-auth` has no CORS middleware for a
cross-origin `fetch` with credentials — both real, separately-scoped
frontend gaps), and ingest itself has no tenant concept for either
storage engine (every record lands in the one shared ClickHouse database
and Tantivy index no matter what — undesigned, not just unbuilt). Full
accounting:
with earlier. Ingest tenant-awareness — the gap this section used to
call "undesigned" — now has a real, if intentionally partial, design:
`ingest` (AGPL core) gained an optional `TenantResolver`
(`ingest/internal/grpcserver`), a per-tenant bearer credential an agent
presents (minted via `enterprise-auth
-create-ingest-credential-tenant=<id>`, validated over the network via a
new `POST /internal/authorize-ingest` endpoint — never an `enterprise/`
import, same boundary shape as `api/authz.Authorizer`), and the
resolved tenant ID is attached to every record as a `tenant_id` Kafka
message header before it's produced. **What's still deferred, clearly**:
nothing downstream reads that header yet — neither `ingest`'s own
ClickHouse writer nor `search`'s independent Redpanda consumer route a
record's write into a per-tenant destination, so every record still
lands in the one shared ClickHouse database/Tantivy index regardless of
which tenant it's now correctly tagged with. That write-routing split
(likely another "second binary," mirroring `enterprise-api`) is real,
scoped, remaining work — attaching a verified tenant identity as early
as possible was deliberately built as a self-contained first step, not
the whole feature. What still keeps this phase from being done: the
actual tenant-picker *page* doesn't exist (`web` has no session/cookie-
handling code at all yet, and `enterprise-auth` has no CORS middleware
for a cross-origin `fetch` with credentials — both real, separately-
scoped frontend gaps), and per-tenant write-routing for ingest per the
above. Full accounting:
`/docs/security/threat-model.md`; step-by-step verification procedure
(not yet run against a live cluster in this environment):
`/docs/phase-4-runbook.md`. The rest of this section describes the exit