Build per-tenant ClickHouse write-routing for ingest (Tantivy still deferred)

ingest tags every record with a tenant_id Kafka header (built previously),
but nothing consumed it to actually route the write. This closes that for
ClickHouse: enterprise/cmd/enterprise-ingest (a second binary, mirroring
enterprise-api) reuses ingest/consumer's own flush loop unchanged, with
enterprise/internal/chwriter.Registry -- a per-tenant clickhousewriter.Writer
registry -- swapped in as the writer. A batch pulled from the single shared
Redpanda topic can mix records from many tenants, so WriteBatch groups by
TenantID and dispatches each group to its own tenant's connection, fail-
closed on an empty or unrecognized tenant_id.

ingest/consumer and ingest/clickhousewriter move out of internal/ (same
reason api/internal/* moved earlier this phase: enterprise/ can't import
anything under another module's internal/). Their New() constructors now
take small local Config structs instead of ingest/internal/config types,
so enterprise/ doesn't need that import either.

Building this surfaced a real bug: tenantprovision.ProvisionClickHouse
only granted SELECT on a tenant's ClickHouse user, correct for chrunner's
read-only use but not enough for chwriter reusing the same credential to
write -- every real per-tenant write would have failed closed with a
permission error. Fixed by widening the grant to SELECT, INSERT; no
cross-tenant boundary is crossed by also allowing INSERT within a
tenant's own database.

Helm gates enterprise-ingest's Deployment on the same
ingest.requireTenantCredential flag that already gates tag validation --
write-routing is meaningless without tagging already being required, so
they're one decision, not two. docker-compose.yml's version is a
disclosed, weaker approximation: it can't achieve Helm's genuine
-mode=server/-mode=consumer split, so with the enterprise profile active
both ingest and enterprise-ingest independently consume every message
via different consumer groups -- harmless duplication for local
verification only.

Not built: Tantivy's independent Redpanda consumer (search/src/consumer.rs)
still doesn't read the tenant_id header at all -- every record still lands
in the one shared index regardless of tenant. Not run: the live-ClickHouse-
gated tests (chwriter's cross-tenant routing test, tenantprovision's INSERT
regression test) -- no Docker/database access in this environment; they're
correct Go that has never executed, disclosed as such in docs/security/
threat-model.md and docs/phase-4-runbook.md §14.
This commit is contained in:
2026-08-14 19:26:09 -07:00
parent 17fdc212c2
commit 1de77b969f
26 changed files with 1355 additions and 267 deletions
@@ -0,0 +1,64 @@
{{/*
Only rendered once per-tenant write-routing is actually turned on (see
ingest.yaml's -mode=server comment) -- this Deployment is what takes
over consuming sentry.logs.raw once ingest.yaml's own consumer half
stops, routing each record to its own tenant's dedicated ClickHouse
database (enterprise/internal/chwriter) instead of the one shared table.
No Service: this is a pure background worker, nothing calls it, only
kubelet's own probes talk to its /healthz.
*/}}
{{- if and .Values.enterprise.enabled .Values.ingest.requireTenantCredential }}
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ .Release.Name }}-enterprise-ingest
labels:
{{- include "sentry.labels" . | nindent 4 }}
{{- include "sentry.selectorLabels" (list $ "enterprise-ingest") | nindent 4 }}
spec:
replicas: {{ .Values.ingest.replicas }}
selector:
matchLabels:
{{- include "sentry.selectorLabels" (list $ "enterprise-ingest") | nindent 6 }}
template:
metadata:
labels:
{{- include "sentry.selectorLabels" (list $ "enterprise-ingest") | nindent 8 }}
spec:
initContainers:
{{- include "sentry.waitForTCP" (list "redpanda" (printf "%s-redpanda" .Release.Name) "9092") | nindent 8 }}
{{- include "sentry.waitForTCP" (list "clickhouse" (printf "%s-clickhouse" .Release.Name) "9000") | nindent 8 }}
{{- include "sentry.waitForTCP" (list "postgres" (printf "%s-postgres" .Release.Name) "5432") | nindent 8 }}
containers:
- name: enterprise-ingest
image: "{{ .Values.enterprise.ingestImage.repository }}:{{ .Values.enterprise.ingestImage.tag }}"
imagePullPolicy: {{ .Values.global.imagePullPolicy }}
env:
- name: REDPANDA_BROKERS
value: "{{ .Release.Name }}-redpanda:9092"
- name: CLICKHOUSE_ADDR
value: "{{ .Release.Name }}-clickhouse:9000"
- name: POSTGRES_ADDR
value: "{{ .Release.Name }}-postgres:5432"
- name: POSTGRES_DATABASE
value: sentry_metadata
- name: POSTGRES_USERNAME
value: sentry
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: {{ .Release.Name }}-postgres
key: password
readinessProbe:
exec:
command: ["/enterprise-ingest", "-healthcheck"]
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
exec:
command: ["/enterprise-ingest", "-healthcheck"]
initialDelaySeconds: 10
periodSeconds: 10
resources:
{{- toYaml .Values.ingest.resources | nindent 12 }}
{{- end }}
+12
View File
@@ -22,6 +22,18 @@ spec:
- name: ingest
image: "{{ .Values.ingest.image.repository }}:{{ .Values.ingest.image.tag }}"
imagePullPolicy: {{ .Values.global.imagePullPolicy }}
{{- if and .Values.enterprise.enabled .Values.ingest.requireTenantCredential }}
# -mode=server only: this Deployment stops running the
# ClickHouse-writing consumer half (the default -mode=all)
# once per-tenant write-routing is on -- enterprise-ingest.yaml
# (below) takes over consuming sentry.logs.raw instead, so it
# can write each tenant's records to their own database rather
# than the one shared table ingest's own consumer always
# writes to. The agent-facing server half (PushBatch, tenant
# tagging via TenantResolver) keeps running here unconditionally
# either way -- only which process consumes the topic changes.
args: ["-mode=server"]
{{- end }}
env:
- name: REDPANDA_BROKERS
value: "{{ .Release.Name }}-redpanda:9092"
+16 -1
View File
@@ -76,7 +76,13 @@ ingest:
# its own deliberate opt-in, not folded into enterprise.enabled
# directly: turning it on requires every agent to already present a
# valid ingest credential (`enterprise-auth
# -create-ingest-credential-tenant=<id>`) or be refused outright.
# -create-ingest-credential-tenant=<id>`) or be refused outright. Also
# controls per-tenant write-routing: true switches ingest.yaml's
# Deployment to -mode=server only and renders enterprise-ingest.yaml
# to take over consuming sentry.logs.raw, writing each tenant's
# records into their own ClickHouse database instead of the one
# shared table -- both flags gate together since write-routing is only
# meaningful once records actually carry a tenant_id to route on.
requireTenantCredential: false
search:
@@ -156,6 +162,15 @@ enterprise:
apiImage:
repository: sentry-enterprise-api
tag: latest
# enterprise-ingest (templates/enterprise-ingest.yaml) -- only
# rendered when ingest.requireTenantCredential is also true (see that
# value's comment); takes over consuming sentry.logs.raw from
# ingest.yaml's own consumer once per-tenant write-routing is on. Same
# "repo root build context" reasoning as apiImage above -- see
# enterprise/cmd/enterprise-ingest/Dockerfile.
ingestImage:
repository: sentry-enterprise-ingest
tag: latest
replicas: 1
resources: {}
# Leave empty to auto-generate (>= 32 bytes) and persist across