Files
cairnobs/ingest
jcoffey-dev 17fdc212c2 Give ingest a real tenant identity (write-routing deferred, disclosed)
Ingest tenant-awareness was named "undesigned, not just unbuilt" across
CLAUDE.md/threat-model.md/the runbook since early Phase 4 -- the last
major standing gap. Scoping was agreed via AskUserQuestion: a
config-supplied tenant_id + shared-secret token ingest validates
(smaller real implementation, no new PKI), over per-tenant mTLS
certs. This change builds that identity mechanism end to end and
attaches it to every record at the point it enters the system; it
deliberately does NOT build per-tenant write-routing for ClickHouse or
Tantivy -- that's real, separately-scoped follow-up work, disclosed
explicitly everywhere this was previously called undesigned, not
silently left half-done.

New pieces:

- metadata/migrations/0034 + enterprise/internal/rbacstore/
  ingest_credentials.go: a per-tenant bearer credential, only its
  SHA-256 hash ever persisted (same reasoning a password gets hashed,
  not stored raw) -- CreateIngestCredential returns the plaintext
  exactly once, ValidateIngestCredential/RevokeIngestCredential/
  ListIngestCredentialsForTenant round it out.
- enterprise-auth gains -create-ingest-credential-tenant/
  -list-ingest-credentials-tenant/-revoke-ingest-credential (same
  offline-operator-flag shape as every other credential-minting flag in
  this binary) and a new POST /internal/authorize-ingest endpoint
  (internal/authhandler) validating a presented token and resolving its
  tenant -- a genuinely different credential type from session-backed
  /internal/authorize, so it doesn't touch session.Manager at all.
- ingest (AGPL core) gains an optional TenantResolver
  (internal/grpcserver, nil by default) and its HTTP client
  implementation (internal/tenantresolver.HTTPResolver) -- a plain HTTP
  call to enterprise-auth's new endpoint, never an enterprise/ import,
  same "network boundary, not import boundary" shape
  api/authz.HTTPAuthorizer already uses for the query path.
  PushBatch now requires an `authorization: Bearer <token>` gRPC
  metadata entry once a resolver is configured, fails the whole batch
  closed on a missing/invalid credential (never falls back to "no
  tenant"), and attaches the resolved tenant ID to every record as a
  `tenant_id` Kafka message header before producing it.

Verified with real round trips at every layer, no Docker needed:
rbacstore's credential CRUD (skip-gated on live Postgres, same as every
other rbacstore integration test this phase), authhandler's new
endpoint (real HTTP via httptest, including the regression test that a
session token must not validate as an ingest credential), tenantresolver
(real HTTP client against httptest, same pattern as
authz.HTTPAuthorizer's own tests), and grpcserver's PushBatch (fake
resolver/producer -- no resolver leaves messages unchanged, a configured
resolver attaches the right header or fails closed on a bad/missing
token).

Helm: ingest.requireTenantCredential (default false) is a deliberate,
separate opt-in from enterprise.enabled -- turning ENTERPRISE_AUTH_URL
on for ingest requires every agent to already hold a credential or be
refused outright, so it must not default on just because
enterprise.enabled does (same reasoning api.yaml's ENTERPRISE_AUTH_URL
isn't tied to enterprise.enabled directly either). docker-compose.yml
leaves it unset, same as ever.

Docs updated everywhere this was called "undesigned": CLAUDE.md,
docs/architecture.md, docs/security/threat-model.md (including its
summary table, now split into "identity: built" vs "write-routing: not
yet"), docs/phase-4-runbook.md (new §13), enterprise/README.md.
2026-08-14 15:21:55 -07:00
..

ingest

Go service sitting between the Rust agent and ClickHouse. Two halves in one binary, selected with --mode:

  • server — mTLS gRPC front end (LogIngest.PushBatch) that agents connect to. Assigns each record a server-side record_id (a UUID, overwriting whatever the agent sent — agents always send it empty) and otherwise forwards records proto-encoded onto Redpanda unchanged. Still kept thin — one field assignment, no real normalization — so agent- facing latency isn't coupled to ClickHouse write performance. record_id has to be assigned exactly once, here, rather than independently by each downstream consumer: Phase 1's Tantivy indexer and the ClickHouse writer both read the same Redpanda messages and need to agree on the same ID for the same record to join search hits back to rows — two consumers generating their own IDs would produce mismatched ones for what's supposed to be the same record.
  • consumer — reads back off Redpanda, normalizes into the ClickHouse row shape (internal/normalize), and batch-writes via the native protocol driver. Commits Redpanda offsets only after a successful ClickHouse write, so a ClickHouse outage causes redelivery on restart rather than data loss.
  • all (default) — both, in one process. This is what docker-compose runs. Splitting into two deployments later (e.g. to scale them independently in k8s) is a manifest change, not a code change — see --mode.

Why Redpanda stays in the path

Confirmed with the project owner during Phase 0 planning: the gRPC front end produces to Redpanda rather than writing ClickHouse directly. This exercises the pinned transport layer from day one and keeps agents from ever needing Kafka credentials — mTLS to ingest is the only network egress an agent has. See /docs/architecture.md.

Dependencies worth knowing about

  • github.com/segmentio/kafka-go — pure Go, no cgo, chosen over franz-go/confluent-kafka-go specifically to keep the distroless build simple (confirmed with the project owner; see git history / PR discussion for the tradeoffs considered).
  • github.com/ClickHouse/clickhouse-go/v2 — official client, native protocol, pure Go (no cgo).
  • golang.org/x/sync/errgroup — used in cmd/ingest/main.go to run the server and consumer halves concurrently and propagate the first error.
  • github.com/google/uuid — was already in the dependency graph transitively (via clickhouse-go); promoted to a direct dependency for record_id generation in internal/grpcserver, so not a new addition to the transitive tree.

Configuration

All via environment variables (see internal/config/config.go for the full list and defaults) — no config file format for Phase 0:

Var Default Purpose
GRPC_LISTEN_ADDR :4317 Agent-facing gRPC listen address
TLS_CERT_FILE / TLS_KEY_FILE /etc/sentry-ingest/server{,-key}.pem ingest's own mTLS identity
TLS_CLIENT_CA_FILE /etc/sentry-ingest/ca.pem CA used to verify agent client certs
REDPANDA_BROKERS localhost:9092 Comma-separated broker list
REDPANDA_TOPIC sentry.logs.raw Must match the topic provisioned in /transport
REDPANDA_CONSUMER_GROUP sentry-ingest Consumer group id
CLICKHOUSE_ADDR localhost:9000 Native protocol port, not HTTP
CLICKHOUSE_DATABASE / _USERNAME / _PASSWORD sentry / default / ``
CONSUMER_BATCH_MAX_SIZE 500 Records per ClickHouse batch insert
CONSUMER_BATCH_FLUSH_INTERVAL_MS 2000 Max time a partial batch waits before flushing

Building & testing

go build ./...
go vet ./...
go test ./...

Requires google.golang.org/protobuf/cmd/protoc-gen-go and google.golang.org/grpc/cmd/protoc-gen-go-grpc only if you're regenerating /proto's Go bindings — ingest itself just imports the already-generated github.com/sentry/sentry/proto module (see the replace directive in go.mod, pointing at ../proto).

# from the repo root, not ingest/
docker build -f ingest/Dockerfile -t sentry-ingest .

Testing notes

internal/consumer and internal/grpcserver depend on Redpanda and ClickHouse only through small interfaces (reader/chWriter in consumer, batchProducer in grpcserver), so the flush/commit/error-handling logic is unit-tested against fakes — no embedded broker or database needed. What's not covered by these tests: the real kafka.Reader/kafka.Writer wiring and the ClickHouse native-protocol driver itself. Those are only exercised by the docker-compose end-to-end flow described in /docs/phase-0-runbook.md.