End-to-end log pipeline for Linux hosts, per /docs/architecture.md: - proto: shared gRPC contract (agent <-> ingest), Go bindings checked in - agent: Rust, musl-targeted, journald/file sourcing, RFC5424 parser, mTLS gRPC client, no required config for the common case - ingest: Go, single binary with --mode server|consumer|all; gRPC front end forwards to Redpanda unchanged, consumer normalizes and batch-writes to ClickHouse with at-least-once delivery - storage: ClickHouse schema + a plain SQL-file migration runner - api: minimal SELECT-only query endpoint, plain REST (not gRPC+gateway yet -- see api/README.md) - web: SvelteKit static SPA, one query page - transport: Redpanda compose + topic provisioning - cli: sentryctl ping stub - hack/dev-certs: throwaway CA + cert generation for local mTLS - root docker-compose.yml + docs/phase-0-runbook.md tie it together Not yet run end-to-end against real Docker/ClickHouse/Redpanda -- see the runbook's caveats section before relying on this working as-is.
storage
ClickHouse schema and migration tooling for Sentry's analytical store.
Schema
One table for Phase 0, logs:
CREATE TABLE logs
(
`timestamp` DateTime64(9, 'UTC'),
`host` String,
`service` String,
`severity` LowCardinality(String),
`message` String,
`attributes` Map(String, String)
)
ENGINE = MergeTree
PARTITION BY toDate(timestamp)
ORDER BY (service, timestamp)
Notes on choices that weren't fully specified by the task description:
DateTime64(9, 'UTC')(nanosecond precision) rather than second or millisecond precision, to match the agent'stimestamp_unix_nanofield end to end without truncation.severityasLowCardinality(String), not a numeric OTelSeverityNumber./ingest'snormalizepackage writes short text values (TRACE/DEBUG/INFO/WARN/ERROR/FATAL/UNSPECIFIED).LowCardinalitygets you most of the storage/query efficiency of an enum without committing to one at the schema level. Splitting into a properSeverityNumber+SeverityTextpair (full OTel shape) is one of the open questions already flagged in/docs/architecture.md.PARTITION BY toDate(timestamp)(daily partitions) andORDER BY (service, timestamp)are exactly what the task asked for — service-scoped queries over a time range are the dominant access pattern this is optimized for.- No TTL/retention clause yet — also an open question in architecture.md, deferred until storage sizing is a real concern.
Migration tooling: a plain SQL-file runner, not golang-migrate
migrate.sh applies migrations/*.sql in filename order over
ClickHouse's HTTP interface, tracking what's applied in a
schema_migrations table. Chosen over golang-migrate for Phase 0
because there's exactly one migration to run — pulling in a migration
framework (another dependency, another thing to configure/vendor) for a
single CREATE TABLE is exactly the kind of premature machinery this
project's conventions say to avoid. Revisit golang-migrate once there's
real schema churn across environments (rollback support, checksums,
concurrent-apply safety become worth their cost at that point, not before).
Convention: one DDL statement per migration file. The ClickHouse HTTP
interface isn't reliably multi-statement, so migrate.sh doesn't try to
split multi-statement files — keep each migration to a single statement.
Running
docker compose up -d # starts a standalone ClickHouse for local work
./migrate.sh # applies migrations/*.sql
Environment variables migrate.sh reads (all optional, matching
/ingest's ClickHouse defaults so the two stay in sync out of the box):
| Var | Default |
|---|---|
CLICKHOUSE_HTTP |
http://localhost:8123 |
CLICKHOUSE_USER |
default |
CLICKHOUSE_PASSWORD |
(empty) |
CLICKHOUSE_DATABASE |
sentry |
There's also a Dockerfile (bash + curl baked in, migrations/ copied in
at build time) used by the root-level docker-compose.yml as a one-shot
init service — no runtime package install, no host volume mount needed.
Adding a migration
Add migrations/000N_description.sql with the next sequential number and
a single DDL statement. migrate.sh picks it up automatically — no
registration step.