Phase 1: Windows log collection + full-text search

Extends the agent, ingest, storage, api, and web with Windows Event
Log/ETW sourcing and Tantivy-backed free-text search, per the approved
Phase 1 plan.

- CLAUDE.md: materialized on disk (never existed as a file before) with
  a new Phase 1 "done looks like" section.
- agent: Windows Event Log (EvtSubscribe) and ETW sources, Windows
  service wrapper (install/uninstall/run-service), both feature- and
  target_os-gated so Linux builds/tests/clippy stay unaffected. Also
  fixed two pre-existing Phase 0 clippy gaps (dead-code on
  default-features-only builds, a type-inference edge case) found while
  testing every feature combination properly for the first time.
  UNVERIFIED on real Windows -- no Windows toolchain existed anywhere in
  the build environment; flagged prominently in three places.
- proto/ingest: new record_id field, assigned once server-side in
  ingest's gRPC front end so ClickHouse and Tantivy agree on the same ID
  for the same record.
- storage: record_id column + bloom filter index, verified against a
  live ClickHouse.
- search: new service, Tantivy index, rskafka consumer as an independent
  second consumer group on the same Redpanda topic ingest already reads.
- api/web: new /search endpoint and page, sharing the query page's
  result-table shape and component.
- hack/windows-fixture: sends realistic Windows-shaped data straight to
  ingest, so the pipeline's handling of it is verifiable without a
  Windows host.

Verified end-to-end on the live docker-compose stack: the same record_id
comes back from both /query and /search for the same log line, including
for windows-fixture's synthetic Windows Event Log data. Real bugs found
and fixed along the way: api/Dockerfile missing proto/ in its build
context, search's logs being completely silent (RUST_LOG gap), and
search/target/ missing from .gitignore/.dockerignore.
This commit is contained in:
2026-08-13 11:27:35 -07:00
parent fe854b1091
commit cd8aa290ca
66 changed files with 6084 additions and 171 deletions
+21 -2
View File
@@ -4,7 +4,8 @@ ClickHouse schema and migration tooling for Sentry's analytical store.
## Schema
One table for Phase 0, `logs`:
One table, `logs` (Phase 0 columns plus `record_id`, added in
`migrations/0002_add_record_id.sql`):
```sql
CREATE TABLE logs
@@ -14,13 +15,31 @@ CREATE TABLE logs
`service` String,
`severity` LowCardinality(String),
`message` String,
`attributes` Map(String, String)
`attributes` Map(String, String),
`record_id` UUID DEFAULT generateUUIDv4()
)
ENGINE = MergeTree
PARTITION BY toDate(timestamp)
ORDER BY (service, timestamp)
-- plus: INDEX record_id_idx record_id TYPE bloom_filter GRANULARITY 4
```
**`record_id`** (Phase 1) is the stable per-record identifier Tantivy's
full-text search joins back to this table with — `/ingest`'s gRPC front
end assigns it once, server-side, before a record is produced to
Redpanda (see `/ingest/README.md` for why it has to happen exactly once,
upstream of both the ClickHouse-writer and Tantivy-indexer consumers).
Added via `ALTER TABLE ... ADD COLUMN` + `ADD INDEX` rather than changing
`ORDER BY`: `ORDER BY (service, timestamp)` is the proven time-range-scan
access pattern from Phase 0 and shouldn't be disturbed for a
fundamentally different access pattern (point lookups by ID). A
data-skipping bloom filter index on `record_id` serves the `WHERE
record_id IN (...)` lookup Tantivy-backed search results need, without
touching the primary sort order. The `DEFAULT generateUUIDv4()` is a
safety net, not the normal path — every row `/ingest` writes explicitly
supplies its own `record_id` from the proto message; the default only
matters for rows written some other way.
Notes on choices that weren't fully specified by the task description:
- **`DateTime64(9, 'UTC')`** (nanosecond precision) rather than second or
@@ -0,0 +1,3 @@
ALTER TABLE logs
ADD COLUMN record_id UUID DEFAULT generateUUIDv4(),
ADD INDEX record_id_idx record_id TYPE bloom_filter GRANULARITY 4