Phase 4: real per-tenant ClickHouse isolation via a new enterprise-api binary

Closes the threat model's headline finding for the SQL query path:
enterprise/internal/tenantprovision does real CREATE DATABASE/USER/GRANT
against ClickHouse, and enterprise/internal/chrunner is a per-tenant
connection registry implementing api's SQLRunner interface, resolving
the tenant from the authenticated request identity -- never a
caller-suppliable parameter. Both are wired into a new binary,
enterprise/cmd/enterprise-api, alongside the unchanged single-tenant
api/cmd/api, since AGPL core can never import enterprise/ and Go's own
internal/ package visibility rules meant enterprise/ couldn't implement
core's SQLRunner interface without importing the package that defines
it. That required moving api/internal/{authz,queryapi,dashboards,
querylang/executor,searchclient,httpserver} out of internal/ -- the
minimal set enterprise-api needs to import; querylang's compiler
internals (planner/lexer/parser/ast/ir) and api's own config stay
internal, since nothing outside api needs them directly.

Also finally wires enterprise/internal/audit into queryapi.AuditLogger
(nil since Phase 4 task 4) via a new adapter, and adds live-ClickHouse
integration tests for two of the four adversarial probes named in
docs/phase-4-isolation-design.md's verification plan.

Corrected several overclaims in the docs while writing this up: an
earlier claim that rbacstore's CRUD was "verified against a live
Postgres" was never actually true in this environment (only
internal/audit was, earlier in this phase, before Docker access was
lost) -- threat-model.md, phase-4-runbook.md, CLAUDE.md, and
enterprise/README.md all now distinguish "a real integration test
exists" from "this was confirmed against a live database."

Still not built: Tantivy/free-text tenant isolation
(enterprise/internal/searchclient), and any deployment-topology
mechanism that actually routes traffic to enterprise-api instead of
plain api -- both binaries exist side by side today with nothing
enforcing or flagging which one a deployment runs.
This commit is contained in:
2026-08-13 22:48:38 -07:00
parent 3eb0f4c589
commit 1d57e697b1
49 changed files with 2003 additions and 237 deletions
+58
View File
@@ -0,0 +1,58 @@
package executor
import (
"context"
"fmt"
"reflect"
"github.com/ClickHouse/clickhouse-go/v2/lib/driver"
)
// ChRunner runs arbitrary (pre-validated) SELECT statements against
// ClickHouse and shapes the result into JSON-friendly columns/rows,
// discovering the result's column set at query time via reflection since
// the query itself is arbitrary. Ported from Phase 0/1's
// api/queryapi.Executor, which this replaces (see task 4) --
// same logic, moved here since it's the query-execution layer's
// plumbing, not specific to the old placeholder /query handler.
type ChRunner struct {
conn driver.Conn
}
func NewChRunner(conn driver.Conn) *ChRunner {
return &ChRunner{conn: conn}
}
func (r *ChRunner) RunSQL(ctx context.Context, sql string) (*Result, error) {
rows, err := r.conn.Query(ctx, sql)
if err != nil {
return nil, fmt.Errorf("executing query: %w", err)
}
defer rows.Close()
columnTypes := rows.ColumnTypes()
result := &Result{
Columns: rows.Columns(),
Rows: [][]any{},
}
for rows.Next() {
dest := make([]any, len(columnTypes))
for i, ct := range columnTypes {
dest[i] = reflect.New(ct.ScanType()).Interface()
}
if err := rows.Scan(dest...); err != nil {
return nil, fmt.Errorf("scanning row: %w", err)
}
row := make([]any, len(dest))
for i, d := range dest {
row[i] = reflect.ValueOf(d).Elem().Interface()
}
result.Rows = append(result.Rows, row)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("iterating rows: %w", err)
}
return result, nil
}
+72
View File
@@ -0,0 +1,72 @@
// Package executor runs a compiled ir.Plan and returns results in a
// shape consistent regardless of which backend(s) were hit -- the point
// of compiling to one IR in the first place. See
// /docs/query-language-design.md's "Execution" section for the four
// routing cases implemented here.
package executor
import (
"context"
"fmt"
"github.com/sentry/sentry/api/internal/querylang/ir"
)
type Result struct {
Columns []string
Rows [][]any
}
// SQLRunner executes a raw SQL statement against ClickHouse. *ChRunner
// (chrunner.go) is the production implementation; tests use a fake --
// same narrow-interface pattern used throughout /ingest and /api.
type SQLRunner interface {
RunSQL(ctx context.Context, sql string) (*Result, error)
}
// SearchClient resolves a Tantivy query into matching record_ids.
type SearchClient interface {
Search(ctx context.Context, query string, limit uint32) ([]string, error)
}
// textSearchLimit caps how many record_ids a Tantivy prefilter can feed
// into a ClickHouse `IN (...)` clause. See /docs/query-language-design.md's
// "Known scaling limitation" -- this is a real, disclosed limit on result
// completeness for very broad text searches, not an oversight.
//
// 5000, not 10000: confirmed by actually running the Phase 2 benchmark
// (see /docs/phase-2-runbook.md) that 10000 quoted UUIDs (~39 bytes each
// including the comma) produces a ~390KB query string, which exceeds
// ClickHouse's default max_query_size (262144 bytes / 256KiB) and fails
// outright with a syntax error rather than degrading gracefully. 5000
// UUIDs is ~195KB, safely under that default with headroom for the rest
// of the query. This was a real failure caught by running the benchmark,
// not a value chosen from first-principles estimation.
const textSearchLimit = 5000
// Execute runs plan against the given backends. The four cases (per the
// design doc): RawSQL passthrough; pure ClickHouse (no TextSearch); text
// search alone (Tantivy prefilter -> ClickHouse row fetch); text search
// plus aggregation (Tantivy prefilter -> ClickHouse aggregate). Cases 2-4
// share the same buildSQL/buildWhereClause code (sql.go) -- the only
// difference is whether a record_id filter is threaded in.
func Execute(ctx context.Context, plan *ir.Plan, sqlRunner SQLRunner, search SearchClient) (*Result, error) {
if plan.RawSQL != "" {
return sqlRunner.RunSQL(ctx, plan.RawSQL)
}
var recordIDFilter []string
if len(plan.TextSearch) > 0 {
ids, err := search.Search(ctx, plan.TextSearch[0].Query, textSearchLimit)
if err != nil {
return nil, fmt.Errorf("full-text search failed: %w", err)
}
if len(ids) == 0 {
return &Result{Columns: []string{}, Rows: [][]any{}}, nil
}
recordIDFilter = ids
}
sql := buildSQL(plan, recordIDFilter)
return sqlRunner.RunSQL(ctx, sql)
}
+333
View File
@@ -0,0 +1,333 @@
package executor
import (
"context"
"errors"
"strings"
"testing"
"time"
"github.com/sentry/sentry/api/internal/querylang/ir"
)
func mustParseTime(t *testing.T, s string) time.Time {
t.Helper()
tm, err := time.Parse(time.RFC3339, s)
if err != nil {
t.Fatalf("parsing time %q: %v", s, err)
}
return tm
}
type fakeSQLRunner struct {
gotSQL string
result *Result
err error
calls int
}
func (f *fakeSQLRunner) RunSQL(_ context.Context, sql string) (*Result, error) {
f.gotSQL = sql
f.calls++
if f.err != nil {
return nil, f.err
}
if f.result != nil {
return f.result, nil
}
return &Result{Columns: []string{}, Rows: [][]any{}}, nil
}
type fakeSearchClient struct {
gotQuery string
gotLimit uint32
ids []string
err error
calls int
}
func (f *fakeSearchClient) Search(_ context.Context, query string, limit uint32) ([]string, error) {
f.gotQuery = query
f.gotLimit = limit
f.calls++
if f.err != nil {
return nil, f.err
}
return f.ids, nil
}
func TestExecuteRawSQLBypassesEverythingElse(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{}
plan := &ir.Plan{RawSQL: "SELECT 1"}
_, err := Execute(context.Background(), plan, sqlRunner, search)
if err != nil {
t.Fatalf("Execute() error = %v", err)
}
if sqlRunner.gotSQL != "SELECT 1" {
t.Fatalf("gotSQL = %q, want %q", sqlRunner.gotSQL, "SELECT 1")
}
if search.calls != 0 {
t.Fatalf("expected search not to be called for RawSQL, got %d calls", search.calls)
}
}
func TestExecutePureClickHousePathSkipsSearch(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{}
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "service", Op: "=", Value: "api"}}}
_, err := Execute(context.Background(), plan, sqlRunner, search)
if err != nil {
t.Fatalf("Execute() error = %v", err)
}
if search.calls != 0 {
t.Fatalf("expected no search calls, got %d", search.calls)
}
if !strings.Contains(sqlRunner.gotSQL, "FROM logs") || !strings.Contains(sqlRunner.gotSQL, "`service` = 'api'") {
t.Fatalf("unexpected SQL: %s", sqlRunner.gotSQL)
}
}
func TestExecuteTextSearchPrefiltersThenQueriesClickHouse(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{ids: []string{"id-1", "id-2"}}
plan := &ir.Plan{TextSearch: []ir.TextPredicate{{Query: "connection refused"}}}
_, err := Execute(context.Background(), plan, sqlRunner, search)
if err != nil {
t.Fatalf("Execute() error = %v", err)
}
if search.gotQuery != "connection refused" {
t.Fatalf("search query = %q", search.gotQuery)
}
if search.gotLimit != textSearchLimit {
t.Fatalf("search limit = %d, want %d", search.gotLimit, textSearchLimit)
}
if !strings.Contains(sqlRunner.gotSQL, "record_id IN ('id-1','id-2')") {
t.Fatalf("unexpected SQL: %s", sqlRunner.gotSQL)
}
}
func TestExecuteTextSearchNoMatchesSkipsClickHouseEntirely(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{ids: nil}
plan := &ir.Plan{TextSearch: []ir.TextPredicate{{Query: "nothing matches"}}}
result, err := Execute(context.Background(), plan, sqlRunner, search)
if err != nil {
t.Fatalf("Execute() error = %v", err)
}
if sqlRunner.calls != 0 {
t.Fatalf("expected ClickHouse not to be queried when search finds nothing, got %d calls", sqlRunner.calls)
}
if len(result.Columns) != 0 || len(result.Rows) != 0 {
t.Fatalf("expected empty result, got %+v", result)
}
}
func TestExecuteTextSearchWithAggregation(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{ids: []string{"id-1"}}
plan := &ir.Plan{
TextSearch: []ir.TextPredicate{{Query: "connection refused"}},
Aggregation: &ir.Aggregation{
Funcs: []ir.AggFunc{{Func: "count", Alias: "count"}},
GroupBy: []string{"host"},
},
}
_, err := Execute(context.Background(), plan, sqlRunner, search)
if err != nil {
t.Fatalf("Execute() error = %v", err)
}
if !strings.Contains(sqlRunner.gotSQL, "record_id IN ('id-1')") {
t.Fatalf("expected the text-search prefilter in the WHERE clause: %s", sqlRunner.gotSQL)
}
if !strings.Contains(sqlRunner.gotSQL, "GROUP BY `host`") {
t.Fatalf("expected GROUP BY: %s", sqlRunner.gotSQL)
}
if !strings.Contains(sqlRunner.gotSQL, "count() AS `count`") {
t.Fatalf("expected count() AS `count`: %s", sqlRunner.gotSQL)
}
}
func TestExecuteSearchErrorPropagates(t *testing.T) {
sqlRunner := &fakeSQLRunner{}
search := &fakeSearchClient{err: errors.New("search unavailable")}
plan := &ir.Plan{TextSearch: []ir.TextPredicate{{Query: "x"}}}
_, err := Execute(context.Background(), plan, sqlRunner, search)
if err == nil {
t.Fatal("expected the search error to propagate")
}
if sqlRunner.calls != 0 {
t.Fatalf("expected ClickHouse not to be queried after a search error, got %d calls", sqlRunner.calls)
}
}
func TestBuildSQLNumericCastOnAttributesField(t *testing.T) {
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "status", Op: ">=", Value: "500"}}}
sql := buildSQL(plan, nil)
want := "toFloat64OrZero(attributes['status']) >= 500"
if !strings.Contains(sql, want) {
t.Fatalf("SQL = %q, want it to contain %q", sql, want)
}
}
func TestBuildSQLStringComparisonOnAttributesField(t *testing.T) {
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "status", Op: "=", Value: "unknown"}}}
sql := buildSQL(plan, nil)
want := "attributes['status'] = 'unknown'"
if !strings.Contains(sql, want) {
t.Fatalf("SQL = %q, want it to contain %q", sql, want)
}
}
func TestBuildSQLTopLevelFieldNeverCast(t *testing.T) {
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "service", Op: "=", Value: "123"}}}
sql := buildSQL(plan, nil)
if strings.Contains(sql, "toFloat64OrZero") {
t.Fatalf("top-level field should never be numeric-cast: %s", sql)
}
if !strings.Contains(sql, "`service` = '123'") {
t.Fatalf("unexpected SQL: %s", sql)
}
}
func TestBuildSQLEscapesInjectionAttemptInValue(t *testing.T) {
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "service", Op: "=", Value: "x'; DROP TABLE logs; --"}}}
sql := buildSQL(plan, nil)
// The whole attacker-controlled value must land inside exactly one
// quoted literal, with its embedded quote backslash-escaped so it
// can't terminate the literal early -- checking for the escaped
// form directly, not just the absence of the raw substring (which
// is a weaker check: "\\'; DROP TABLE" still *contains* "'; DROP
// TABLE" as a substring, so that alone doesn't prove escaping
// happened).
want := `'x\'; DROP TABLE logs; --'`
if !strings.Contains(sql, want) {
t.Fatalf("expected the literal %q in SQL, got: %s", want, sql)
}
}
func TestBuildSQLDefaultLimitAppliedWhenNoneGiven(t *testing.T) {
plan := &ir.Plan{Filters: []ir.FilterPredicate{{Field: "service", Op: "=", Value: "api"}}}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "LIMIT 100") {
t.Fatalf("expected the default row limit, got: %s", sql)
}
}
func TestBuildSQLExplicitLimitOverridesDefault(t *testing.T) {
plan := &ir.Plan{
Filters: []ir.FilterPredicate{{Field: "service", Op: "=", Value: "api"}},
Limit: &ir.Limit{N: 5},
}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "LIMIT 5") || strings.Contains(sql, "LIMIT 100") {
t.Fatalf("expected LIMIT 5, got: %s", sql)
}
}
func TestBuildSQLTailWithoutSortOrdersAscending(t *testing.T) {
plan := &ir.Plan{Limit: &ir.Limit{N: 10, Tail: true}}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "ORDER BY `timestamp` ASC") {
t.Fatalf("expected ascending order for tail, got: %s", sql)
}
}
func TestBuildSQLNoSortDefaultsNewestFirst(t *testing.T) {
plan := &ir.Plan{}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "ORDER BY `timestamp` DESC") {
t.Fatalf("expected newest-first default, got: %s", sql)
}
}
func TestBuildSQLSortByAggregateAlias(t *testing.T) {
plan := &ir.Plan{
Aggregation: &ir.Aggregation{
Funcs: []ir.AggFunc{{Func: "count", Alias: "count"}},
GroupBy: []string{"host"},
},
Sort: []ir.SortField{{Field: "count", Desc: true}},
}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "ORDER BY `count` DESC") {
t.Fatalf("expected ORDER BY on the aggregate alias, got: %s", sql)
}
}
func TestBuildSQLSortByGroupByField(t *testing.T) {
plan := &ir.Plan{
Aggregation: &ir.Aggregation{
Funcs: []ir.AggFunc{{Func: "count", Alias: "count"}},
GroupBy: []string{"host"},
},
Sort: []ir.SortField{{Field: "host", Desc: false}},
}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "ORDER BY `host` ASC") {
t.Fatalf("expected ORDER BY on the group-by column, got: %s", sql)
}
}
func TestBuildSQLAggregationOnAttributesFieldAlwaysCasts(t *testing.T) {
plan := &ir.Plan{
Aggregation: &ir.Aggregation{
Funcs: []ir.AggFunc{{Func: "avg", Field: "latency_ms", Alias: "avg_latency"}},
},
}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "AVG(toFloat64OrZero(attributes['latency_ms'])) AS `avg_latency`") {
t.Fatalf("unexpected SQL: %s", sql)
}
}
func TestBuildSQLProjectionFields(t *testing.T) {
plan := &ir.Plan{Fields: []string{"host", "message"}}
sql := buildSQL(plan, nil)
if !strings.Contains(sql, "SELECT `host` AS `host`, `message` AS `message` FROM logs") {
t.Fatalf("unexpected SQL: %s", sql)
}
}
func TestBuildSQLTimeRange(t *testing.T) {
from := mustParseTime(t, "2026-08-14T00:00:00Z")
to := mustParseTime(t, "2026-08-14T01:00:00Z")
plan := &ir.Plan{TimeRange: &ir.TimeRange{From: from, To: to}}
sql := buildSQL(plan, nil)
// Space-separated, no 'T'/'Z' -- ClickHouse's implicit string->DateTime64
// cast for a column-vs-literal comparison is strict and rejects
// RFC3339/ISO-8601 shaped literals ("code: 53, Cannot convert string...
// to type DateTime64(9, 'UTC')"), confirmed by actually running a
// dashboard panel with earliest= against live ClickHouse -- this test
// previously asserted the RFC3339 shape that ClickHouse rejects, which
// is exactly how the bug went unnoticed: nothing here ever executed the
// SQL against a real database.
if !strings.Contains(sql, "`timestamp` >= '2026-08-14 00:00:00'") {
t.Fatalf("missing From bound: %s", sql)
}
if !strings.Contains(sql, "`timestamp` <= '2026-08-14 01:00:00'") {
t.Fatalf("missing To bound: %s", sql)
}
}
func TestFormatClickHouseDateTime64OmitsTrailingZeroFraction(t *testing.T) {
// time.Time's default zero-value fractional seconds must not leave a
// stray "." with nothing after it -- Format's `.999999999` verb
// already handles this (trims to nothing when the fraction is zero),
// but it's worth pinning down given how easy the RFC3339Nano mistake
// was to miss in the first place.
got := formatClickHouseDateTime64(mustParseTime(t, "2026-08-14T00:00:00Z"))
if got != "2026-08-14 00:00:00" {
t.Fatalf("got %q, want no trailing fractional-seconds dot", got)
}
got = formatClickHouseDateTime64(mustParseTime(t, "2026-08-14T00:00:00.223505479Z"))
if got != "2026-08-14 00:00:00.223505479" {
t.Fatalf("got %q", got)
}
}
+247
View File
@@ -0,0 +1,247 @@
package executor
import (
"fmt"
"regexp"
"strings"
"time"
"github.com/sentry/sentry/api/internal/querylang/ir"
)
// defaultRowLimit is the safety net when a raw-row query has neither an
// explicit head/tail nor an aggregation -- without it, a bare `service=api`
// with no other pipe stages would return every matching row unbounded.
// Independent of planner's own defaultLimit (same value, different
// concern: that one fills in `head`/`tail` with no N given; this one
// guards queries that never mention head/tail at all).
const defaultRowLimit = 100
// logs' real columns, per /storage. Anything else maps to
// attributes['field'] -- see /docs/query-language-design.md's "Field
// mapping" section.
var topLevelFields = map[string]bool{
"timestamp": true,
"host": true,
"service": true,
"severity": true,
"message": true,
"record_id": true,
}
func buildSQL(plan *ir.Plan, recordIDFilter []string) string {
var sb strings.Builder
sb.WriteString("SELECT ")
sb.WriteString(selectClause(plan))
sb.WriteString(" FROM logs")
if where := buildWhereClause(plan, recordIDFilter); where != "" {
sb.WriteString(" WHERE ")
sb.WriteString(where)
}
if plan.Aggregation != nil && len(plan.Aggregation.GroupBy) > 0 {
sb.WriteString(" GROUP BY ")
cols := make([]string, len(plan.Aggregation.GroupBy))
for i, g := range plan.Aggregation.GroupBy {
cols[i] = columnExpr(g)
}
sb.WriteString(strings.Join(cols, ", "))
}
writeOrderBy(&sb, plan)
if plan.Limit != nil {
fmt.Fprintf(&sb, " LIMIT %d", plan.Limit.N)
} else if plan.Aggregation == nil {
fmt.Fprintf(&sb, " LIMIT %d", defaultRowLimit)
}
return sb.String()
}
func writeOrderBy(sb *strings.Builder, plan *ir.Plan) {
switch {
case len(plan.Sort) > 0:
sb.WriteString(" ORDER BY ")
parts := make([]string, len(plan.Sort))
for i, s := range plan.Sort {
dir := "ASC"
if s.Desc {
dir = "DESC"
}
parts[i] = sortColumnExpr(plan, s.Field) + " " + dir
}
sb.WriteString(strings.Join(parts, ", "))
case plan.Limit != nil && plan.Limit.Tail:
// `tail N` with no explicit sort: order ascending so LIMIT N
// takes the chronologically *last* N rows. Callers wanting
// strict newest-first display order re-sort client-side --
// documented in the query language reference.
sb.WriteString(" ORDER BY `timestamp` ASC")
case plan.Aggregation == nil:
// Raw-row queries with no explicit sort default to newest-first,
// matching the Phase 0/1 UI default.
sb.WriteString(" ORDER BY `timestamp` DESC")
}
}
// sortColumnExpr resolves a sort field against an aggregation's own
// output columns (alias or group-by field) before falling back to the
// normal top-level/attributes mapping -- `sort -count` after `stats
// count` refers to the aggregate's alias, not a raw column.
func sortColumnExpr(plan *ir.Plan, field string) string {
if plan.Aggregation != nil {
for _, f := range plan.Aggregation.Funcs {
if f.Alias == field {
return quoteIdent(field)
}
}
for _, g := range plan.Aggregation.GroupBy {
if g == field {
return columnExpr(field)
}
}
}
return columnExpr(field)
}
func selectClause(plan *ir.Plan) string {
if plan.Aggregation != nil {
parts := make([]string, 0, len(plan.Aggregation.GroupBy)+len(plan.Aggregation.Funcs))
for _, g := range plan.Aggregation.GroupBy {
parts = append(parts, columnExpr(g)+" AS "+quoteIdent(g))
}
for _, f := range plan.Aggregation.Funcs {
parts = append(parts, aggExpr(f)+" AS "+quoteIdent(f.Alias))
}
return strings.Join(parts, ", ")
}
if len(plan.Fields) > 0 {
parts := make([]string, len(plan.Fields))
for i, f := range plan.Fields {
parts[i] = columnExpr(f) + " AS " + quoteIdent(f)
}
return strings.Join(parts, ", ")
}
return "*"
}
// aggExpr always numeric-casts non-top-level (attributes-map) fields for
// sum/avg/min/max, unlike comparison predicates where casting is
// conditional on whether the compared value looks numeric -- an
// aggregate function is inherently a numeric (or, for min/max,
// order-comparable) operation, so there's no "maybe string" case the way
// there is for `field=value`. Known Phase 2 limitation: min/max on a
// non-top-level field always compares numerically, not lexicographically
// -- string min/max on attributes isn't supported this phase.
func aggExpr(f ir.AggFunc) string {
if f.Func == "count" {
return "count()"
}
col := columnExpr(f.Field)
if !topLevelFields[f.Field] {
col = "toFloat64OrZero(" + col + ")"
}
return strings.ToUpper(f.Func) + "(" + col + ")"
}
func buildWhereClause(plan *ir.Plan, recordIDFilter []string) string {
var conds []string
if len(recordIDFilter) > 0 {
quoted := make([]string, len(recordIDFilter))
for i, id := range recordIDFilter {
quoted[i] = quoteLiteral(id)
}
conds = append(conds, "record_id IN ("+strings.Join(quoted, ",")+")")
}
for _, f := range plan.Filters {
conds = append(conds, buildComparisonSQL(f))
}
if plan.TimeRange != nil {
if !plan.TimeRange.From.IsZero() {
conds = append(conds, "`timestamp` >= "+quoteLiteral(formatClickHouseDateTime64(plan.TimeRange.From)))
}
if !plan.TimeRange.To.IsZero() {
conds = append(conds, "`timestamp` <= "+quoteLiteral(formatClickHouseDateTime64(plan.TimeRange.To)))
}
}
return strings.Join(conds, " AND ")
}
// formatClickHouseDateTime64 formats t the way ClickHouse's implicit
// string->DateTime64 CAST expects for a WHERE-clause comparison:
// "YYYY-MM-DD HH:MM:SS[.fractional]", space-separated, no 'T'/'Z'. This
// is a real, measured requirement, not a guess: an ISO-8601/RFC3339Nano
// literal (e.g. "2026-08-12T20:17:40.223505479Z", what time.RFC3339Nano
// produces) fails at query time with "code: 53, Cannot convert string
// ... to type DateTime64(9, 'UTC')" -- ClickHouse's *implicit* cast used
// for column-vs-literal comparisons is strict, unlike the lenient
// parseDateTimeBestEffort used elsewhere in ClickHouse. Found by
// actually running a dashboard panel with a relative earliest= against
// live ClickHouse (Phase 2's own unit tests never caught this: they
// assert against a fake SQLRunner that checks the generated SQL string,
// not that ClickHouse accepts it, and none of Phase 2's own live-stack
// runbook queries happened to use earliest=/latest= at all).
func formatClickHouseDateTime64(t time.Time) string {
return t.UTC().Format("2006-01-02 15:04:05.999999999")
}
// buildComparisonSQL numeric-casts a non-top-level field only when the
// compared value itself looks numeric -- `status>=500` casts (numeric
// comparison intent), `status="unknown"` doesn't (string comparison
// intent). Top-level fields are never cast; ClickHouse compares them
// against a string literal natively (LowCardinality(String)/String
// compare as-is; DateTime64 columns need formatClickHouseDateTime64's
// exact literal shape, handled in buildWhereClause above, not here).
func buildComparisonSQL(f ir.FilterPredicate) string {
if !topLevelFields[f.Field] && isNumericLiteral(f.Value) {
return "toFloat64OrZero(" + columnExpr(f.Field) + ") " + f.Op + " " + f.Value
}
return columnExpr(f.Field) + " " + f.Op + " " + quoteLiteral(f.Value)
}
func columnExpr(field string) string {
if topLevelFields[field] {
return quoteIdent(field)
}
return "attributes[" + quoteLiteral(field) + "]"
}
func quoteIdent(name string) string {
return "`" + strings.ReplaceAll(name, "`", "``") + "`"
}
// quoteLiteral is the actual injection defense for every user-controlled
// string embedded in generated SQL (filter values, attribute keys, time
// bounds, record_ids). Field/keyword tokens from the lexer are already
// constrained to [a-zA-Z0-9_.] by construction (see lexer.isIdentPart)
// and can't carry SQL metacharacters at all, but quoted-string *values*
// can contain anything, so this can't be skipped for them.
func quoteLiteral(s string) string {
var sb strings.Builder
sb.WriteByte('\'')
for _, r := range s {
switch r {
case '\\':
sb.WriteString(`\\`)
case '\'':
sb.WriteString(`\'`)
default:
sb.WriteRune(r)
}
}
sb.WriteByte('\'')
return sb.String()
}
var numericLiteralRe = regexp.MustCompile(`^-?\d+(\.\d+)?$`)
func isNumericLiteral(s string) bool {
return numericLiteralRe.MatchString(s)
}