f5ff19327bb66b54ae494ab2141915fa2d172304
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d05ebe046c |
Make the demo reset survive a slow start, and drop the SQL panels
Two faults, both found by running the reset against the demo rather than by reading it, and both fixed on the box before this commit existed. The reset raced its own alerting container. Writing .env changes alerting's environment, so `docker compose up -d alerting` recreates it -- and the next line posted notification targets to it with no wait. On a busy box that lost: curl returned nothing, json.load threw on an empty string, and set -e killed the script. The damage is in the ordering: `docker compose down -v` runs near the top, so any failure after it leaves the public demo up, empty, and with the simulator still stopped, because the unit is only restarted on the last line. It has been surviving nightly on timing alone. It now polls /healthz for up to 60 seconds and fails loudly before the seed rather than after the wipe. Five dashboard panels were raw ClickHouse SQL, which dashboards refuse: validatePanel rejects query_language "sql" outright, because the time-range picker is injected as leading query terms and a SELECT has nowhere to put them. They were written that way because the pipe syntax has no time bucketing -- no timechart, no bin -- so a genuine time series is not available to a dashboard panel at all. Each is now the breakdown the panel was actually asking for: kubelet events by kind and host, DNS by type and result, queue depth by queue and host, IIS by site and host, HAProxy by backend and balancer. Both mistakes were mine and both were avoidable by reading: the rule about pipe-syntax-only dashboards is stated in terraform/README.md, and I had already read the line that says it. Verified on the demo: 50 hosts, 24 services, 1,922,128 records, all 13 dashboards and all 11 alert rules applied, simulator active. `system` appears on exactly 31 hosts, which is the Linux count -- no Windows host was given a journald stream. |
||
|
|
e10289ee36 |
Grow the demo fleet to fifty hosts and a real Windows tier
Twelve hosts running six services read as somebody's side project. Fifty hosts across twenty-one services read as an estate, which is what a visitor is trying to see themselves in. Thirty-one Linux, eighteen Windows, one Linux host whose agent is gone. The proportions are the point: Windows now carries Active Directory, IIS, SQL Server, Exchange, file shares, Remote Desktop, print, WSUS and SCCM rather than appearing as a Security channel on one box. Linux gains a load-balancer tier, an outbound proxy, MySQL beside Postgres, RabbitMQ, Elasticsearch, three Kubernetes nodes, CI, Vault, OpenLDAP, BIND and a backup server. Fifteen new generators, each writing what the real daemon writes -- HAProxy's timing quintuple, MySQL slow-query blocks, W3C extended format for IIS, kubelet PLEG lines, BIND query logging with NXDOMAIN, Squid's TCP_DENIED, SQL Server deadlock and I/O-stall messages -- with the structured fields carried in attributes so both halves of the query language have something to work on. Three things found while doing it, each of which would have shipped as a quiet wrongness: linuxHosts() decided Windows by `service == "eventlog"`. That held while eventlog was the only Windows role; with IIS and SQL Server on Windows it would have given every one of them a journald system stream -- sshd and UFW lines on a Windows box. It now decides from `os`. worker-02 lost its filling disk when the fleet was rewritten, which is the story worker-disk-filling thresholds on. Restored at its original rate, with a comment saying why it cannot move. Five dashboard panels were written as `timechart`, which this query language does not have -- its stages are where/stats/sort/fields/head/ tail. They are raw ClickHouse SQL now, which the reference recommends for exactly this, and a `pie` panel was dropped before it shipped because the web's VizType union has no such member even though the API accepts one. Five new dashboards: platform/Kubernetes, directory and DNS, messaging and search, the Windows server estate, and edge/proxy. Every panel was checked against the language's real stage list and the web's real viz types. Volume roughly quadruples: about 316 records/minute at rate-scale 1, and 1.9M per nightly reset at the demo's own settings against about 0.5M before. ClickHouse will not notice; reset time and disk on the demo box might, so the README says so and names RATE_SCALE as the lever. |
||
|
|
bcb9a01cd6 |
Give the demo a live synthetic fleet, dashboards, and alert rules
The demo had 75k generic records across eight host-0N/service pairs, one dashboard, one alert rule, and -- because nothing ever called AgentControl.CheckIn -- a completely empty Agents page. /hack/demo-simulator replaces the generic data with a fictional but coherent fleet: 14 hosts running nginx, an API tier, workers, Postgres, Redis, mail, Linux journals and Windows event logs, whose messages and attributes look like what those services actually write. It backfills a week (~370k records, ~20s) and then keeps running. Running continuously is the point, not an implementation detail. Three things the demo has to show are only true if data keeps arriving: the Agents page marks a host stale once check-ins stop, alert rules evaluate over trailing windows and would freeze in one state against a static dataset, and any "last 15 minutes" view is empty on data that stopped growing overnight. It also emits metrics/heartbeats and answers CheckIn faithfully enough that the remote-config editor's pending -> applied transition works end to end. Seeded incidents give the data something to find: an api-02 outage with matching slow queries on db-01, 5xx at the edge and cascading job failures; an SSH probe burst; a spam wave; a disk filling up; and one decommissioned host left deliberately stale. /hack/demo-seed holds the rest of the deployment -- the nightly reset, eight dashboards (64 panels, every viz type but line), eleven alert rules across three notification targets, and the systemd unit. Rule thresholds are calibrated against what the simulator actually produces: the first pass had four rules whose thresholds the traffic could never reach and one that fired during normal operation. No line charts: dashboard panels reject the raw-SQL escape hatch, and the pipe language has no time-bucketing, so a real time axis isn't expressible today. Noted in demo-seed/README.md rather than papered over. |