Siglake / benchmarksGitHub ↗

How we measure

Methodology

Fairness notes (http-logs)

The board carries eight series: Siglake, Quickwit, Elasticsearch, ClickHouse and clickhouse-s3, DuckDB and duckdb-s3, and VictoriaLogs.

Every Siglake row says whether the result cache was off

Result caches are off; data caches are warm. Each shape runs warmups and then 10 measured repetitions; the board shows the median with the first post-settle execution (cold) in parentheses. A result cache memoizes the answer to a query, so leaving one on measures the cache rather than the engine — repeated queries collapse onto a cache-hit floor and stop being comparable to anything. Every engine that has one and lets us disable it therefore runs with it off: Siglake with SIGLAKE_QUERY_RESULT_CACHE=off (its snapshot-keyed result cache, default-on in production) and Elasticsearch with request_cache=false (its shard request cache, which memoizes aggregation results). ClickHouse's query cache and DuckDB have no result cache in play by default. Data caches — OS page cache, block/mark caches, Siglake's per-file footer cache — are left warm everywhere, since every engine has equivalents and disabling them would measure cold I/O instead. Siglake's production result cache is a real feature and makes repeated queries far faster than the board shows; it is disabled here precisely so the board measures execution.

"Result cache off for every measured query" is the methodology claim the whole board rests on, and the 2026-07-27 round measured Siglake's whole-answer memo table without anyone noticing. Since 2026-09-03 the Siglake arm reads SIGLAKE_QUERY_RESULT_CACHE back from the live siglake-query-server container after its pre-suite cold restart (bench/siglake_on_node.sh), writes it to results/<date>-aws/siglake-query-config.json, and run_queries.py --config embeds it as a config block in siglake-core.json. emit_public_record.py then:

lint_records.py enforces the two allowed values whenever the key is present and requires it on every Siglake row dated 2026-09-03 or later. A cache-on row is accepted for deliberately cached datasets, but the lint prints a notice.

A cache-on row measures the cache, so it takes no part in the comparisons that claim to measure execution. The board, the charts and their over-time lines, the cost table and the count and head_ts cross-checks all draw from the rows that do not say result_cache: "on" — one rule, make_charts.eligible, which site/build_site.py and check_counts.py share so they cannot disagree about which round is the latest. A cached diagnostic therefore cannot take a measured column's place by being newer, however it was appended. Two things it does not change: the row stays in records.jsonl and in the published data/records.jsonl, which remain append-only and complete; and a row from a round that recorded no mode (everything before 2026-09-03) is selected exactly as before, because unverified is not the same as known-on. The chart, site and count-check runs print how many rows the rule left out. No comparison renders a cached row; include_cached=True is the opt-in a comparison view would have to ask for, and no renderer asks today.

The Runs page is the inventory of rounds rather than a comparison, so it lists the cached rounds along with the measured ones and marks them cache-on diagnostic beside the engine and version that ran them. A reader who finds a round there can see why its numbers are nowhere on the board. The mark follows the row, not the date: a date whose Siglake arm was cached and whose other engines were not lists both, and the same engine and version appears twice when one round of it recorded the cache on and another recorded it off. A round that recorded no mode is listed with no mark, which says nothing about its cache either way.

bench/httplogs_run.sh copies the sidecar next to the raw results and passes it; a refusal for one arm is reported and the other arms are still emitted, and the driver exits 1 at the end.

Open data: the full append-only results dataset is published at data/records.jsonl; charts and this site regenerate from it deterministically.

Additional engine coverage

EventData’s first published comparison still targets the existing eight configurations. Work on these additions runs separately and does not hold up publication once those eight are qualified.