How we measure
Methodology
Fairness notes (http-logs)
The board carries eight series: Siglake, Quickwit, Elasticsearch, ClickHouse and clickhouse-s3, DuckDB and duckdb-s3, and VictoriaLogs.
- Exact-point candidate rounds are excluded by default. A round launched with
EXACT_POINT_STEP=1runs the primary writer from a feature-built image and writes its table in anexact_point_experiment_*namespace. Its ingest, compaction, layout and query results therefore do not describe the ordinary Siglake build. Keep those rounds out ofrecords.jsonl; they are retained as experimental evidence for the exact-point falsifiers. There is no automatic override that makes such a round a public baseline. - Every arm ingests the same frozen corpus: the 1998 World Cup HTTP server logs via the Elastic Rally
http_logsdistribution — 247,249,096 real-world documents (~33 GB raw) with typed fields (status/size integers), spanning 45 days with heavily skewed traffic (77% of the documents fall in the last quarter of the time span). Every arm reads the same frozen NDJSON shards. - Every arm runs single-node on the same instance class (AWS m6i.4xlarge, 16 vCPU / 64 GiB), queried through its standard front door, measured post-ingest at an idle settled state (merges/compaction quiesced; the Siglake arm drains its freshness markers first so no arm measures with rows in flight).
- Storage is a first-class attribute of the comparison (the board shows it under each engine's column header): Siglake and Quickwit serve from S3 object storage; Elasticsearch, VictoriaLogs and the base ClickHouse/DuckDB series from local EBS (which, at this corpus size, effectively means OS page cache when warm). Two same-storage-class variants close the loop: clickhouse-s3 (native MergeTree-over-S3 disk, with ClickHouse's local filesystem cache deliberately disabled so the series really does serve from S3 — a setting that costs ClickHouse latency and is not its default) and duckdb-s3 (the same plain Parquet read via httpfs from S3 — the exact storage class Siglake's warehouse uses). Elasticsearch has no OSS S3 query path (searchable snapshots are Enterprise-licensed), which is itself part of the comparison.
- Explicit typed mappings on every engine that has a schema: status/size are integer (fast/LowCardinality-adjacent) columns on Siglake, Quickwit, Elasticsearch, ClickHouse and the Parquet baseline; the request line is the full-text field (tokenbf_v1 skip index on ClickHouse). VictoriaLogs is schemaless — there is no index to create and fields are typed on read, so status/size come back as strings and its LogsQL shapes compare them numerically (
status:=404,status:>=500). - VictoriaLogs is a single
victoria-logscontainer on local EBS. Ingest is/insert/jsonline, which takes the corpus NDJSON verbatim (_time_field=timestamp,_msg_field=message). It runs with no_stream_fields: the obvious candidate (host) has over a million distinct values and VictoriaLogs warns against high-cardinality stream fields, so every row lands in one stream — the fair analogue of the other arms' single flat table.-retentionPeriod=100yis required because the corpus is from 1998 and the default retention would drop every row at ingest. Queries are LogsQL (victorialogs_queryper shape inbench/queries-httplogs.json), all answered through/select/logsql/query, which streams JSON lines for both browse and| statsshapes. - Windowed shapes use identical absolute window bounds on every arm (synthesized from the corpus manifest); cross-arm count equality is the correctness check, and approximate cells (Elasticsearch's HLL
count_distinct_host) take no part in it. The verdict for the current board is generated fromrecords.jsonlrather than typed here — by the same selectioncheck_counts.pygates in CI, so this text cannot claim an agreement the checker does not see: On the current board every count cell agrees across all 8 series (11 count-carrying shapes compared) with one known exception:count_distinct_host, where Siglake reports 1,149,520 against 1,149,519 on clickhouse, clickhouse-s3, duckdb, duckdb-s3 and victorialogs. The extra host matches thefreshness-probemarker rows the Siglake arm ingests alongside the corpus (the same rows that put its ingested row count a few dozen above the manifest); it is a harness artifact, recorded here rather than papered over. count_distinct_hostis exact on Siglake (count(DISTINCT host)), ClickHouse (uniqExact), DuckDB (count(DISTINCT host)) and VictoriaLogs (count_uniq); approximate on Elasticsearch (cardinality HLL, so its cell takes no part in the count cross-check); refused by this Quickwit build, which has no cardinality aggregation. Latencies live on the board, not in this text.- Refusals are results: an arm that errors on a shape is recorded as unsupported, not skipped. ES and QW both cap deep pagination (10K window), and QW refuses the cardinality aggregation; those render as "unsupported" cells.
- DuckDB is the vanilla-parquet baseline: the NDJSON shards converted to PLAIN Parquet (default writer settings, arrival order, no sorting, no indexes, no serving process) queried in-process by stock DuckDB via read_parquet. Text search is LIKE substring — vanilla Parquet has no token index. It calibrates what "just files + an OSS query engine" buys before any log-analytics system is involved.
Every Siglake row says whether the result cache was off
Result caches are off; data caches are warm. Each shape runs warmups and then 10 measured repetitions; the board shows the median with the first post-settle execution (cold) in parentheses. A result cache memoizes the answer to a query, so leaving one on measures the cache rather than the engine — repeated queries collapse onto a cache-hit floor and stop being comparable to anything. Every engine that has one and lets us disable it therefore runs with it off: Siglake with SIGLAKE_QUERY_RESULT_CACHE=off (its snapshot-keyed result cache, default-on in production) and Elasticsearch with request_cache=false (its shard request cache, which memoizes aggregation results). ClickHouse's query cache and DuckDB have no result cache in play by default. Data caches — OS page cache, block/mark caches, Siglake's per-file footer cache — are left warm everywhere, since every engine has equivalents and disabling them would measure cold I/O instead. Siglake's production result cache is a real feature and makes repeated queries far faster than the board shows; it is disabled here precisely so the board measures execution.
"Result cache off for every measured query" is the methodology claim the whole board rests on, and the 2026-07-27 round measured Siglake's whole-answer memo table without anyone noticing. Since 2026-09-03 the Siglake arm reads SIGLAKE_QUERY_RESULT_CACHE back from the live siglake-query-server container after its pre-suite cold restart (bench/siglake_on_node.sh), writes it to results/<date>-aws/siglake-query-config.json, and run_queries.py --config embeds it as a config block in siglake-core.json. emit_public_record.py then:
- refuses (exit 1, nothing appended) a results file whose
config(or the--configsidecar, when the file has none) saysresult_cacheis anything butoff— an unset value counts as on, because that is Siglake's default; - appends it anyway under
--allow-cached, for a deliberately cached dataset; - stamps
result_cache("off"/"on") on every row it appends when the mode is known. Rows from rounds that did not record the mode (everything before 2026-09-03) carry noresult_cachekey, and the emitter warns that the mode is unverified rather than guessing.
lint_records.py enforces the two allowed values whenever the key is present and requires it on every Siglake row dated 2026-09-03 or later. A cache-on row is accepted for deliberately cached datasets, but the lint prints a notice.
A cache-on row measures the cache, so it takes no part in the comparisons that claim to measure execution. The board, the charts and their over-time lines, the cost table and the count and head_ts cross-checks all draw from the rows that do not say result_cache: "on" — one rule, make_charts.eligible, which site/build_site.py and check_counts.py share so they cannot disagree about which round is the latest. A cached diagnostic therefore cannot take a measured column's place by being newer, however it was appended. Two things it does not change: the row stays in records.jsonl and in the published data/records.jsonl, which remain append-only and complete; and a row from a round that recorded no mode (everything before 2026-09-03) is selected exactly as before, because unverified is not the same as known-on. The chart, site and count-check runs print how many rows the rule left out. No comparison renders a cached row; include_cached=True is the opt-in a comparison view would have to ask for, and no renderer asks today.
The Runs page is the inventory of rounds rather than a comparison, so it lists the cached rounds along with the measured ones and marks them cache-on diagnostic beside the engine and version that ran them. A reader who finds a round there can see why its numbers are nowhere on the board. The mark follows the row, not the date: a date whose Siglake arm was cached and whose other engines were not lists both, and the same engine and version appears twice when one round of it recorded the cache on and another recorded it off. A round that recorded no mode is listed with no mark, which says nothing about its cache either way.
bench/httplogs_run.sh copies the sidecar next to the raw results and passes it; a refusal for one arm is reported and the other arms are still emitted, and the driver exits 1 at the end.
Open data: the full append-only results dataset is published at data/records.jsonl; charts and this site regenerate from it deterministically.
Additional engine coverage
- Parseable — adapter and correctness qualification in progress; no measurements yet.
- DuckDB + DuckLake — adapter and correctness qualification in progress; no measurements yet.
EventData’s first published comparison still targets the existing eight configurations. Work on these additions runs separately and does not hold up publication once those eight are qualified.