Benchmarking¶
SignalDB ships a set of Criterion
micro-benchmarks covering the hot paths that performance work keeps
touching: OTLP decode, WAL encoding, the Iceberg append, schema
materialization, compaction, and the querier read paths (both raw scan cost
and the full querier service). This page is how to
run them, how to compare a branch against main, and where the nightly trend
lives.
What the numbers mean. Every bench runs against in-memory catalogs and object stores on a single machine. They measure CPU, planning, encoding, and commit-protocol cost — not S3 latency, not network, not a loaded cluster. Treat them as a relative regression signal ("did this change make it worse?"), never as sizing guidance or a production latency claim.
Latest nightly results¶
The nightly workflow appends every run to the
benchmark-data
branch and this table is regenerated from it when the docs site is built.
Full per-benchmark trend charts: cedricziel.github.io/signaldb/benchmarks.
Latest run: 2026-09-27 04:18 UTC on commit a705ffd1 · trend charts · 43 runs recorded
| Benchmark | Mean | vs. previous run |
|---|---|---|
acceptor_ingest/otlp_decode_and_convert |
962.96 µs | -44.4% |
acceptor_ingest/otlp_convert_only |
562.65 µs | -46.3% |
wal/record_batch_roundtrip |
375.50 µs | -49.2% |
acceptor_ingest_logs/otlp_decode_and_convert |
723.01 µs | -47.0% |
acceptor_ingest_logs/otlp_convert_only |
343.24 µs | -49.0% |
acceptor_ingest_metrics/otlp_decode_and_convert |
942.77 µs | -46.3% |
acceptor_ingest_metrics/otlp_convert_only |
610.41 µs | -46.7% |
single_batch_writes/100_rows_0.0MB |
857.96 µs | -34.0% |
single_batch_writes/1000_rows_0.4MB |
1.54 ms | -39.1% |
single_batch_writes/10000_rows_3.5MB |
7.76 ms | -38.9% |
single_batch_writes/100000_rows_37.4MB |
62.58 ms | -42.6% |
multi_batch_writes/2_batches_2000_rows |
2.40 ms | -39.3% |
multi_batch_writes/5_batches_5000_rows |
4.97 ms | -39.2% |
multi_batch_writes/10_batches_10000_rows |
9.20 ms | -39.3% |
multi_batch_writes/20_batches_20000_rows |
18.19 ms | -39.2% |
writer/creation |
686.67 µs | -41.8% |
concurrent_writes/2_writers |
1.46 ms | -31.4% |
concurrent_writes/4_writers |
2.45 ms | -27.1% |
concurrent_writes/8_writers |
4.67 ms | -28.8% |
ingest_sort/in_order/1000 |
28.89 µs | -49.4% |
ingest_sort/shuffled/1000 |
47.67 µs | -50.7% |
ingest_sort/in_order/10000 |
251.26 µs | -50.0% |
ingest_sort/shuffled/10000 |
546.82 µs | -50.3% |
ingest_sort/in_order/100000 |
2.80 ms | -48.6% |
ingest_sort/shuffled/100000 |
13.27 ms | -36.3% |
schema_transform/transform_trace_v1_to_v2 |
986.89 µs | -26.0% |
compactor/rewrite_6_files |
10.36 ms | -42.9% |
querier_read/trace_lookup_by_id_unbounded |
11.30 ms | -48.7% |
querier_read/trace_lookup_by_id_cold_without_cache |
10.98 ms | -48.2% |
querier_read/trace_lookup_by_id_cold_with_cache |
11.09 ms | -48.1% |
querier_read/trace_lookup_by_id_warm_with_cache |
10.84 ms | -48.2% |
querier_read/trace_lookup_by_id_windowed |
2.88 ms | -46.4% |
querier_read/trace_lookup_by_id_via_index |
6.62 ms | -46.7% |
querier_read/trace_search_groups |
10.84 ms | -48.3% |
declared_ordering/recent_first_topk/attested |
4.02 ms | -46.6% |
declared_ordering/recent_first_topk/attested_split_off |
3.98 ms | -46.4% |
declared_ordering/recent_first_topk/unattested |
3.97 ms | -45.8% |
declared_ordering/oldest_first_topk/attested |
3.35 ms | -46.8% |
declared_ordering/oldest_first_topk/attested_split_off |
3.38 ms | -46.7% |
declared_ordering/oldest_first_topk/unattested |
3.80 ms | -46.3% |
declared_ordering/ordered_full_scan/attested |
6.11 ms | -50.1% |
declared_ordering/ordered_full_scan/attested_split_off |
8.09 ms | -52.1% |
declared_ordering/ordered_full_scan/unattested |
6.55 ms | -49.2% |
querier_service/find_trace_by_id |
11.63 ms | -45.9% |
querier_service/find_trace_by_id_hinted |
2.92 ms | -43.1% |
querier_service/search_traces_recent |
34.48 ms | -44.9% |
querier_service/promql_range_avg_by_service |
54.63 ms | -48.1% |
querier_service/logql_line_filter |
57.84 ms | -46.7% |
trace_index_scaling/10000 |
698.63 µs | -35.1% |
trace_index_scaling/100000 |
666.42 µs | -38.3% |
trace_index_scaling/1000000 |
687.60 µs | -37.8% |
What is benchmarked¶
All benches live behind a per-crate benchmarks feature (harness = false
Criterion targets) so they never enter a normal build. CI's clippy step on core PRs
(--all-targets --all-features) compiles them, so a bench that stops
building fails CI.
| Crate | Bench target | Measures |
|---|---|---|
common |
ingest_and_wal |
Acceptor CPU per OTLP trace request: protobuf decode + OTLP→Arrow; WAL record_batch_to_bytes/bytes_to_record_batch round-trip |
common |
signal_decode |
Same decode + convert for logs and metrics requests |
writer |
schema_transform_benchmarks |
transform_trace_v1_to_v2 — the wire→storage materialization plan |
writer |
iceberg_benchmarks |
IcebergTableWriter::append_batches_with_marker across batch sizes, multi-batch commits, and concurrent tenants; writer creation cost separately; ingest_sort isolates the per-commit-group sort by the declared key on in-order and shuffled input |
tests-integration |
querier_read_paths |
Trace lookup by id (unbounded, time-windowed, via a point index, and with/without the Parquet footer cache) and the trace-groups listing over a seeded Iceberg table; recent_first_topk times ORDER BY timestamp DESC LIMIT n over sequential files in every attestation state and prints what each scan opened and pruned (see Declared sort orders) |
tests-integration |
querier_service_read_paths |
The real querier (QuerierFlightService::do_get with router ticket formats): bloom-pruned find_trace with and without a time hint, search_traces, a PromQL range query and a LogQL line filter through the actual engines |
tests-integration |
compaction |
CompactionExecutor::execute_candidate rewriting a set of small files |
tests-integration |
trace_index_scaling |
Point lookup against a prefix-sharded, bloom-filtered Parquet index at 10k → 1M traces |
The inputs come from shared fixtures: common::testing::sample_trace_request
and friends for OTLP payloads, and tests_integration::generators for seeded
Iceberg tables. Reuse those rather than hand-building data in a new bench.
Running locally¶
# Everything, with Criterion's HTML report under target/criterion/report/
scripts/run-benches.sh
# One crate
scripts/run-benches.sh -p writer
# One target directly
cargo bench -p common --features benchmarks --bench ingest_and_wal
scripts/run-benches.sh discovers [[bench]] targets through
cargo metadata, so a new bench only needs its Cargo.toml entry. Anything
after -- is passed to Criterion; a quick smoke run looks like
scripts/run-benches.sh -p common -- --warm-up-time 0.5 --measurement-time 1 --sample-size 10
Close other heavy processes first: the writer and querier benches are sensitive to CPU contention, and Criterion will happily report a 30% "change" that is your IDE indexing.
Comparing a branch against main¶
Criterion baselines make an A/B comparison two commands:
git switch main
scripts/run-benches.sh -- --save-baseline main
git switch my-branch
scripts/run-benches.sh -- --baseline main
The second run prints each benchmark's change against the saved baseline
with a significance verdict (No change in performance detected /
Performance has improved / Performance has regressed). Baselines are
stored under target/criterion/<bench>/<baseline>/ and survive cargo
clean-free rebuilds; use --save-baseline with a branch name so several can
coexist.
For a PR that claims a performance win, paste the relevant --baseline main
lines into the description — see #1245 for the pattern.
Nightly trend¶
.github/workflows/benchmarks.yml runs the full suite once a night (and on
workflow_dispatch) in release mode on a single runner, feeds Criterion's
--output-format bencher output to
github-action-benchmark,
and pushes the accumulated series to the benchmark-data branch under
dev/bench/. The docs build copies that directory into the site as
cedricziel.github.io/signaldb/benchmarks and regenerates the table above with
scripts/render-bench-summary.py.
The workflow points CRITERION_HOME at the runner's temp directory so
Criterion's own data never lands in the cached target/. rust-cache strips
the files but not the directories from target/criterion/ before saving, and
a restored tree of empty base/ directories makes Criterion attempt a
baseline comparison, fail, and print the error into the bencher output the
next step parses.
An alert-threshold of 150% flags a benchmark whose mean is 1.5× the
previous run in the job summary. It does not fail the workflow yet: shared
runners are noisy, and the threshold needs a few weeks of observed variance
before it becomes a gate.
Declared sort orders: what the benchmark shows¶
querier_read_paths's declared_ordering group is the measurement gate of
the declared-sort-orders work (#936, #1317). It seeds one traces table of 60
sequential ingest files (1,000 spans each, two minutes apart, one row group
per file, 2.5 MB) twice — once through ingest, so every file attests the
declared (timestamp, trace_id) order, and once through the plain write path,
so none does — and runs three ordered shapes against the querier's real
session. Before timing, it prints what each scan did, from DataFusion's own
metrics, because wall-clock alone cannot tell an elided sort from a quiet
machine. These are the numbers from the run that closed #1317 (in-memory
object store, warm footer cache, 4 vCPUs):
| Shape | Files | Reached | Pruned | Read | Bytes | Sort | Time |
|---|---|---|---|---|---|---|---|
ORDER BY timestamp DESC LIMIT 20 |
attested | 60 | 56 | 4 | 23,641 | kept | 13.3 ms |
| attested, split off | 60 | 58 | 2 | 11,820 | kept | 13.2 ms | |
| unattested | 60 | 59 | 2 | 11,820 | kept | 12.0 ms | |
ORDER BY timestamp ASC LIMIT 20 |
attested | 4 | 0 | 4 | 23,641 | elided | 10.6 ms |
| attested, split off | 2 | 0 | 2 | 11,824 | elided | 10.2 ms | |
| unattested | 60 | 58 | 2 | 11,817 | kept | 12.1 ms | |
ORDER BY timestamp ASC (all 60,000 rows) |
attested | 60 | 0 | 60 | 354,626 | elided | 20.9 ms |
| attested, split off | 60 | 0 | 60 | 354,626 | elided | 26.7 ms | |
| unattested | 60 | 0 | 60 | 354,626 | kept | 25.4 ms |
"Reached" counts files the scan got as far as preparing to open; "pruned"
counts the reached files it then skipped whole on their statistics without
reading them; "read" is the rest, one row group each in this layout; "split
off" is [querier.datafusion].split_file_groups_by_statistics = false. The
pruned/read split of a TopK arm moves by one file between runs: the dynamic
filter tightens as partitions race to fill the heap.
What the mechanism columns say:
- Recent-first (
DESC) gains nothing from attestation. DataFusion 54 cannot serve the reverse of a declared order without a sort, so every arm keeps aSortExec: TopK. What makes the shape cheap is the TopK's dynamic filter: the scan reads files newest-first, fills the heap from the first file per group, and prunes the other 56–59 files on statistics. That works on any file with statistics, attested or not — which is why the unattested arm is not slower. - Oldest-first (
ASC) is where attestation pays. The request matches the declared order, so the attested scan declares it, the sort is elided, and the limit stops the scan after the first file of each group: 4 files reached instead of 60. On an object store each of those 56 unreached footers is a round-trip the query never makes; in memory it is the ~1.5 ms between the arms. - A full ordered scan elides the sort outright. Attested: rows stream out in file order. Unattested: 60,000 rows are sorted first. The difference is the sort itself (~4.5 ms here, and the sort's memory).
- Every timing includes ~10 ms of planning — the Iceberg provider reads manifests and builds the scan on every query — which is why the ratios look small next to the I/O ratios. The rows are the same across arms; the bench asserts it.
The write-path cost of the sort that makes attestation possible is
iceberg_benchmarks's ingest_sort group: the columnar sort of one commit
group by the table's key, on the same metrics batches
single_batch_writes appends. Ingest's usual input arrives close to time
order, which is the cheap case; a shuffled group is the expensive one.
| Rows | Append (single_batch_writes) |
Sort, in order | Sort, shuffled |
|---|---|---|---|
| 1,000 | 4.6 ms | 75 µs | 111 µs |
| 10,000 | 15.3 ms | 748 µs | 1.45 ms |
| 100,000 | 125.6 ms | 20.0 ms | 35.6 ms |
The sort is 2–5% of the append for the commit groups ingest actually forms (thousands of rows) and reaches 16% (in order) to 28% (shuffled) only at 100,000-row groups, where the append is already 125 ms. It is bounded by the group, not the table.
Adding a bench¶
- Put the file under the owning crate's
benches/(ortests-integration/benches/when it needs the writer to seed a table plus the querier or compactor). - Add a
[[bench]]entry withharness = falseandrequired-features = ["benchmarks"]. - Use
Throughput::Elementsfor per-record work,iter_customwhen the measured operation mutates state that must be rebuilt outside the timed region (seetests-integration/benches/compaction.rs), and assert on the result once up front so the bench cannot silently time an empty scan. - Keep the header comment honest about what is and is not measured.
- Add a row to the table on this page.