Skip to content

feat(blockdev): EIP-7870 mixed workloads and QD1 latency in the storage probe - #347

Merged
skylenet merged 3 commits into
masterfrom
feat/probe-eip-mixed-and-qd1
Sep 30, 2026
Merged

skylenet merged 3 commits into
masterfrom
feat/probe-eip-mixed-and-qd1

Conversation

@skylenet

@skylenet skylenet commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Problem

  1. The probe did not match the EIP-7870 fio commands:

    fio --randrepeat=1 --ioengine=libaio --direct=1 --gtod_reduce=1 --name=test --filename=test --bs=4k --iodepth=64 --size=4G --readwrite=randrw --rwmixread=75
    fio --name=seq --filename=test --direct=1 --ioengine=libaio --iodepth=64 --bs=1M --size=4G --rw=readwrite
    
    • EIP-7870 uses one mixed run for each pair of values. The probe ran separate read and write runs.
    • EIP-7870 uses one thread with libaio at depth 64. The probe used 64 threads with blocking I/O.
    • EIP-7870 stops after 4 GiB of I/O. The probe ran for a fixed time.

    On benchmark-ci-8, the old method gave about 4× the random IOPS of the EIP command on the same disk.

  2. The probe did not measure QD1. A client reads state mostly one key at a time, so each read waits for the drive. Drive datasheets give a QD1 line for this reason. For example, a Samsung 970 EVO lists 15,000 read and 50,000 write IOPS at QD1, and 500,000 / 480,000 at QD32.

Changes

  • EIP-7870 workloads: they now use the method of the fio commands.
    • On Linux, one goroutine keeps io_depth I/Os in flight with native AIO (io_setup, io_submit, io_getevents), as fio --ioengine=libaio with one job does. It submits one I/O for each io_submit call, as fio does by default (iodepth_batch_submit=1).
    • Each I/O is a read at random with a chance of 75% (random run) or 50% (sequential run). A sequential run keeps one offset for reads and one for writes, as fio does (checked with the fio I/O log).
    • A run stops after file_size bytes of reads and writes, as fio --size does, with a cap of 2 minutes for a very slow disk.
    • The existing rand_* and seq_* fields hold the values. The new rand_read_percent and seq_read_percent fields mark a mixed result.
    • On other OS, the goroutine method stays for the tests. The probe runs only on Linux.
  • QD1 workloads: 4 KiB random read and 4 KiB random write, with one I/O in flight and blocking I/O (as fio psync). New fields qd1_rand_read and qd1_rand_write with iops, p50_us and p99_us. A histogram with 1 µs buckets up to 20 ms records the latency, and it keeps each slower latency as an exact value. With one I/O in flight, a run has at most duration / 20 ms slow values (250 in 5s), so the memory stays small, and a p99 above 20 ms (HDD, or a drive busy with GC) is still a measured value. The writes do not use fsync, as in the datasheets. duration (default 5s) now applies only to these two workloads.
  • Order: QD1 read, QD1 write, random mixed, sequential mixed. The QD1 runs go first, so the drive is not busy with garbage collection from the deep runs.
  • Default file size: 4GB, as in EIP-7870. The probe now needs 5 GiB free in the data mount.
  • UI:
    • The Storage panel shows the mix in the row labels (for example "75/25 mix") and a new QD1 table with IOPS, p50 and p99.
    • The note under the table describes the method, or tells that the values come from an older probe.
    • The compare view has a new "Disk QD1 Random IOPS (R/W)" row.
  • Markdown export and runner log: they show the QD1 values and the mix.
  • Docs: docs/configuration.md and config.example.yaml.

Check against fio

On benchmark-ci-8, which was idle (2× Samsung PM9A1 512 GB in md RAID, ext4, Ryzen 7 PRO 8700GE), 3 runs each on the same disk:

Workload EIP fio command Probe, this PR Probe, 64 threads (before)
Random 4K 75/25, read IOPS 116,865–117,829 135,277–136,990 458,673–459,922
Random 4K 75/25, write IOPS 39,057–39,379 45,114–45,715 153,116–153,431
Sequential 1M 50/50, read MB/s 1,124–1,221 781–3,595 1,463–3,389
QD1 workload fio psync, 1 job Probe, this PR
Random read 4K 18,429 IOPS, p50 54 µs, p99 56 µs 18,122–18,139 IOPS, p50 54 µs, p99 64 µs
Random write 4K 42,786 IOPS, p50 23 µs 42,754–42,844 IOPS, p50 23 µs
  • Random: the probe is now 1.16× the EIP command (before: 3.9×). Both are one thread limited by CPU (fio used 87% of one core). The rest of the difference is fio's own CPU cost for each I/O.
  • Sequential: the method is the same, but these consumer drives give very different values from run to run. In the same session, the EIP fio command gave 1.1–2.8 GB/s, and a 5 s time-based fio run gave 2.3–3.0 GB/s. The spread comes from the drive state (SLC cache), not from the method.
  • QD1: the probe and fio agree within 2%.

Compatibility

Older results have no rand_read_percent, seq_read_percent or qd1_* fields. The UI shows them as before, with a note that their values are higher than the EIP-7870 mixed runs.

Tests

  • go test ./pkg/blockdev/ ./pkg/executor/ passes on macOS.
  • The Linux test binary passes on benchmark-ci-8, with the AIO tests: random mix, sequential mix, cancel, bad file, and stop after ioBytes.
  • go test -tags containers_image_openpgp ./pkg/runner/ passes.
  • golangci-lint: 0 issues for macOS and Linux. go vet passes for linux/arm64.
  • UI: tsc, eslint and npm run build pass.
  • pkg/config has a test that fails on macOS (TestValidateDataDirMethods_SchelkBinary, it opens /proc/mounts). It fails on master too.

…ge probe

The probe compared separate read and write runs with the EIP-7870
floors. The EIP-7870 fio commands use one mixed run: 4 KiB random I/O
with 75% reads, and 1 MiB sequential I/O with 50% reads, at I/O depth
64 on a 4 GiB file. Separate runs give each direction the whole drive,
so the values were too high for that comparison.

The probe also did not measure QD1. A client reads state mostly one
key at a time, so the QD1 latency is closer to the client workload.

- The deep workloads are now the two EIP-7870 mixed runs.
  rand_read_percent and seq_read_percent mark a mixed result.
- Two new QD1 workloads: 4 KiB random read and write, with IOPS, p50
  and p99 latency (1 µs histogram, fixed memory).
- The default file size is 4GB, as in EIP-7870.
- The UI shows the mix in the labels, a QD1 table, and a QD1 row in
  the compare view. It marks the results of older probes.
- The Markdown export and the runner log show the QD1 values.

@redpandabot redpandabot Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

The new commit replaces the 64-thread psync deep workloads with a single-thread native-AIO workload matching fio's libaio method, which is well built: correct iocb/io_event ABI layouts with compile-time assertions, safe slot reuse, and clean exit paths with io_destroy plus KeepAlive on every return. The one previously open finding remains: the 20 ms histogram cap on QD1 p99 latency is still unmentioned in the docs and UI.

Issues

  • 🟡 pkg/blockdev/probe.go:410 — 20 ms p99 cap is silent in the UI and docs — see the thread on that line

Reviewed @ bcc04464
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts." — Fred Brooks

Comment thread pkg/blockdev/probe.go Outdated
func (h *latencyHistogram) add(d time.Duration) {
us := d.Microseconds()
if us > maxLatencyUs {
us = maxLatencyUs

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 20 ms p99 cap is silent in the UI and docs

Latencies above maxLatencyUs are clamped into the top bucket, so on an HDD or a drive busy with GC the reported p99 shows as 20.0 ms while the real value can be several times higher; neither the QD1 table (StorageInfoPanel.tsx) nor the updated configuration.md mentions the cap. A saturation indicator or a docs note would prevent misreading a clamped value as measured.

The mixed workloads used 64 goroutines with blocking I/O. The EIP-7870
fio commands use one job with libaio: one thread keeps 64 I/Os in
flight. On a fast drive, that thread is limited by its CPU time. A
check against fio on benchmark-ci-8 gave 460K random reads for the
probe and 118K for the EIP command.

- On Linux, the mixed workloads use native AIO (io_setup, io_submit,
  io_getevents) from one goroutine. They submit one I/O for each
  io_submit call, as fio does by default.
- They stop after file_size bytes of I/O, as fio --size does, with a
  cap of 2 minutes. duration applies only to the QD1 workloads.
- Other OS keep the goroutine method, for the tests.
The latency histogram put each I/O slower than 20 ms into its top
bucket. On an HDD, or on a drive busy with garbage collection, the p99
then showed as 20 ms, and the real value could be several times higher.

The histogram now keeps each latency above 20 ms as an exact value. It
records one I/O in flight, so a run of duration d has at most d / 20 ms
such values (250 in 5s), and the memory stays small.

@redpandabot redpandabot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The probe now runs the EIP-7870 fio method (one thread, native AIO, mixed read/write runs bounded by file size) and adds QD1 4 KiB random read/write latency with a fixed-size 1 µs histogram. The previously open finding — p99 values above 20 ms being silently clamped — is fixed by d041446, which keeps slow latencies as exact values with test coverage. The AIO lifecycle, error/cancel paths, and UI/docs/markdown surfacing all check out; no new problems found.


Reviewed 14 changed file(s) @ d041446f — no blocking issues found.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts." — Fred Brooks

@redpandabot

redpandabot Bot commented Sep 30, 2026

Copy link
Copy Markdown

CI failed @ d041446f

1 failed job: the Go CI job failed on a single flaky test in pkg/api/indexer, a package untouched by this PR.

  • Check - CI Go / lint-and-test — flaky — Race in TestIndexer_RecordsManualTrigger: the test's require.Eventually polls only for the pass row to appear in the store, then asserts idx.State().Running is false. In pkg/api/indexer/indexer.go, runPassInner writes the pass row via idx.recordPass(ctx, pass) (line 329) and only afterwards does the deferred idx.running.Store(false) (line 262) clear Running, so the test can observe the row while Running is still true. The provided log tail did not contain the failure (its bare FAIL is go's summary); the full job log fetched via gh shows the only failure: '--- FAIL: TestIndexer_RecordsManualTrigger (0.11s)' with 'indexer_pass_test.go:126: Error: Should be false / Messages: the pass is done' in package github.com/ethpandaops/benchmarkoor/pkg/api/indexer. The PR diff touches only pkg/blockdev, pkg/config, pkg/executor, pkg/runner/storage.go and ui/, nothing reaching pkg/api/indexer; master's most recent Check - CI (Go) run (36712685817) is green, so this is an inherent timing race that had not fired yet, not a deterministic pre-existing failure. (pkg/api/indexer/indexer_pass_test.go:126) — Make the test's wait cover the condition it asserts: change the require.Eventually at pkg/api/indexer/indexer_pass_test.go:119-121 to 'return len(f.passes()) == 1 && !f.idx.State().Running' (or add a require.Eventually for !f.idx.State().Running before line 125). Do not clear running before recordPass, which would break the documented invariant that a running pass is invisible in the pass history.

Diagnosed Check - CI Go @ d041446f
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts." — Fred Brooks

@skylenet
skylenet merged commit 3a17faa into master Sep 30, 2026
9 of 10 checks passed
@skylenet
skylenet deleted the feat/probe-eip-mixed-and-qd1 branch September 30, 2026 13:20
@envuladu

Copy link
Copy Markdown

New

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants