feat(blockdev): EIP-7870 mixed workloads and QD1 latency in the storage probe - #347
Conversation
…ge probe The probe compared separate read and write runs with the EIP-7870 floors. The EIP-7870 fio commands use one mixed run: 4 KiB random I/O with 75% reads, and 1 MiB sequential I/O with 50% reads, at I/O depth 64 on a 4 GiB file. Separate runs give each direction the whole drive, so the values were too high for that comparison. The probe also did not measure QD1. A client reads state mostly one key at a time, so the QD1 latency is closer to the client workload. - The deep workloads are now the two EIP-7870 mixed runs. rand_read_percent and seq_read_percent mark a mixed result. - Two new QD1 workloads: 4 KiB random read and write, with IOPS, p50 and p99 latency (1 µs histogram, fixed memory). - The default file size is 4GB, as in EIP-7870. - The UI shows the mix in the labels, a QD1 table, and a QD1 row in the compare view. It marks the results of older probes. - The Markdown export and the runner log show the QD1 values.
There was a problem hiding this comment.
Summary
The new commit replaces the 64-thread psync deep workloads with a single-thread native-AIO workload matching fio's libaio method, which is well built: correct iocb/io_event ABI layouts with compile-time assertions, safe slot reuse, and clean exit paths with io_destroy plus KeepAlive on every return. The one previously open finding remains: the 20 ms histogram cap on QD1 p99 latency is still unmentioned in the docs and UI.
Issues
- 🟡
pkg/blockdev/probe.go:410— 20 ms p99 cap is silent in the UI and docs — see the thread on that line
Reviewed @ bcc04464
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts." — Fred Brooks
| func (h *latencyHistogram) add(d time.Duration) { | ||
| us := d.Microseconds() | ||
| if us > maxLatencyUs { | ||
| us = maxLatencyUs |
There was a problem hiding this comment.
🟢 20 ms p99 cap is silent in the UI and docs
Latencies above maxLatencyUs are clamped into the top bucket, so on an HDD or a drive busy with GC the reported p99 shows as 20.0 ms while the real value can be several times higher; neither the QD1 table (StorageInfoPanel.tsx) nor the updated configuration.md mentions the cap. A saturation indicator or a docs note would prevent misreading a clamped value as measured.
The mixed workloads used 64 goroutines with blocking I/O. The EIP-7870 fio commands use one job with libaio: one thread keeps 64 I/Os in flight. On a fast drive, that thread is limited by its CPU time. A check against fio on benchmark-ci-8 gave 460K random reads for the probe and 118K for the EIP command. - On Linux, the mixed workloads use native AIO (io_setup, io_submit, io_getevents) from one goroutine. They submit one I/O for each io_submit call, as fio does by default. - They stop after file_size bytes of I/O, as fio --size does, with a cap of 2 minutes. duration applies only to the QD1 workloads. - Other OS keep the goroutine method, for the tests.
The latency histogram put each I/O slower than 20 ms into its top bucket. On an HDD, or on a drive busy with garbage collection, the p99 then showed as 20 ms, and the real value could be several times higher. The histogram now keeps each latency above 20 ms as an exact value. It records one I/O in flight, so a run of duration d has at most d / 20 ms such values (250 in 5s), and the memory stays small.
There was a problem hiding this comment.
The probe now runs the EIP-7870 fio method (one thread, native AIO, mixed read/write runs bounded by file size) and adds QD1 4 KiB random read/write latency with a fixed-size 1 µs histogram. The previously open finding — p99 values above 20 ms being silently clamped — is fixed by d041446, which keeps slow latencies as exact values with test coverage. The AIO lifecycle, error/cancel paths, and UI/docs/markdown surfacing all check out; no new problems found.
Reviewed 14 changed file(s) @ d041446f — no blocking issues found.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts." — Fred Brooks
CI failed @
|
|
New |
Problem
The probe did not match the EIP-7870 fio commands:
libaioat depth 64. The probe used 64 threads with blocking I/O.On benchmark-ci-8, the old method gave about 4× the random IOPS of the EIP command on the same disk.
The probe did not measure QD1. A client reads state mostly one key at a time, so each read waits for the drive. Drive datasheets give a QD1 line for this reason. For example, a Samsung 970 EVO lists 15,000 read and 50,000 write IOPS at QD1, and 500,000 / 480,000 at QD32.
Changes
io_depthI/Os in flight with native AIO (io_setup,io_submit,io_getevents), as fio--ioengine=libaiowith one job does. It submits one I/O for eachio_submitcall, as fio does by default (iodepth_batch_submit=1).file_sizebytes of reads and writes, as fio--sizedoes, with a cap of 2 minutes for a very slow disk.rand_*andseq_*fields hold the values. The newrand_read_percentandseq_read_percentfields mark a mixed result.psync). New fieldsqd1_rand_readandqd1_rand_writewithiops,p50_usandp99_us. A histogram with 1 µs buckets up to 20 ms records the latency, and it keeps each slower latency as an exact value. With one I/O in flight, a run has at mostduration/ 20 ms slow values (250 in 5s), so the memory stays small, and a p99 above 20 ms (HDD, or a drive busy with GC) is still a measured value. The writes do not usefsync, as in the datasheets.duration(default 5s) now applies only to these two workloads.4GB, as in EIP-7870. The probe now needs 5 GiB free in the data mount.docs/configuration.mdandconfig.example.yaml.Check against fio
On benchmark-ci-8, which was idle (2× Samsung PM9A1 512 GB in md RAID, ext4, Ryzen 7 PRO 8700GE), 3 runs each on the same disk:
psync, 1 jobCompatibility
Older results have no
rand_read_percent,seq_read_percentorqd1_*fields. The UI shows them as before, with a note that their values are higher than the EIP-7870 mixed runs.Tests
go test ./pkg/blockdev/ ./pkg/executor/passes on macOS.ioBytes.go test -tags containers_image_openpgp ./pkg/runner/passes.golangci-lint: 0 issues for macOS and Linux.go vetpasses for linux/arm64.tsc,eslintandnpm run buildpass.pkg/confighas a test that fails on macOS (TestValidateDataDirMethods_SchelkBinary, it opens/proc/mounts). It fails onmastertoo.