Skip to content

fix(stats): count a stacked block device once in the disk metrics - #346

Merged
skylenet merged 1 commit into
masterfrom
fix/stats-stacked-device-double-count
Sep 30, 2026
Merged

skylenet merged 1 commit into
masterfrom
fix/stats-stacked-device-double-count

Conversation

@skylenet

Copy link
Copy Markdown
Member

Problem

The disk bytes and IOPS of a client are 2× too high on hosts with a stacked block device (dm, LVM, md). This includes all the schelk hosts.

The kernel charges an I/O to the cgroup on each block layer that it goes through. The cgroup reader added all the lines of io.stat. Example from a client container on benchmark-ci-11:

259:0 rbytes=3943170048 wbytes=184811520 rios=962688 wios=45120   <- nvme2n1 (under dm-0)
253:0 rbytes=3943170048 wbytes=184811520 rios=962688 wios=45120   <- dm-0 (bench_era)
259:1 rbytes=97472512   wbytes=11702272  rios=1470   wios=100     <- nvme0n1 (under md2)
259:2 rbytes=99926016   wbytes=10448896  rios=1580   wios=96      <- nvme1n1 (under md2)
9:2   rbytes=197398528  wbytes=22151168  rios=2732   wios=180     <- md2 (root RAID)
  • The dm-0 line and the nvme2n1 line are identical.
  • The md2 line is the sum of nvme0n1 and nvme1n1.
  • io.stat lists the whole disk (nvme2n1), not the partition that dm-0 uses (nvme2n1p2).

One ethrex test showed 8.1 GB/s of reads. That is more than a PCIe 4.0 x4 drive can do.

Fix

  • The new stackedDevices helper in pkg/stats/stacked.go reads /sys/dev/block/<maj:min>. It finds a device that another block device holds, directly or through a partition (holders). It caches the answer for each device.
  • readIOStats (cgroup reader) and extractBlkioBytes / extractBlkioOps (Docker reader) skip such a device.
  • Without sysfs (not Linux, or no /sys), every device counts, as before.

Limit: a disk with one partition under dm-0 and another partition in direct use is skipped as a whole. io.stat does not split the I/O by partition, so the direct partition I/O is not counted.

Effect

The disk values of new runs on stacked devices go down by half. Old runs keep their values. Thus a compare between old and new runs shows a drop in disk I/O that is not real.

Tests

  • go test ./pkg/stats/ passes, with a fake sysfs tree and the real io.stat lines from benchmark-ci-11.
  • golangci-lint run ./pkg/stats/: 0 issues.
  • A Linux build of stacked.go ran on benchmark-ci-11 against the live container io.stat. It marked 259:0, 259:1 and 259:2 as lower devices, and it counted 253:0 and 9:2.

The kernel charges an I/O to the cgroup on each block layer that it
goes through. On a schelk host, io.stat lists the datadir I/O once for
dm-0 and once for the NVMe under it, and the root filesystem I/O once
for md2 and once for each member disk. The cgroup reader added all the
lines, so the disk bytes and IOPS were 2x too high.

- A new stackedDevices helper reads sysfs. It finds a device that
  another block device holds, directly or through a partition.
- The cgroup reader and the Docker reader skip such a device.
- Without sysfs (not Linux), every device counts, as before.
@redpandabot

redpandabot Bot commented Sep 30, 2026

Copy link
Copy Markdown

Summary

The PR skips block devices that a dm/md device holds (directly or via a partition) when summing cgroup io.stat and Docker blkio stats, fixing the 2× double-count of disk bytes/IOPS on stacked-device hosts. The implementation is correct and behavior-preserving where sysfs is unavailable: parsing is unchanged, the cache is mutex-guarded and nil-safe, and all callers were updated. One narrow edge: the docker fallback reader classifies devices using the local /sys, which is the wrong host if DOCKER_HOST points at a remote daemon.

Issues

  • 🟢 pkg/stats/docker_reader.go:76 — docker reader assumes the daemon host's sysfs is the local /sys — The docker reader is the fallback when the cgroup path isn't detectable, and the client is built with client.FromEnv (pkg/docker/docker.go:145), which honors a remote DOCKER_HOST. With a remote daemon, stackedDevices consults the local machine's /sys/dev/block and could wrongly skip the remote container's devices, silently undercounting disk I/O. Docs only document same-host unix sockets, so this is an edge case; a debug log when a sysfs lookup finds nothing might help diagnosis.

Reviewed @ 53f67949
"Debugging is twice as hard as writing the code in the first place." — Brian Kernighan

@skylenet
skylenet merged commit be2df5a into master Sep 30, 2026
8 checks passed
@skylenet
skylenet deleted the fix/stats-stacked-device-double-count branch September 30, 2026 12:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant