Repository navigation
fix(stats): count a stacked block device once in the disk metrics - #346
Merged
Merged
Conversation
The kernel charges an I/O to the cgroup on each block layer that it goes through. On a schelk host, io.stat lists the datadir I/O once for dm-0 and once for the NVMe under it, and the root filesystem I/O once for md2 and once for each member disk. The cgroup reader added all the lines, so the disk bytes and IOPS were 2x too high. - A new stackedDevices helper reads sysfs. It finds a device that another block device holds, directly or through a partition. - The cgroup reader and the Docker reader skip such a device. - Without sysfs (not Linux), every device counts, as before.
SummaryThe PR skips block devices that a dm/md device holds (directly or via a partition) when summing cgroup io.stat and Docker blkio stats, fixing the 2× double-count of disk bytes/IOPS on stacked-device hosts. The implementation is correct and behavior-preserving where sysfs is unavailable: parsing is unchanged, the cache is mutex-guarded and nil-safe, and all callers were updated. One narrow edge: the docker fallback reader classifies devices using the local /sys, which is the wrong host if DOCKER_HOST points at a remote daemon. Issues
Reviewed @ |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The disk bytes and IOPS of a client are 2× too high on hosts with a stacked block device (dm, LVM, md). This includes all the schelk hosts.
The kernel charges an I/O to the cgroup on each block layer that it goes through. The cgroup reader added all the lines of
io.stat. Example from a client container on benchmark-ci-11:dm-0line and thenvme2n1line are identical.md2line is the sum ofnvme0n1andnvme1n1.io.statlists the whole disk (nvme2n1), not the partition thatdm-0uses (nvme2n1p2).One ethrex test showed 8.1 GB/s of reads. That is more than a PCIe 4.0 x4 drive can do.
Fix
stackedDeviceshelper inpkg/stats/stacked.goreads/sys/dev/block/<maj:min>. It finds a device that another block device holds, directly or through a partition (holders). It caches the answer for each device.readIOStats(cgroup reader) andextractBlkioBytes/extractBlkioOps(Docker reader) skip such a device./sys), every device counts, as before.Limit: a disk with one partition under dm-0 and another partition in direct use is skipped as a whole.
io.statdoes not split the I/O by partition, so the direct partition I/O is not counted.Effect
The disk values of new runs on stacked devices go down by half. Old runs keep their values. Thus a compare between old and new runs shows a drop in disk I/O that is not real.
Tests
go test ./pkg/stats/passes, with a fake sysfs tree and the realio.statlines from benchmark-ci-11.golangci-lint run ./pkg/stats/: 0 issues.stacked.goran on benchmark-ci-11 against the live containerio.stat. It marked259:0,259:1and259:2as lower devices, and it counted253:0and9:2.