Skip to content

dbus: report the kernel's cgroup freezer files when TestFreezer fails - #531

Open
amccabe wants to merge 1 commit into
coreos:mainfrom
amccabe:freezer-test-kernel-state
Open

amccabe wants to merge 1 commit into
coreos:mainfrom
amccabe:freezer-test-kernel-state

Conversation

@amccabe

@amccabe amccabe commented Sep 19, 2026 •

Copy link
Copy Markdown

Description

Three out of fifty-four recent runs of TestFreezer on ubuntu:24.04 (one on each of #528, #529, and #530) failed with

unit is not frozen after calling FreezeUnit(), FreezerState=running

All three ran on 2026-09-18, all three on runners the job log places in Azure centralus, which was the only region with a TestFreezer failure that day. Other centralus jobs in the same hour passed, so while it appears to be regional, there is another factor in the failure.

Broken down by distro and systemd version, across the same runs (ubuntu:22.04 and debian:bookworm from #530's branch, the rest from the matrix on main):

Distro systemd TestFreezer
ubuntu:24.04 255.4 3 failed of 54
debian:trixie 257 0 of 54
fedora 259 0 of 54
debian:bullseye 247 0 of 39 that reached the tests
debian:bookworm 252 0 of 3
ubuntu:22.04 249 0 of 3
ubuntu:20.04 245, no freezer skipped

So it appeared to be specific to 255 in centralus. The three failing jobs ran the same container systemd build as August's passes, and the same runner image and base image as the twelve ubuntu:24.04 jobs that passed the same day. The failing tests completed faster than passing ones (0.01 to 0.02 s against 0.03 to 0.05 s), so the reply came back early rather than the system being slow.

Suspected cause

Reading upstream systemd v255's freezer handling: FreezeUnit writes the cgroup's cgroup.freeze, marks the unit freezing, and defers its reply until a notification on cgroup.events makes it re-read that file. The handler treats frozen 0 read while the unit is freezing as the unit having thawed, replies success, and sets the state to running; a later frozen 1 is then ignored as not initiated by systemd. That is the only path in v255 to a successful reply with FreezerState=running. The notification that triggers the premature read is most likely the populated change from the service having just started, which the kernel delivers through a workqueue; if that delivery is delayed past the D-Bus round trip, for example by a vCPU stall on a busy host, it lands after the freeze has been written and before the sleeping task has taken its freezer trap. I was unable to reproduce it locally in about 3,000 iterations under CPU, steal, and notification load, which fits a trigger that lives on the host.

What this change does

The assertions are unchanged. Both failure messages now include the kernel's view of the unit's cgroup, cgroup.freeze and cgroup.events, resolved from the unit's ControlGroup property. With this change, the next failure would read something like

FreezerState=running, cgroup.freeze="1" cgroup.events="populated 1 frozen 1"

That logging indicates cgroup.freeze: 1, which would mean systemd wrote the freeze and its bookkeeping disagrees with the kernel, which I'd expect from these failed tests. cgroup.freeze: 0 would mean the freeze was never written or that it was reverted (some other cause). The ThawUnit message also said "not frozen" where it meant "not running".

This is extra diagnosis rather than a test fix intentionally. If the suspected cause is right, the test failure is honest. FreezeUnit returned success for a unit systemd then reports as running, and waiting or retrying in the test would hide it. The fix would belong in systemd, and the added output is what a report there needs.

Verified: go vet, and the test passing in an ubuntu:24.04 container running systemd 255.4 set up the way scripts/ci-runner.sh does; the failure path checked once by forcing the assertion.

TestFreezer asserts systemd's FreezerState right after FreezeUnit and
ThawUnit return. When that assertion fails, the message now includes the
unit's cgroup.freeze and cgroup.events, so a FreezerState that disagrees
with the kernel can be told apart from a freeze that never happened.
The ThawUnit message said "not frozen" where it meant "not running".

Assisted-by: Claude:claude-fable-5-1 [claude-code]
Signed-off-by: Andrew McCabe <amccabe@users.noreply.github.com>
@amccabe

amccabe commented Sep 21, 2026 •

Copy link
Copy Markdown
Author

After some further investigation, it looks like it was preexisting to v255, but it was address in v257 and then differently in v258 (and the v258 change was backported to v56). But the logging change should just be benign and helpful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant