Conversation
TestFreezer asserts systemd's FreezerState right after FreezeUnit and ThawUnit return. When that assertion fails, the message now includes the unit's cgroup.freeze and cgroup.events, so a FreezerState that disagrees with the kernel can be told apart from a freeze that never happened. The ThawUnit message said "not frozen" where it meant "not running". Assisted-by: Claude:claude-fable-5-1 [claude-code] Signed-off-by: Andrew McCabe <amccabe@users.noreply.github.com>
Author
|
After some further investigation, it looks like it was preexisting to v255, but it was address in v257 and then differently in v258 (and the v258 change was backported to v56). But the logging change should just be benign and helpful. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Three out of fifty-four recent runs of TestFreezer on ubuntu:24.04 (one on each of #528, #529, and #530) failed with
All three ran on 2026-09-18, all three on runners the job log places in Azure centralus, which was the only region with a TestFreezer failure that day. Other centralus jobs in the same hour passed, so while it appears to be regional, there is another factor in the failure.
Broken down by distro and systemd version, across the same runs (ubuntu:22.04 and debian:bookworm from #530's branch, the rest from the matrix on main):
So it appeared to be specific to 255 in centralus. The three failing jobs ran the same container systemd build as August's passes, and the same runner image and base image as the twelve ubuntu:24.04 jobs that passed the same day. The failing tests completed faster than passing ones (0.01 to 0.02 s against 0.03 to 0.05 s), so the reply came back early rather than the system being slow.
Suspected cause
Reading upstream systemd v255's freezer handling:
FreezeUnitwrites the cgroup'scgroup.freeze, marks the unit freezing, and defers its reply until a notification oncgroup.eventsmakes it re-read that file. The handler treatsfrozen 0read while the unit is freezing as the unit having thawed, replies success, and sets the state to running; a laterfrozen 1is then ignored as not initiated by systemd. That is the only path in v255 to a successful reply withFreezerState=running. The notification that triggers the premature read is most likely thepopulatedchange from the service having just started, which the kernel delivers through a workqueue; if that delivery is delayed past the D-Bus round trip, for example by a vCPU stall on a busy host, it lands after the freeze has been written and before the sleeping task has taken its freezer trap. I was unable to reproduce it locally in about 3,000 iterations under CPU, steal, and notification load, which fits a trigger that lives on the host.What this change does
The assertions are unchanged. Both failure messages now include the kernel's view of the unit's cgroup,
cgroup.freezeandcgroup.events, resolved from the unit'sControlGroupproperty. With this change, the next failure would read something likeThat logging indicates
cgroup.freeze:1, which would mean systemd wrote the freeze and its bookkeeping disagrees with the kernel, which I'd expect from these failed tests.cgroup.freeze:0would mean the freeze was never written or that it was reverted (some other cause). TheThawUnitmessage also said "not frozen" where it meant "not running".This is extra diagnosis rather than a test fix intentionally. If the suspected cause is right, the test failure is honest.
FreezeUnitreturned success for a unit systemd then reports as running, and waiting or retrying in the test would hide it. The fix would belong in systemd, and the added output is what a report there needs.Verified:
go vet, and the test passing in an ubuntu:24.04 container running systemd 255.4 set up the wayscripts/ci-runner.shdoes; the failure path checked once by forcing the assertion.