Repository navigation
Conversation
Upfront partitioning tells agents what to read, but nothing checks what they actually read. Parse agent session logs for Read tool calls and a deterministic subset of shell reads, merge line ranges per file and agent, and report coverage against an optional target list. Gaps can be emitted as a follow-up prompt, and --fail-under gives CI an exit code. Unparsed read-like shell commands are counted, never guessed, so coverage is a lower bound. Closes protocol-security#65. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
|
||
| # Post-hoc read coverage for swarm sessions. | ||
| # | ||
| # Parses agent JSONL session logs (Claude Code `stream-json`), |
There was a problem hiding this comment.
Interesting! Is this 'driver agnostic' or works just with Claude?
The reason I am asking is because I am planning to take claude, gemini, codex, pi, all out of swarm-core and instead maintain them in swarm-drivers (a hypothetical, new repo).
Having coverage.sh tightly coupled with a specific driver can be fine, but then it means we have to take this into account when moving stuff into swarm-drivers.
From a quick scan it seems this works well with Claude Code but has to be adjusted to work with other CLIs.
There was a problem hiding this comment.
This is Claude-only. Making it driver agnostic is a small change: move the log parsing into a per-driver function (like agent_activity_jq) and keep the rest in core.
Codex will need more work, since every read goes through the shell.
If you want this, I'll turn this PR into a draft and work on it, or you can merge this as Claude-only and I can work on others in another PR; your choice.
Sample Gemini/Codex session logs would help me test it against real output.
Summary
coverage.sh, which reports which files and line ranges agents actually read, parsed from their session logs (closes Post-hoc coverage analysis: identify what agents actually read vs. what missed #65).Readtool calls and a deterministic subset of shell reads; other read-like commands are counted as unresolved, never guessed, so coverage is a lower bound.--targetsreports never-read and partially read files,--prompt-outturns the gaps into a follow-up prompt, and--fail-undergives CI an exit code.--logs, with table or--jsonoutput./bin/bash3.2 (no associative arrays, BSDwcpadding stripped).Changes
coverage.sh: CLI entry point; collects agent logs from containers or--logs, renders table/JSON, writes the gap prompt, enforces--fail-under.lib/coverage.sh: log parsing (jq over Claude Codestream-json), EOF/tail range resolution against file sizes, per-file range merging, target pathspec expansion.tests/test_coverage.sh: 24 unit tests covering Read metadata vs. input, the shell-read subset, unresolved counting, merging, targets, and the prompt and exit-code paths..github/workflows/ci.yml: addcoverage.shto the shellcheck file list.USAGE.md: new "Coverage analysis" section; add the test to the unit test list.CHANGELOG.md: entry under Unreleased.Test plan
./tests/test.sh --unit: 14/14 files, 1389 tests pass locally (macOS)./bin/bash tests/test_coverage.shpasses under bash 3.2 and bash 5.shellcheck -s bash --severity=warning coverage.sh lib/coverage.sh tests/test_coverage.shis clean../coverage.shprints a per-file table;./coverage.sh --jsonemits JSON../coverage.sh --targets targets.txt --prompt-out /tmp/followup.md --fail-under 80lists gaps, writes the prompt, and exits 2 when a target is below 80%.The last two are unchecked because Docker wasn't available locally, so they couldn't be run against real swarm containers.
Co-Authored-By: Claude Opus 5.5 noreply@anthropic.com