Skip to content

Add pack-audit: readiness scanner, rubric, and auditor agent - #50

Draft
brandonrc wants to merge 5 commits into
nebari-dev:mainfrom
brandonrc:pack-audit-tool
Draft

brandonrc wants to merge 5 commits into
nebari-dev:mainfrom
brandonrc:pack-audit-tool

Conversation

@brandonrc

@brandonrc brandonrc commented Oct 9, 2026 •

Copy link
Copy Markdown

Closes #49.

What this adds

  • tools/pack-audit/audit.py: scans a pack repo, renders the chart with the NebariApp toggle on and off, inspects manifests (probes, securityContext, image pinning, NetworkPolicy, scrape annotations, literal secrets), validates the NebariApp spec against the operator CRD and pack-metadata.yaml against the dashboard schema, runs kubeconform (or kubernetes-validate) over the render, greps docs and CI for the checklist's requirements, and writes reports/<pack>.{json,md} with a weighted score, achieved maturity level, and blockers per level. report rescores after judgment edits; summary compares several packs; summary_html.py renders a shareable page.
  • tools/pack-audit/checklist.yaml: docs/release-readiness-checklist.md as data. Each item carries level, category, check type (auto|judgment|manual), applicability, expectation, and source/example links to the rule and to a first-party pack file that satisfies it. One auditor-added item (NA-07, NebariApp spec validates and resolves) predicts the cluster-only NA-01.
  • .claude/agents/software-pack-auditor.md: Claude Code agent that runs the scanner, reads the pack, decides judgment items with cited evidence, corrects scanner heuristics, rescores, and writes a verdict and fix list. Read-only on the pack by default; fix mode only after explicit permission.
  • tools/pack-audit/reference/template-conventions.md: what a healthy pack looks like, distilled from this template and the first-party packs.
  • docs/superpowers/specs/2026-10-09-pack-audit-design.md: design notes and decisions (gate-based scoring, excluded process rows, telemetry judged against the NIC OTel collector's discovery mechanism, read-only audit).

Runs on Nebari (three states)

The score does not say whether a pack will run on Nebari, so each report also answers that separately: verified (installed on a cluster with the operator and the NebariApp reached Ready), likely (static pre-check: CRD-only fields, explicit routing, resolvable Service), or no. audit.py verify moves a pack from likely to verified by reusing dev/Makefile's cluster target, installing with the NebariApp enabled at <pack>.nebari.local, waiting for Ready, probing the hostname through the gateway, and writing NA-01/NA-04 back with the operator's condition reasons. --dry-run prints the plan. It has not yet been exercised end to end on a kind cluster from this tool (no Docker access on the authoring machine); the dry-run plan has been checked against the Makefile and the operator's dev scripts.

Export for other tools

audit.py export --format sarif|junit|csv renders the same results JSON as SARIF 2.1.0 (rule id = checklist id, helpUri = the rule link; for GitHub code scanning and VS Code), JUnit XML (one suite per maturity level; for GitLab/Jenkins/Azure test reports), or CSV.

Scoring in one paragraph

Items are weighted by the level at which they block promotion (E=4, A=3, B=2, GA=1). PASS=1, PARTIAL=0.5, FAIL=0. MANUAL (cluster/person) and NA items are excluded from the denominator; pre-sales and sign-off rows are listed but never scored. The repo-readiness level is the highest level with no FAIL/PARTIAL at or below it.

Validation

  • Self-test from this branch: python3 -I tools/pack-audit/audit.py scan examples/auth-fastapi and examples/wrap-existing-chart both render, lint, and package; wrap-existing-chart fails NA-07 because its values.yaml omits routing (see the issue for why that matters on operator v0.1.1).
  • Calibrated on nebari-dev/data-science-pack (declared beta): 78.7%, Experimental by the letter of the rubric; remaining gaps are mostly where content lives (README headings, examples/), plus a stale appVersion and no metrics story.
  • Used on four packs built outside the org with very different shapes (a GitLab-hosted generated chart, a multi-chart Istio/KServe bundle, an EKS/ALB chart with no NebariApp, a generated GPU inference pack). The heuristics that misfired on those and on the baseline are fixed in this version.

Requirements

Python 3.10+ with PyYAML (plus jsonschema for schema validation of pack-metadata.yaml); helm 3.8+ on PATH; optional kubeconform or pip install kubernetes-validate.

Not in this PR

  • Wiring the self-test into lint.yaml (open question in the issue).
  • The two example fixes and the CRD reference refresh noted in the issue; they are independent of this tool.

Add tools/pack-audit/ with a deterministic scanner and scorer (audit.py),
the release-readiness checklist encoded as data (checklist.yaml, with a
source and first-party example link per item), a cross-pack HTML summary,
and a vendored pack-metadata schema. Add a Claude Code agent
(.claude/agents/software-pack-auditor.md) that runs the scanner, decides
the items that need a reader, rescores, and writes a verdict with a
file-level fix list. Audit mode is read-only on the pack; fix mode is
opt-in. The design spec explains how the rubric, scoring, and agent were
derived from this template, the first-party packs, and nebari-operator
v0.1.1.

Refs nebari-dev#49
…/CSV export

The score alone does not say whether a pack will run on Nebari, so every
report now carries a three-state answer: verified (installed on a cluster
with the operator and the NebariApp reached Ready), likely (static CRD
pre-check passes), or no. `audit.py verify` moves packs from likely to
verified by reusing dev/Makefile's cluster target, installing the pack
with the NebariApp enabled, waiting for Ready, probing the hostname
through the gateway, and writing NA-01/NA-04 back with the operator's
condition reasons. `audit.py export` renders the same results as SARIF
2.1.0, JUnit XML, or CSV for code scanning and CI test reports. `report`
now preserves an auditor narrative placed above the generated section.

Refs nebari-dev#49
The operator attaches the Keycloak groups client scope and
group-membership mapper only when the NebariApp lists `groups` in
auth.scopes. A NebariApp that sets auth.groups without requesting the
scope gets a deny-by-default SecurityPolicy on a claim the token never
carries, so every login is rejected. Seen on a live NIC cluster; the
scanner now reports it as a blocking NebariApp spec issue (NA-07), and the
conventions doc shows the correct scope list.

Refs nebari-dev#49

@oldsj oldsj left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Really useful work. The checklist as data, deterministic scoring with the agent only deciding judgment items, not scoring cluster/person items, and SARIF/JUnit export are all the right calls. Feedback below, roughly in priority order.

1. Lean on existing scanners for generic Kubernetes hygiene. kube-linter, kube-score and Trivy already cover non-root, image pinning, privileged containers, hostNetwork, hostPath, resource requests, RBAC scope, and "every pod is selected by a NetworkPolicy", and they keep adding rules. Suggest audit.py shells out to one of them and maps its findings onto checklist IDs, keeping its own code for what only Nebari can check: NebariApp, pack-metadata.yaml, docs, and release conventions. That would shrink audit.py considerably and keep the generic checks current. chart-testing (ct) is the usual tool for the lint/install/upgrade items.

2. SEC-06 is easy to pass. It passes if any NetworkPolicy renders, so a single allow-all policy scores PASS. A better bar: every workload is selected by a policy and policies set policyTypes explicitly, with a warning for namespace-wide or all-port rules (the vista report notes its ingress admits all of envoy-gateway-system, which this would flag).

3. Checks not covered today. Whether these come from an external scanner per point 1 or from audit.py itself, a vendor-pack audit should cover: privileged containers, hostNetwork, hostPath, added capabilities, ClusterRole/ClusterRoleBinding and CRD installs, and automountServiceAccountToken. Missing resource requests are only an observation right now; on a shared cluster they're worth scoring (Beta seems right).

4. Checklist wording (upstream doc).

  • SEC-06 at GA feels late. On a strict-mode cluster (default-deny until a policy allows traffic) a pack without policies doesn't run at all, so it's closer to an Alpha/Beta concern.
  • +1 to rewording the Telemetry row to "scrape annotations or OTLP".
  • +1 to giving "deploys without modification" a testable definition.

5. "Verified" before verify has run. Since audit.py verify hasn't been exercised end to end yet, either one run on kind before merge, or label it experimental in the README so the "verified" state isn't over-trusted.

6. Self-test in CI. Yes to the open question: scanning examples/* in lint.yaml keeps the scanner and checklist from drifting apart.

7. Reach. Living in tools/ here works for packs built from the template, but vendor packs are the main audience and rarely are. A reusable GitHub Action or a published CLI would let any pack repo run it.

8. Site profiles. Could checklist.yaml support loading extra checklists (--profile <site>), scored separately from maturity? Deployment sites often have hard requirements beyond maturity: strict NetworkPolicy shapes, how workloads get cloud credentials, custom CA trust, reachable registries. A profile hook lets a site add those without forking the tool. We'd use it for our ATEP requirements.

9. PR size. At ~3.3k lines across nine files, splitting the scanner and checklist from the agent and HTML summary would make it easier to review and land.

🤖 claude-opus-5-5 (medium) · reviewed by @oldsj

…l on a CI file

candidate_values_files now ranks examples/*nebari* ahead of dev/local
profiles, so EX-01 is judged on the file installers are told to use.
OI-01 no longer lists a missing CI workflow as a template-layout gap; CI
wiring is scored by IN-06, IN-07 and RE-07 and noted as an observation.
A values file that the NebariApp-enabled render needed is a Nebari
example (PARTIAL) even when it is not named nebari-values.yaml; PASS
still requires the conventional name.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a readiness auditor (tools/pack-audit + Claude Code agent) that scores packs against the release-readiness checklist

2 participants