Repository navigation
Conversation
Add tools/pack-audit/ with a deterministic scanner and scorer (audit.py), the release-readiness checklist encoded as data (checklist.yaml, with a source and first-party example link per item), a cross-pack HTML summary, and a vendored pack-metadata schema. Add a Claude Code agent (.claude/agents/software-pack-auditor.md) that runs the scanner, decides the items that need a reader, rescores, and writes a verdict with a file-level fix list. Audit mode is read-only on the pack; fix mode is opt-in. The design spec explains how the rubric, scoring, and agent were derived from this template, the first-party packs, and nebari-operator v0.1.1. Refs nebari-dev#49
…/CSV export The score alone does not say whether a pack will run on Nebari, so every report now carries a three-state answer: verified (installed on a cluster with the operator and the NebariApp reached Ready), likely (static CRD pre-check passes), or no. `audit.py verify` moves packs from likely to verified by reusing dev/Makefile's cluster target, installing the pack with the NebariApp enabled, waiting for Ready, probing the hostname through the gateway, and writing NA-01/NA-04 back with the operator's condition reasons. `audit.py export` renders the same results as SARIF 2.1.0, JUnit XML, or CSV for code scanning and CI test reports. `report` now preserves an auditor narrative placed above the generated section. Refs nebari-dev#49
The operator attaches the Keycloak groups client scope and group-membership mapper only when the NebariApp lists `groups` in auth.scopes. A NebariApp that sets auth.groups without requesting the scope gets a deny-by-default SecurityPolicy on a claim the token never carries, so every login is rejected. Seen on a live NIC cluster; the scanner now reports it as a blocking NebariApp spec issue (NA-07), and the conventions doc shows the correct scope list. Refs nebari-dev#49
oldsj
left a comment
There was a problem hiding this comment.
Really useful work. The checklist as data, deterministic scoring with the agent only deciding judgment items, not scoring cluster/person items, and SARIF/JUnit export are all the right calls. Feedback below, roughly in priority order.
1. Lean on existing scanners for generic Kubernetes hygiene. kube-linter, kube-score and Trivy already cover non-root, image pinning, privileged containers, hostNetwork, hostPath, resource requests, RBAC scope, and "every pod is selected by a NetworkPolicy", and they keep adding rules. Suggest audit.py shells out to one of them and maps its findings onto checklist IDs, keeping its own code for what only Nebari can check: NebariApp, pack-metadata.yaml, docs, and release conventions. That would shrink audit.py considerably and keep the generic checks current. chart-testing (ct) is the usual tool for the lint/install/upgrade items.
2. SEC-06 is easy to pass. It passes if any NetworkPolicy renders, so a single allow-all policy scores PASS. A better bar: every workload is selected by a policy and policies set policyTypes explicitly, with a warning for namespace-wide or all-port rules (the vista report notes its ingress admits all of envoy-gateway-system, which this would flag).
3. Checks not covered today. Whether these come from an external scanner per point 1 or from audit.py itself, a vendor-pack audit should cover: privileged containers, hostNetwork, hostPath, added capabilities, ClusterRole/ClusterRoleBinding and CRD installs, and automountServiceAccountToken. Missing resource requests are only an observation right now; on a shared cluster they're worth scoring (Beta seems right).
4. Checklist wording (upstream doc).
- SEC-06 at GA feels late. On a strict-mode cluster (default-deny until a policy allows traffic) a pack without policies doesn't run at all, so it's closer to an Alpha/Beta concern.
- +1 to rewording the Telemetry row to "scrape annotations or OTLP".
- +1 to giving "deploys without modification" a testable definition.
5. "Verified" before verify has run. Since audit.py verify hasn't been exercised end to end yet, either one run on kind before merge, or label it experimental in the README so the "verified" state isn't over-trusted.
6. Self-test in CI. Yes to the open question: scanning examples/* in lint.yaml keeps the scanner and checklist from drifting apart.
7. Reach. Living in tools/ here works for packs built from the template, but vendor packs are the main audience and rarely are. A reusable GitHub Action or a published CLI would let any pack repo run it.
8. Site profiles. Could checklist.yaml support loading extra checklists (--profile <site>), scored separately from maturity? Deployment sites often have hard requirements beyond maturity: strict NetworkPolicy shapes, how workloads get cloud credentials, custom CA trust, reachable registries. A profile hook lets a site add those without forking the tool. We'd use it for our ATEP requirements.
9. PR size. At ~3.3k lines across nine files, splitting the scanner and checklist from the agent and HTML summary would make it easier to review and land.
🤖 claude-opus-5-5 (medium) · reviewed by @oldsj
…l on a CI file candidate_values_files now ranks examples/*nebari* ahead of dev/local profiles, so EX-01 is judged on the file installers are told to use. OI-01 no longer lists a missing CI workflow as a template-layout gap; CI wiring is scored by IN-06, IN-07 and RE-07 and noted as an observation.
A values file that the NebariApp-enabled render needed is a Nebari example (PARTIAL) even when it is not named nebari-values.yaml; PASS still requires the conventional name.
Closes #49.
What this adds
tools/pack-audit/audit.py: scans a pack repo, renders the chart with the NebariApp toggle on and off, inspects manifests (probes, securityContext, image pinning, NetworkPolicy, scrape annotations, literal secrets), validates the NebariApp spec against the operator CRD andpack-metadata.yamlagainst the dashboard schema, runs kubeconform (orkubernetes-validate) over the render, greps docs and CI for the checklist's requirements, and writesreports/<pack>.{json,md}with a weighted score, achieved maturity level, and blockers per level.reportrescores after judgment edits;summarycompares several packs;summary_html.pyrenders a shareable page.tools/pack-audit/checklist.yaml:docs/release-readiness-checklist.mdas data. Each item carries level, category, check type (auto|judgment|manual), applicability, expectation, andsource/examplelinks to the rule and to a first-party pack file that satisfies it. One auditor-added item (NA-07, NebariApp spec validates and resolves) predicts the cluster-only NA-01..claude/agents/software-pack-auditor.md: Claude Code agent that runs the scanner, reads the pack, decides judgment items with cited evidence, corrects scanner heuristics, rescores, and writes a verdict and fix list. Read-only on the pack by default; fix mode only after explicit permission.tools/pack-audit/reference/template-conventions.md: what a healthy pack looks like, distilled from this template and the first-party packs.docs/superpowers/specs/2026-10-09-pack-audit-design.md: design notes and decisions (gate-based scoring, excluded process rows, telemetry judged against the NIC OTel collector's discovery mechanism, read-only audit).Runs on Nebari (three states)
The score does not say whether a pack will run on Nebari, so each report also answers that separately: verified (installed on a cluster with the operator and the NebariApp reached
Ready), likely (static pre-check: CRD-only fields, explicitrouting, resolvable Service), or no.audit.py verifymoves a pack from likely to verified by reusingdev/Makefile'sclustertarget, installing with the NebariApp enabled at<pack>.nebari.local, waiting forReady, probing the hostname through the gateway, and writing NA-01/NA-04 back with the operator's condition reasons.--dry-runprints the plan. It has not yet been exercised end to end on a kind cluster from this tool (no Docker access on the authoring machine); the dry-run plan has been checked against the Makefile and the operator's dev scripts.Export for other tools
audit.py export --format sarif|junit|csvrenders the same results JSON as SARIF 2.1.0 (rule id = checklist id,helpUri= the rule link; for GitHub code scanning and VS Code), JUnit XML (one suite per maturity level; for GitLab/Jenkins/Azure test reports), or CSV.Scoring in one paragraph
Items are weighted by the level at which they block promotion (E=4, A=3, B=2, GA=1). PASS=1, PARTIAL=0.5, FAIL=0. MANUAL (cluster/person) and NA items are excluded from the denominator; pre-sales and sign-off rows are listed but never scored. The repo-readiness level is the highest level with no FAIL/PARTIAL at or below it.
Validation
python3 -I tools/pack-audit/audit.py scan examples/auth-fastapiandexamples/wrap-existing-chartboth render, lint, and package;wrap-existing-chartfails NA-07 because itsvalues.yamlomitsrouting(see the issue for why that matters on operator v0.1.1).nebari-dev/data-science-pack(declared beta): 78.7%, Experimental by the letter of the rubric; remaining gaps are mostly where content lives (README headings,examples/), plus a staleappVersionand no metrics story.Requirements
Python 3.10+ with PyYAML (plus
jsonschemafor schema validation ofpack-metadata.yaml);helm3.8+ on PATH; optionalkubeconformorpip install kubernetes-validate.Not in this PR
lint.yaml(open question in the issue).