Skip to content

[agentbricks] Add project status, safe cleanup, and evaluation starters - #683

Closed
shivam5 wants to merge 3 commits into
codex/bugbash-local-statefrom
codex/bugbash-project-lifecycle
Closed

shivam5 wants to merge 3 commits into
codex/bugbash-local-statefrom
codex/bugbash-project-lifecycle

Conversation

@shivam5

@shivam5 shivam5 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #682, which follows #681. This diff contains only project lifecycle/evaluation work. Merge the preceding PRs, then retarget this PR to main.

Users have no project-wide view of configured resources, no ownership-aware cleanup, and no runnable evaluation starter. This adds agentbricks status, conservative agentbricks cleanup, and an extendable MLflow evaluation starter for the OpenAI Agents SDK and LangGraph managed-runtime templates.

Addresses ML-70357.

What changes

  • Status: an offline inventory of the selected profile/workspace, framework, bindings and provisioning receipts. --verify performs read-only checks, reports identifiers/available URLs and individual errors, and distinguishes configuration from resources the caller can read. It never provisions resources; successful reads do not claim that the App can invoke every model/tool.
  • Cleanup: preview by default, explicit --apply and confirmation (or --yes). Deploy records workspace-scoped creation/reuse receipts as provisioning succeeds. Cleanup verifies the current App service-principal identity against its creation receipt, validates managed Runtime Store ownership, deletes the Runtime Store before the App, and persists each outcome for safe partial retries. Failed Runtime Store deletion retains the App.
  • Evaluations: generated evals/cases.jsonl, evals/run.py and extension instructions call the actual project's /api/invocations route, preserving its model, instructions and tools. Each case gets a fresh session. MLflow records input, final assistant output, errors, scorer feedback and aggregate metrics; failed invocations/expectations exit nonzero. The copied standalone runner works without importing a newly added module from an older released runtime. Both OpenAI role=assistant and LangGraph type=ai outputs are supported.
  • UI overview scope: docs/project-overview-design.md specifies resource, cost, evaluation and deployed-version data sources, unavailable states, authorization requirements and an initial resources-and-links slice. This PR supplies the design; it does not ship those four UI panels.

Feedback addressed

Source Exact section Change
Chang Shi Lim — Agentbricks_CLI_Feedback Wish List, item 1: overall project status/state Read-only project inventory and explicit verification
Chang Shi Lim — Agentbricks_CLI_Feedback Wish List, item 2: cleanup Creation receipts, preview/confirmation and identity-checked cleanup
Fabian Nobis — Production Grade Document Chatbot Paragraph beginning “One thing I’m missing is scaffolding for an eval set” Runnable dataset, real agent invocations and MLflow results
Chang Shi Lim — Agentbricks_CLI_Feedback Wish List, item 3: UGW Resources, Cost, Evals, Version Bounded overview design with truthful data sources and dependencies

Before / after evidence

These are lifecycle enhancements; the “before” state is missing functionality, not a claim that all resource operations were broken.

Scenario Before (334da797) After
status --source <project> Exit 1: unknown command status Exit 0 with versioned JSON inventory; no auth initialization or file changes without --verify
cleanup --source <project> Exit 1: unknown command cleanup Exit 0 with App/Runtime Store deletion candidates and explicit retained resources; no deletion during preview
Live workspace status No consolidated command App, Memory Store, Session Store and experiment all reported accessible; App state RUNNING; recorded created resources distinguished from unknown/adopted tracing ownership
Cancel live cleanup No project cleanup workflow Entering n exits 1 with Aborted!; the temporary App remained callable afterward
Evaluation starter No generated dataset/runner Fresh generated entrypoints passed through both deployed frameworks: OpenAI 2/2 starter + 3/3 extended, LangGraph 2/2 starter + 3/3 extended, with local MLflow results
Live negative evaluation No runnable starter A successful real-agent answer with an intentionally wrong expectation produced FAIL / exit 1; the corrected extended dataset then produced PASS / exit 0
Evaluation failure and extension No runnable starter to extend Real local HTTP runtime + MLflow: an intentionally unmet expectation scored 0%, then an extended three-case dataset scored 100%; persisted traces and both scorer assessments verified; exactly four calls for four cases

The final live evaluation checks executed freshly generated evals/run.py entrypoints against both deployed agents pinned to combined SHA 4389014570af6a6b14ca70bd0bbe51dd6066c772. Both copied scripts were byte-for-byte matches for the published evaluator source. They ran directly with the existing framework dependency environment, without a PYTHONPATH override or an import of the new evaluator module. Agent/model calls were real; evaluation results were stored in local SQLite MLflow databases. Saved per-case traces and both scorer assessments were verified for all 11 cases across the five runs.

Validation

  • 203 targeted tests passed across deploy, init, status/cleanup safety, evaluations, CLI/help and Runtime Store ownership tests.
  • Changed Python files pass Ruff and ty; git diff --check passes.
  • tests/e2e/evaluation_smoke.py passed against a real loopback DurableAgentServer and local SQLite MLflow backend. Its deterministic local agent verifies the evaluator/runtime/MLflow path; it is separate from the real-model workspace evaluation above.
  • Live checks used temporary resources in an explicitly selected staging workspace. No pre-existing resources were deleted for validation.
  • Real deployed evaluation results: OpenAI starter run 231fe82ac7b540e9b8270fb02aed41be and extended run cc64b2efb601451b8358614358017eee; LangGraph starter run 7bf4a0090f034153bfc9ae3626ce4dbb and extended run 2852c08ceddd4a50b3cbc85456d0d24e. Intentional negative run 2ebe827aa0d94a32bda26c91871cc6bf preserved a successful invocation and failed expectation separately. These run IDs refer to the local test databases, not workspace-hosted MLflow runs.

Reproduce the targeted tests from integrations/agentbricks:

.venv/bin/pytest tests/unit_tests/{deploy,init,project_lifecycle,evaluation,cli,cli_ergonomics,lakebase_runtime_store}_test.py -q --disable-warnings --maxfail=2

With full MLflow installed (included by either framework extra), reproduce the local end-to-end test:

PYTHONPATH=src python tests/e2e/evaluation_smoke.py --output /tmp/agentbricks-eval-smoke

From a generated project, evaluate a live deployment and inspect its results:

uv run python evals/run.py --app <temporary-test-app> --profile <oauth-profile>
# Add the README's third arithmetic example to a copy of evals/cases.jsonl:
uv run python evals/run.py --app <temporary-test-app> --profile <oauth-profile> --data evals/extended.jsonl
uv run mlflow ui --backend-store-uri sqlite:///.agentbricks/evaluations.db

Cleanup boundary and limitations

Cleanup always retains shared-capable Memory/Session Stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files—even if this project originally created them. Current APIs cannot prove those resources have no other consumers. Adopted Apps, Apps without a creation receipt and Apps whose identity changed are also retained. If an App is already absent, any residual Runtime Store is retained for manual owner inspection. Naming similarity is never ownership evidence.

Keep .agentbricks/resources.json to retain provisioning receipts. The starter's deterministic checks are smoke tests, not a production-quality evaluation benchmark. Custom HTTP-server templates require an evaluator matching their own contract.

Final verification

  • PR head: 8834dc93. Its complete tree is byte-identical to combined test commit 143365d0; application source matches deployed commit 4389014570af6a6b14ca70bd0bbe51dd6066c772 (the subsequent commit only fixes test loading/types).
  • 1,538 unit tests passed, 14 skipped; six fresh-scaffold functional checks, 25 UI tests per framework, Ruff/format/type checks, and five browser e2e runs passed for the combined tree.
  • Fresh generated evaluation entrypoints passed starter and extended datasets on both fixed deployments in ml-inference-staging; an intentional failed expectation exited 1 and persisted the failing assessment.
  • Live cleanup on both task-created Apps deleted the App and its owner-validated Runtime Store. Follow-up GETs returned NOT_FOUND for all four resources. Repeated cleanup exited 0 and preserved the recorded outcome.
  • Read-only checks after cleanup confirmed all four Memory/Session Stores, both experiments and both synced source folders remained accessible. This verifies the retention boundary on real resources, in addition to unit coverage of adopted/recreated Apps and partial failures.

@shivam5

shivam5 commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #687, an independent 22-line documentation change targeting main. It explains existing inspection and cleanup commands. New status/cleanup infrastructure, evaluation scaffolding, and overview UI are deferred. The original branch and commits are preserved.

@shivam5 shivam5 closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant