Conversation
…ifecycle # Conflicts: # integrations/agentbricks/src/databricks_agentbricks/cli/init.py
Contributor
Author
|
Superseded by #687, an independent 22-line documentation change targeting main. It explains existing inspection and cleanup commands. New status/cleanup infrastructure, evaluation scaffolding, and overview UI are deferred. The original branch and commits are preserved. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #682, which follows #681. This diff contains only project lifecycle/evaluation work. Merge the preceding PRs, then retarget this PR to
main.Users have no project-wide view of configured resources, no ownership-aware cleanup, and no runnable evaluation starter. This adds
agentbricks status, conservativeagentbricks cleanup, and an extendable MLflow evaluation starter for the OpenAI Agents SDK and LangGraph managed-runtime templates.Addresses ML-70357.
What changes
--verifyperforms read-only checks, reports identifiers/available URLs and individual errors, and distinguishes configuration from resources the caller can read. It never provisions resources; successful reads do not claim that the App can invoke every model/tool.--applyand confirmation (or--yes). Deploy records workspace-scoped creation/reuse receipts as provisioning succeeds. Cleanup verifies the current App service-principal identity against its creation receipt, validates managed Runtime Store ownership, deletes the Runtime Store before the App, and persists each outcome for safe partial retries. Failed Runtime Store deletion retains the App.evals/cases.jsonl,evals/run.pyand extension instructions call the actual project's/api/invocationsroute, preserving its model, instructions and tools. Each case gets a fresh session. MLflow records input, final assistant output, errors, scorer feedback and aggregate metrics; failed invocations/expectations exit nonzero. The copied standalone runner works without importing a newly added module from an older released runtime. Both OpenAIrole=assistantand LangGraphtype=aioutputs are supported.docs/project-overview-design.mdspecifies resource, cost, evaluation and deployed-version data sources, unavailable states, authorization requirements and an initial resources-and-links slice. This PR supplies the design; it does not ship those four UI panels.Feedback addressed
Before / after evidence
These are lifecycle enhancements; the “before” state is missing functionality, not a claim that all resource operations were broken.
334da797)status --source <project>unknown command status--verifycleanup --source <project>unknown command cleanupRUNNING; recorded created resources distinguished from unknown/adopted tracing ownershipnexits 1 withAborted!; the temporary App remained callable afterwardThe final live evaluation checks executed freshly generated
evals/run.pyentrypoints against both deployed agents pinned to combined SHA4389014570af6a6b14ca70bd0bbe51dd6066c772. Both copied scripts were byte-for-byte matches for the published evaluator source. They ran directly with the existing framework dependency environment, without aPYTHONPATHoverride or an import of the new evaluator module. Agent/model calls were real; evaluation results were stored in local SQLite MLflow databases. Saved per-case traces and both scorer assessments were verified for all 11 cases across the five runs.Validation
git diff --checkpasses.tests/e2e/evaluation_smoke.pypassed against a real loopbackDurableAgentServerand local SQLite MLflow backend. Its deterministic local agent verifies the evaluator/runtime/MLflow path; it is separate from the real-model workspace evaluation above.231fe82ac7b540e9b8270fb02aed41beand extended runcc64b2efb601451b8358614358017eee; LangGraph starter run7bf4a0090f034153bfc9ae3626ce4dbband extended run2852c08ceddd4a50b3cbc85456d0d24e. Intentional negative run2ebe827aa0d94a32bda26c91871cc6bfpreserved a successful invocation and failed expectation separately. These run IDs refer to the local test databases, not workspace-hosted MLflow runs.Reproduce the targeted tests from
integrations/agentbricks:.venv/bin/pytest tests/unit_tests/{deploy,init,project_lifecycle,evaluation,cli,cli_ergonomics,lakebase_runtime_store}_test.py -q --disable-warnings --maxfail=2With full MLflow installed (included by either framework extra), reproduce the local end-to-end test:
From a generated project, evaluate a live deployment and inspect its results:
Cleanup boundary and limitations
Cleanup always retains shared-capable Memory/Session Stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files—even if this project originally created them. Current APIs cannot prove those resources have no other consumers. Adopted Apps, Apps without a creation receipt and Apps whose identity changed are also retained. If an App is already absent, any residual Runtime Store is retained for manual owner inspection. Naming similarity is never ownership evidence.
Keep
.agentbricks/resources.jsonto retain provisioning receipts. The starter's deterministic checks are smoke tests, not a production-quality evaluation benchmark. Custom HTTP-server templates require an evaluator matching their own contract.Final verification
8834dc93. Its complete tree is byte-identical to combined test commit143365d0; application source matches deployed commit4389014570af6a6b14ca70bd0bbe51dd6066c772(the subsequent commit only fixes test loading/types).