[agentbricks] Delete runtime store before role in tool-matrix cleanup - #688
Open
elainewang-db wants to merge 1 commit into
Open
elainewang-db wants to merge 1 commit into
elainewang-db wants to merge 1 commit into
Conversation
Also surface failed cleanup entries in the evidence gate output. Co-authored-by: Isaac <no-reply@databricks.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Two changes to the Agent Bricks e2e tool-matrix cleanup (
integrations/agentbricks/tests/e2e/tool_matrix.py):cleanup()tore down each app with rawdatabricks apps delete, which leaves the deploy-created managed runtime storeruntime-stores/{app}and its dedicated Lakebase DB. The subsequentdatabricks postgres delete-rolethen fails because the App SP role still owns that DB. Now, for each app with stores, the harness issuesdatabricks api delete /api/2.0/agents/runtime-stores/{app}(404-tolerant) before the role delete, matching the product teardown order indeployments delete(verified incli/deploy.py+ agentkit_api_client.delete_runtime_store). The role delete only runs once the store delete succeeds.verify_evidence's gate printed only a generic one-liner; it now prints each failed cleanup entry (resource+detail) and flags any app deleted without an absence confirmation, so cleanup failures name themselves in CI (whereevidence.jsonisn't uploaded).Why
On the latest nightly dispatch the matrix was fully green (24 passed, 0 failed) but the job exited 1 solely on the cleanup gate. The failing cleanup resource was not visible in CI output. The leading suspect — grounded in the code path and a previously-seen gotcha — is the Lakebase role delete failing on the leftover runtime-store DB; this PR fixes that teardown ordering and adds the diagnostics so the next run confirms the exact failing entry (rather than guessing again).
Test
uv run --group tests pytest tests/unit_tests -k "cleanup or tool_matrix or evidence or author"→ 82 passed. Addedtool_matrix_cleanup_test.pycases: runtime-store DELETE in the expected command sequence; store-delete-fails → role skipped + gate trips; 404 → treated as deleted + role proceeds.Live-validated: dispatched the Agent Bricks integration suite against this PR's head → the
agentbricks-testsjob passed — the first fully-green end-to-end run (all 24 matrix rows pass and the cleanup gate now clears). This confirms the leftover runtime-store DB blocking the role delete was indeed the cause. Run: https://github.com/databricks-eng/ai-oss-integration-tests-runner/actions/runs/37081527823This pull request and its description were written by Isaac.