Skip to content

Commit b5a6d53

Browse files
authored
fix(mothership): run same-workflow workflow tools in order instead of failing as busy (#8709)
* fix(mothership): run same-workflow workflow tools in order instead of failing as busy When the agent fans out run_block/run_workflow/run_from_block/run_workflow_until_block calls on one workflow in a single step, the browser executor rejected every call after the first with WORKFLOW_EXECUTION_BUSY, because the editor keeps one visible execution per workflow. The agent then spent a round re-issuing them one by one. That was about 46% of run_block failures. Calls behind a run tool already running that workflow in this tab now wait in arrival order and launch when it releases. Stop cancels waiting calls without launching them, a chat recovery treats a waiting call as owned by this tab, and a run the agent did not start (the user's manual run) is still reported as busy. The wait needs no timeout of its own: a call nobody claims within the server's pickup grace is run by the server, and the late launch here gets the execute route's benign 409. * fix(mothership): keep a workflow for a run whose stream dropped until it settles A run tool whose execute stream is interrupted leaves the run executing server-side and keeps its terminal pointer so the editor can re-attach. Admitting the next waiting call right away overwrote that pointer and claimed the workflow, so reconnect skipped it and the still-running execution lost its live output and the editor's Stop. The interrupted execution now holds the workflow until the server reports it settled and any reconnect has drained it. The hold is not run-tool ownership, so reconnect can still claim the pointer. Waiting calls start after it; one still waiting past the server's pickup grace is run by the server, and its late start here gets the benign 409. * fix(mothership): release a dropped run's hold once it settles, even after an abandoned reconnect The hold waited for the execution store to stop showing the interrupted execution as current. Navigating away cancels the editor's reconnect without clearing that, so a settled run kept later calls waiting until the server's pickup fallback ran them. The hold now waits for the server to settle the execution and for no reconnect stream to be open for it, then clears the stale visible state it left behind before admitting the next call. * fix(mothership): re-send an undelivered workflow tool completion while another run owns the workflow A run whose completion report failed keeps it in session storage for recovery to re-send. Recovery refused to touch the call whenever another run tool owned the workflow, which the queue now makes common, so the stored report was never delivered. Re-sending a stored report needs nothing from the workflow; recovery now does it first, and only skips the pointer lookup and cleanup that belong to the run that owns the workflow. Also makes the abandoned-reconnect test track the running flag, so it fails without the stale-state clearing it covers. * improvement(mothership): keep each workflow's run-tool slot in one record and poll a dropped run only when needed Replaces the three parallel maps (owning run tool, waiters, interrupted execution) with one per-workflow record whose owner is a run tool or an interrupted execution. Acquire, admit and release are its only mutators, and every run tool releases through one path. A dropped run is now checked only while a call is actually waiting behind it, never while the editor's reconnect stream is following it or the tab is hidden, stops fetching once the server reports it settled, and gives up the hold after repeated status failures. Also extracts the visible-execution reset, makes waiter removal a single pass and moves the design notes to the PR description. * fix(mothership): clear a dropped run's saved pointer when its hold is released A reconnect cancelled by navigating away leaves the settled run's execution pointer saved. Releasing the hold now clears it when it still names that execution, before the next run is admitted, so recovery cannot later report the finished run as a live one.
1 parent ab6b400 commit b5a6d53

3 files changed

Lines changed: 612 additions & 80 deletions

File tree

‎apps/sim/hooks/use-execution-stream.ts‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -285,6 +285,11 @@ function reconnectStreamKey(workflowId: string, executionId: string): string {
285285
return `${workflowId}:reconnect:${executionId}`
286286
}
287287

288+
/** Whether a reconnect stream for this execution is currently open in this tab. */
289+
export function isReconnectStreamOpen(workflowId: string, executionId: string): boolean {
290+
return sharedAbortControllers.has(reconnectStreamKey(workflowId, executionId))
291+
}
292+
288293
function abortStream(key: string): void {
289294
const controller = sharedAbortControllers.get(key)
290295
if (!controller) return

0 commit comments

Comments
 (0)