Repository navigation
v0.9.15: mship improvements, memory improvements, nextjs bump, plane and ramp integrations, outbox hardening - #8748
Merged
Merged
Conversation
…ead of extra E2B round trips (#8621) * perf(sandbox): grant the run_code session lease in the reconnect instead of extra E2B round trips Reused Mothership workbench calls made five E2B control-plane requests: list, connect, getInfo + setTimeout on acquire (the set always fired), and getInfo on release. Connect now asks for max(5 min, remaining, lease), the handle records the deadline it requested, and acquisition/release skip the provider while that lower bound covers the request. Final deadlines are unchanged; unrequested deadlines are still read back. * test(sandbox): model connect as setting the deadline so a dropped preserve is caught
The simple Chat effort picker labeled medium as Low, high as Medium, and xhigh as High. Derive its options from MOTHERSHIP_EFFORT_OPTIONS so each label names the effort it sends: Medium, High, Extra High. Values, the default (high), and the stored preference are unchanged.
…un --stop-after` (#8622) * feat(workflows): stop a manual v2 run after a block, and `workflows run --stop-after` A manual v2 run can now name `run.stopAfterBlockId`; the run stops once that block completes and downstream blocks do not execute. Combined with a block entry on the same block, it re-runs exactly one block against a prior run's persisted upstream outputs, server-side: sim workflows run W --from-block X --source-run R --stop-after X --select-output X.result Agents verifying an edit no longer re-run every upstream block (often a slow LLM or API call) or toggle blocks off to skip them. - Contract: optional `stopAfterBlockId` on the manual run selection. - Application: both manual operations refuse a block missing from the saved workflow or nested in a loop/parallel (the engine would otherwise run to the end or stop after one iteration), before anything runs. - Execute service: threads the trusted value to the sync and stream paths. - CLI: `--stop-after <blockId>` implies --manual and rejects --async. - E2E: test-workflow-stop-after-e2e.ts against a running app; the http-e2e job gains a Redis service because hosted billing admits runs through a Redis usage reservation. * fix(workflows): refuse stop targets a run cannot reach, and run the E2E self-hosted - The manual operations refuse a stop block the run cannot reach from its entry (an upstream block would let the run finish everything after the entry), and look blocks up as own properties. - The executor fails a run whose stop block is absent from the workflow it executes, instead of running everything; this closes the window between validation and the executor's own draft load, for every caller. - The CLI refuses an empty --stop-after rather than dropping it. - CI: the stop-after E2E gets its own self-hosted app step; the SCIM suite asserts PostgreSQL rate-limit storage, so Redis is not added to that app. Fixture cleanup waits for run logs to finalize before deleting. * fix(workflows): refuse a disabled stop block, or one reached only through one The executor omits disabled blocks from its graph, so a disabled stop target, or one whose only path runs through a disabled block, is never reached and the run would finish everything after the entry. * fix(workflows): the executor refuses a disabled stop block too The serialized workflow keeps disabled blocks, but the DAG skips them, so a disabled stop target would never be reached. --------- Co-authored-by: Waleed Latif <waleed@sim.ai>
…pletion (#8620) * fix(executor): stop retaining duplicate copies of loop block outputs * test(executor): guard output sharing in block logs and loop aggregates * fix(logs): size execution data the way JSON.stringify writes it * fix(logs): count JSON string bytes without copying and unbox primitive wrappers * fix(logs): measure execution data iteratively and apply toJSON on functions * fix(executor): drop the full JSON clone of execution state at run completion * fix(executor): normalize only live state for PII masking and walk serializability lazily * fix(executor): snapshot array lengths and unbox wrappers in JSON walks * fix(logs): read boxed boolean and bigint values the way JSON.stringify does
* feat(projects): add project identity and lifecycle foundation * feat(projects): create projects with their initial environment * docs(projects): record project files follow-up * docs(projects): explain project and workspace creation flows * fix(projects): stage activation after compatible writers deploy * refactor(projects): prepare compatible writers for the SQL backfill * fix(projects): clean up automatically created fixture Projects * fix(workflows): guard restore against concurrent workspace archive * fix(projects): close lifecycle races and surface rollout conflicts * fix(workflows): return not found when import loses archive race
* fix(mcp): restrict MCP server destination changes to admins * fix(mcp): compare exact MCP paths and guard concurrent URL changes * fix(mcp): treat setting a URL on a URL-less server as a destination change * fix(mcp): guard re-registration against concurrent URL changes * fix(mcp): narrow re-registration URL before the guarded update * chore(mcp): use absolute import in utils test
…v2 provider discovery (#8632)
…ck (#8635) * fix(executor): fail a stop-after run whose routing skips the stop block A run with stopAfterBlockId only stopped when the stop block completed. When a router, condition, or untaken error path routed the run away from it, the stop never triggered and the run finished every other branch, reporting success as if it had stopped there. A static check before the run cannot see this. - The engine ends the run as soon as every path into the stop block has been deactivated, before any further block starts, and fails it with `Stop block "<name>" (<id>) was not reached: no path this run took leads to it`. - Any run that ends without completing its stop block fails the same way: a stop block missing from the executed graph, or a Response block that ended the run first. - A loop or parallel stop with nothing to run completes at its start sentinel, whose end sentinel never runs, so that exit now counts as reaching it. - The v2 contract and the CLI `--stop-after` help describe the failure. - E2E: a condition fixture checks the stop on the taken branch still stops there, a stop on the skipped branch fails the run before the other branch's slow block finishes, and the CLI exits non-zero. * fix(executor): a skipped stop block fails a run another branch paused, and names a Response ending - A run whose stop block was proven unreachable fails even when another branch paused, instead of returning a paused run that would resume past it. - When a Response block ended the run first, the error says so rather than claiming no path leads to the stop block.
…le (#8631) * fix(projects): restore workspace deletion and tighten Project lifecycle - Archive a Project with its last active environment instead of refusing the workspace delete; account deletion follows the same rule, and the implicit archive is audited - Gate Project APIs on a `projects` AppConfig flag (PROJECT_API_ENABLED fallback) - Run Project reads in a read-only snapshot without locks; list Projects from the caller's grants with batched authorization - Batch workflow archival, Project transfer and owner reassignment; move Project ownership on organization ownership transfer - Reuse shared advisory-lock and text-array helpers; narrow admin-move conflict mapping to ProjectConflictError * improvement(projects): unify environment archive and align with shared patterns - Archive a workspace's workflows atomically with it through one archiveEnvironmentInTransaction shared by workspace delete and Project archive; the workspace row is locked before the sweep so concurrent creates are covered - Scope the Project lock timeout to lock acquisition and map lock timeouts and deadlocks to a retryable conflict; backfill and multi-Project locks use the shared advisory-lock helpers in code-unit order - Shared orchestrationFailureResponse for raw routes; contracts use the ID primitives and export only what is consumed; audit enums and mock in sync - Project restrictions section matches its sibling settings rows - Batch account-deletion Project loads/locks; skip inconsistent Projects in lists - Harden the foundation integration suite (user-keyed cleanup, poll helper, pid-scoped waits, precise assertions) * fix(projects): trim the requested organization id before validating it * fix(projects): address review on Project locking, archive notifications and list policy cost * fix(projects): keep archive retries from re-stamping MCP servers and isolate post-commit notifications
…m, restore Low (#8634) * feat(mothership): keep each chat's reasoning effort, default to medium, restore Low The simple picker offers Low / Medium / High / Extra High again, each sending exactly that effort. New chats and chats never changed run at medium instead of high. An effort the user picks is stored on the chat (copilot_chats.config) through a new PUT /api/mothership/chats/[chatId]/effort and at turn admission, and later turns of that chat keep it. The global last-used effort is no longer persisted. The Sim Chat block defaults to medium. * fix(mothership): keep the latest effort pick through refetches, failed saves and abandoned new chats * fix(mothership): leave a deduplicated send's chat on the pick its first attempt stored * fix(mothership): show a recovered chat's pick while its details load * fix(mothership): hand a withdrawn first send's effort pick back to the new-chat composer * fix(mothership): hand a withdrawn send's pick back only while its new-chat surface is open
… Tools Compared (#8638) Co-authored-by: Sim Pi Agent <pi@sim.ai>
… Integrations (#8639) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…8642) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…gine-optimization-actually-mean (#8647) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
#8643) * fix(mothership): refuse desktop claims for unapproved or stopped calls The desktop authorize route claimed any pending call, so a gated terminal run that was still awaiting approval, or a call on a run the user had stopped, could be claimed and executed. The claim now locks the run row and refuses once tool admission has closed (Stop, a newer turn, or the run's end), and refuses a call held for the user's decision until they allow it. Whether a call is gated depends on the turn, so pre-persist records it on the row (permission_requested_at). Stop now settles the stopped runs' open desktop calls in the transaction that closes admission: unclaimed calls as never started, claimed calls as outcome unknown. A result for an already-settled call (a retry, or one that lost to Stop) is acknowledged with the stored outcome instead of 404/500. * fix(mothership): settle every open desktop call on Stop and answer 410 after it Stop also settles delivered desktop calls and reads of granted local folders. Authorize checks admission before the call's status, so a call Stop already settled answers 410 rather than 404. * test(mothership): assert the acknowledged outcome, not the publish mock * test(mothership): give the hand-built tool call table the new permission column * refactor(mothership): one claim primitive, one desktop-tool classifier, sealed Stop results The desktop claim is now an option of the run-locked tool execution claim (claimSimToolExecution becomes claimToolExecution) instead of a second copy of the admission check. Stop picks the open desktop calls with the shared TS classifier, which moves to lib/mothership/tools/desktop-tools.ts, instead of a SQL restatement of it, and seals each result the way the confirm route does, so a waiter restores what Stop did rather than failing to unseal it. * fix(mothership): a declined call stays unclaimable without its gate marker Calls gated before permission_requested_at existed carry no marker, so a recorded decision that does not allow the call now disqualifies it too. * test(mothership): assert authorize outcomes, not claim mock calls
…#8655) Picks queued behind an in-flight save run onMutate at once, so after low, high, low a failed first save dropped the newest low and the picker fell back to the stale server effort. Each pick now carries a token and a failed save drops only its own pick.
… JSON serializability (#8658) * fix(executor): unwrap boxed primitives by internal slot when checking JSON serializability * fix(executor): check the boxed slot first and convert wrappers with ToNumber/ToString
…or (#8768) A clone from GitHub took a minute or more for 1 in 10 checks jobs (up to 4 min), and with ~19 parallel jobs per run one slow clone set the run's pace. Same-repo pull requests now check out with useblacksmith/checkout, which keeps a git mirror on a sticky disk and fetches only new objects. Fork pull requests and pushes keep actions/checkout: the mirror is one disk per repository that job steps can write to, and Git does not re-hash objects read through alternates, so only runs that already share the pull_request caches share it.
* fix(docs): publish full LLM text as a static asset * fix(docs): regenerate full text during development
…oes overwriting unsynced changes (#8764) * fix(billing): converge Stripe cancel/contact syncs and guard webhook echoes over unsynced DB changes * fix(billing): record each committed sync value on every in-flight event so no stale intent can resurface * fix(billing): close remaining stale-intent paths (enterprise retry, legacy echoes, customer restore) * fix(billing): record Team activation's cleared cancellation under the row lock; drop the settled-event scan * fix(billing): lock the subscription before the outbox row in every sync retry path * fix(billing): push each sync's recorded value and reject Stripe reads older than a newer Sim commit * fix(billing): order Stripe-accepted reconcile values by when Stripe was read * fix(billing): recommit the in-flight value, include the plan in seat convergence, bound the reconcile's outbox reads * fix(billing): stamp every sync a later Stripe read confirms; ignore a moved cancel_at * fix(billing): a fresh idempotency key per seat write so a returning value is applied, not replayed * fix(billing): compare membership-driven seat and cancel changes against the committed value, not a possibly-stale row * fix(billing): a membership writer skips only when both the row and the committed value already match * fix(billing): a seat-row repair records no seat-change audit; deterministic latest-sync test helper * test(billing): latest-sync helper fails loudly on an enqueue-time tie
Presidio's BatchAnalyzerEngine defaults nlp.pipe to batch_size=1, so each 2000-text request ran 2000 separate forward passes. Batching at 256 roughly halves analysis time on prod-shaped chunks with identical detections.
* feat(permission-groups): add agent default * fix(permission-groups): validate defaults before seeding agents * fix(permission-groups): normalize defaults and recover policy reads
* feat(integrations): add Ramp and Microsoft Intune * fix(integrations): preserve agent filters and output contracts
* feat(models): add current frontier models * fix(providers): preserve proxy thinking and Mistral answer history * fix(mistral): preserve native continuation and saved answers
* fix(executor): preserve stop-after across pause and resume * fix(executor): retain stop targets while sibling branches pause
…ose its QA gaps (#8782) * improvement(desktop): ship the background executor without a flag * fix(desktop): wait on the shell's own startup signal before refusing a run A background terminal run on a fresh shell waited a flat 8 s for the first prompt and then failed with NO_SHELL_INTEGRATION. A startup file that stops at a question (oh-my-zsh's update prompt) holds the prompt indefinitely, so the run failed with a message blaming the shell's integration. The generated startup files now send a SimStartup marker before any of the user's files run. A shell that never sends it within 8 s of spawning is not instrumented and is refused as before (an unsupported shell at once). One that did gets 30 s to reach its prompt, which tolerates slow startup files on a busy machine; past that the run is refused with the screen it is stuck on, so the agent can ask the user to answer it. Stop now ends the wait, and a shell that exits mid-wait reports SESSION_CLOSED. * test(desktop): durable E2E reports, hermetic shells, and a real network cut - background-executor and terminal-cancel record each check through one shared helper as it finishes. The background-executor report was written from module state in afterAll, so a worker restart after a failure wiped it and the report could read green. Reports from an earlier run are replaced, not appended to. check-report.spec.ts forces a worker restart in a nested Playwright run and asserts the failure stays. - The executor fixture launches the app with zsh and an empty ZDOTDIR, so a developer's .zshrc (an unanswered oh-my-zsh update prompt) cannot hold the first prompt and fail unrelated scenarios. - Scenario B now cuts the network at the socket level: every open connection, the doorbell stream included, drops and new ones are reset. It asserts the result lands exactly once, the doorbell reopens, and new work is picked up after reconnecting. * test(desktop): a call declined in a background chat is never claimed or run The fixture Sim can now decline a held call the way Sim settles it, and the executor suite checks the device never claims it and its command never runs after the approval notification. * test(desktop): run the live desktop suite with the background executor on The live suite's Sim has Redis, so with the executor shipped always-on every turn now binds to the app and runs in the background. The proxy answers the app's registration as an install without Redis would, so the chat-view tests keep covering that path, and the round trip asserts the app stays dormant. Two tests let Sim's own answer through: a call issued after the user switched chats runs on the desktop, and a result reported across a network cut (every connection dropped, the first report landing late) reaches the agent exactly once. * fix(desktop): measure a shell's startup bound from when its startup files began The 30 s bound ran from spawn, so a shell whose startup marker arrived late on a busy machine got less than the promised time for its startup files. The deadline now starts at the marker. * test(desktop): drop the outer runner's variables with the shared omit helper
* feat(api): expose credential sharing and SSO administration * fix(api): make SSO administration atomic and align CLI actions * fix(auth): align SSO admission locks and reject stale provider links * fix(auth): serialize domain administration and align SSO validation * fix(auth): preserve SSO audit events across application operations
* feat(buffer): complete stable API tool coverage * fix(buffer): preserve optional inputs and sort overrides * fix(buffer): accept media when edit assets are null * fix(buffer): normalize cleared inputs and complete test scope * chore(buffer): simplify live validation and clarify docs * fix(buffer): format nested input documentation * fix(buffer): preserve unset fields and enforce response contracts * fix(buffer): validate OneOf filters and deleted post IDs * fix(buffer): validate projected scalar values
…ep XDG zsh configs integrated (#8785) * fix(desktop): poll desktop activity only for users with a desktop, keep XDG zsh configs integrated * fix(desktop): give user zsh files their own ZDOTDIR, drop stale activity, require a live session * test(desktop): check delivered events and real requests, not mock calls
…rands (#8760) * fix(code-placeholders): harden shell placeholder interpolation in arithmetic contexts * fix(code-placeholders): keep integer attributes across re-declaration and unset -f, guard shell values against subscript expansions * fix(code-placeholders): end assignment-value scope at the word, carry attributes into heredoc bodies * fix(code-placeholders): treat a declaration as pending inside its own command's expansions * fix(code-placeholders): narrow arithmetic guards and preserve literal data * fix(code-placeholders): skip continued newlines before comparison operands * fix(code-placeholders): preserve command scope across process substitutions
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.