An LLM works out how to do a job inside a legacy UI that has no API. The successful run is compiled into a typed, versioned capability artifact. From then on the artifact replays deterministically, with no model in the decision loop, and returns a result an AI agent can actually branch on.
The model discovers. The artifact is the capability. Replay is how the agent invokes it.
The design rationale — schema decisions, the error taxonomy, the control-transfer model, and what was deliberately cut — is in REPORT.md.
A local mock of a fictional core banking console, Meridian Core 4.2
(apps/core-banking/). It is deliberately hostile, because that is what the target
environment looks like:
- a real HTML
<frameset>with a nav frame and a main frame, - nested layout tables,
<font>tags, no test IDs, no ARIA, - form fields with no accessible name at all — the member-number input is
identified only by a
<b>MEMBER NUMBER</b>in the adjacent<td>, - screen identifiers (
MI0200,SA0320) stamped on every screen, the way these systems actually do, - distinct error codes for "no such member" (
MI-0011), "below minimum deposit" (SA-0021), "not entitled" (SEC-0413) and a genuine host crash, - and an admin hook that injects real runtime conditions on demand: session expiry, a six-second hang, a modal maintenance notice, a 500 from the host.
Everything is fictional. No real service, no real credentials, no real PII.
Requires Node 20+.
npm install # also downloads Chromium via Playwright
cp .env.example .env # throwaway credentials for the local mock.env holds the mock's sign-on credentials (teller1 / demo123). They are read
from the environment at execution time and never written into an artifact or a log —
see Safety.
To run a discovery of your own you also need a model. Everything else — replay, the demo, the whole test suite — runs without one.
# in .env
CUA_PROVIDER=openai # or anthropic
OPENAI_API_KEY=sk-...One command. It starts the mock, then invokes the same recorded artifact five times and produces the five outcomes a calling agent has to tell apart.
npm run demo| Scene | What happens | Result |
|---|---|---|
| 1 | Member 12345 is looked up | success + typed regularSavingsBalance |
| 2 | Member 00000 does not exist | business_outcome member_not_found |
| 3 | A maintenance modal blocks the screen | recovered — detected, dismissed, retried |
| 4 | The host returns an unrecoverable error | failed + screenshot and DOM snapshot |
| 5 | A vendor renamed the button; nothing resolves | escalates, a human clicks it in the same live session, run resumes → recovered |
Evidence for each lands in evidence/demo/<scene>-<runid>/.
Scene 5 uses a scripted stand-in for the operator so the demo is reproducible. To do it yourself, with a real browser and a real console:
npm run mock # terminal 1
npm run replay -- \
--artifact evidence/capabilities/lookup-savings-balance.json \
--params '{"memberNumber":"12345"}' \
--headed # terminal 2When a run gets stuck it prints an operator console URL. Open it: you get the reason, the blocked step, a screenshot, and Resume / Abort. Drive the already-open browser window yourself, then press Resume — the automation picks the flow back up from that step. Or from a third terminal:
npm run operator -- status
npm run operator -- resume --note "clicked it by hand"npm run mock # terminal 1
npm run discover -- \
--goal "Look up member 12345 and read their current regular savings balance." \
--headedThe agent signs on, navigates, fills, reads, and calls finish. Every action it
proposes goes through the policy gate; every element it touches gets a locator
cascade synthesised and verified against the live page before it is written down.
The artifact is saved to evidence/capabilities/<id>.json, the transcript, screenshots
and structured log to evidence/local/<run-id>/.
The model is given no navigate tool and no way to read a credential. It has to reach
every screen by clicking, and it asks for fill_credential(ref, field) without ever
seeing the value.
npm run replay -- \
--artifact evidence/capabilities/lookup-savings-balance.json \
--params '{"memberNumber":"12345"}'Useful flags:
| Flag | Effect |
|---|---|
--inject <mode>[:count] |
Ask the mock for a runtime condition: session_expired, slow, app_error, interstitial, not_authorized |
--overlay <path> |
Specialise the capability for one tenant |
--require-approved |
Refuse to run a draft artifact unattended |
--approve-risk irreversible |
Pre-approve a risk class for this invocation only |
--no-escalate |
Fail instead of asking a human |
--update-stability |
Record the attempt against the artifact's stability counters |
--headed |
Watch it |
# a legitimate business answer, not an error — exit code 0
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
--params '{"memberNumber":"00000"}'
# the session drops mid-flow; it signs back on and re-walks
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
--params '{"memberNumber":"12345"}' --inject session_expired
# the host falls over; it stops and says exactly where
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
--params '{"memberNumber":"12345"}' --inject app_errorThis is the surface an AI agent programs against: discover a capability, read a JSON Schema for its arguments, invoke it.
npm run capabilities -- list
npm run capabilities -- describe lookup-savings-balance
npm run capabilities -- invoke lookup-savings-balance --params '{"memberNumber":"12345"}'describe returns the parameter schema, the typed outputs, and the business
outcomes the capability may legitimately return — so the calling agent can plan for
"no such member" instead of treating it as a bug.
npm run replay -- \
--artifact evidence/capabilities/lookup-savings-balance.json \
--overlay config/tenants/harbor-point-cu.overlay.json \
--params '{"memberNumber":"12345"}'The overlay is a delta, not a fork: one relabelled button and one tenant-specific runtime condition. Its locator is prepended to the base cascade, so when the override is stale — as it is against this mock — the run falls back to the base locators, succeeds, and emits a drift warning naming the tier that actually matched.
npm test # 65 tests, no API key, no network
npm run typecheck| Suite | Covers |
|---|---|
tests/unit.test.ts |
Schema validation, the invocation contract, risk classification, the policy gate, redaction, tenant overlays |
tests/discovery.test.ts |
The real agent loop with the model swapped for a script: locator verification, relational fallback, credential and PII hygiene, record-time verification of the success condition |
tests/replay.test.ts |
Determinism, business outcomes, dismissal, reauthentication, hard failures, slow screens, approval gating |
tests/handoff.test.ts |
The session lock, control transfer, an operator finishing a step by hand, abort, and the approval gate on an irreversible action |
The discovery and replay suites build their fixture by running real discovery against the real mock with a scripted provider standing in for the model, so the two halves of the system are tested against each other and cannot drift apart.
evidence/ is committed.
evidence/
capabilities/lookup-savings-balance.json the artifact (recorded by gpt-4o)
discovery/discovery-20260911T202612-k9vs/ the genuine LLM discovery run
demo/ the five replay runs from `npm run demo`
Each run directory holds run.jsonl (structured, redacted, one event per line),
screenshots/, snapshots/ (DOM captures taken on failure and at escalation), and
result.json.
evidence/local/ is where your own runs go, and is gitignored.
apps/core-banking/ the hostile mock: frameset, tables, error codes, fault injection
config/
app-profiles/ per-vendor-product config: entry point, screen-ID pattern,
credential refs, and the curated error taxonomy
policy.json allowlist, risk ceiling, what needs human approval
tenants/ tenant overlays
src/
agent/ the observe-decide-act loop, tool schemas, artifact compiler
capability/ the artifact schema, locator model, validation, value binding
replay/ the deterministic executor and the result contract
surface/ the SurfaceAdapter seam; the Playwright web implementation
policy/ allowlist, risk classification, credentials, redaction
session/ the control lock and escalation
operator/ the minimal operator console
evidence/ structured run recorder
scripts/demo.ts the demo path