Skip to content

Repository files navigation

Computer-Use Automation

An LLM works out how to do a job inside a legacy UI that has no API. The successful run is compiled into a typed, versioned capability artifact. From then on the artifact replays deterministically, with no model in the decision loop, and returns a result an AI agent can actually branch on.

The model discovers. The artifact is the capability. Replay is how the agent invokes it.

The design rationale — schema decisions, the error taxonomy, the control-transfer model, and what was deliberately cut — is in REPORT.md.


What it runs against

A local mock of a fictional core banking console, Meridian Core 4.2 (apps/core-banking/). It is deliberately hostile, because that is what the target environment looks like:

  • a real HTML <frameset> with a nav frame and a main frame,
  • nested layout tables, <font> tags, no test IDs, no ARIA,
  • form fields with no accessible name at all — the member-number input is identified only by a <b>MEMBER NUMBER</b> in the adjacent <td>,
  • screen identifiers (MI0200, SA0320) stamped on every screen, the way these systems actually do,
  • distinct error codes for "no such member" (MI-0011), "below minimum deposit" (SA-0021), "not entitled" (SEC-0413) and a genuine host crash,
  • and an admin hook that injects real runtime conditions on demand: session expiry, a six-second hang, a modal maintenance notice, a 500 from the host.

Everything is fictional. No real service, no real credentials, no real PII.


Setup

Requires Node 20+.

npm install          # also downloads Chromium via Playwright
cp .env.example .env # throwaway credentials for the local mock

.env holds the mock's sign-on credentials (teller1 / demo123). They are read from the environment at execution time and never written into an artifact or a log — see Safety.

To run a discovery of your own you also need a model. Everything else — replay, the demo, the whole test suite — runs without one.

# in .env
CUA_PROVIDER=openai        # or anthropic
OPENAI_API_KEY=sk-...

Demo path

One command. It starts the mock, then invokes the same recorded artifact five times and produces the five outcomes a calling agent has to tell apart.

npm run demo
Scene What happens Result
1 Member 12345 is looked up success + typed regularSavingsBalance
2 Member 00000 does not exist business_outcome member_not_found
3 A maintenance modal blocks the screen recovered — detected, dismissed, retried
4 The host returns an unrecoverable error failed + screenshot and DOM snapshot
5 A vendor renamed the button; nothing resolves escalates, a human clicks it in the same live session, run resumes → recovered

Evidence for each lands in evidence/demo/<scene>-<runid>/.

Watch the handoff happen

Scene 5 uses a scripted stand-in for the operator so the demo is reproducible. To do it yourself, with a real browser and a real console:

npm run mock                                   # terminal 1

npm run replay -- \
  --artifact evidence/capabilities/lookup-savings-balance.json \
  --params '{"memberNumber":"12345"}' \
  --headed                                     # terminal 2

When a run gets stuck it prints an operator console URL. Open it: you get the reason, the blocked step, a screenshot, and Resume / Abort. Drive the already-open browser window yourself, then press Resume — the automation picks the flow back up from that step. Or from a third terminal:

npm run operator -- status
npm run operator -- resume --note "clicked it by hand"

The other commands

Discover (needs a model)

npm run mock          # terminal 1

npm run discover -- \
  --goal "Look up member 12345 and read their current regular savings balance." \
  --headed

The agent signs on, navigates, fills, reads, and calls finish. Every action it proposes goes through the policy gate; every element it touches gets a locator cascade synthesised and verified against the live page before it is written down. The artifact is saved to evidence/capabilities/<id>.json, the transcript, screenshots and structured log to evidence/local/<run-id>/.

The model is given no navigate tool and no way to read a credential. It has to reach every screen by clicking, and it asks for fill_credential(ref, field) without ever seeing the value.

Replay (never needs a model)

npm run replay -- \
  --artifact evidence/capabilities/lookup-savings-balance.json \
  --params '{"memberNumber":"12345"}'

Useful flags:

Flag Effect
--inject <mode>[:count] Ask the mock for a runtime condition: session_expired, slow, app_error, interstitial, not_authorized
--overlay <path> Specialise the capability for one tenant
--require-approved Refuse to run a draft artifact unattended
--approve-risk irreversible Pre-approve a risk class for this invocation only
--no-escalate Fail instead of asking a human
--update-stability Record the attempt against the artifact's stability counters
--headed Watch it
# a legitimate business answer, not an error — exit code 0
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
  --params '{"memberNumber":"00000"}'

# the session drops mid-flow; it signs back on and re-walks
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
  --params '{"memberNumber":"12345"}' --inject session_expired

# the host falls over; it stops and says exactly where
npm run replay -- --artifact evidence/capabilities/lookup-savings-balance.json \
  --params '{"memberNumber":"12345"}' --inject app_error

The agent-facing catalog

This is the surface an AI agent programs against: discover a capability, read a JSON Schema for its arguments, invoke it.

npm run capabilities -- list
npm run capabilities -- describe lookup-savings-balance
npm run capabilities -- invoke  lookup-savings-balance --params '{"memberNumber":"12345"}'

describe returns the parameter schema, the typed outputs, and the business outcomes the capability may legitimately return — so the calling agent can plan for "no such member" instead of treating it as a bug.

Multi-tenant specialisation

npm run replay -- \
  --artifact evidence/capabilities/lookup-savings-balance.json \
  --overlay config/tenants/harbor-point-cu.overlay.json \
  --params '{"memberNumber":"12345"}'

The overlay is a delta, not a fork: one relabelled button and one tenant-specific runtime condition. Its locator is prepended to the base cascade, so when the override is stale — as it is against this mock — the run falls back to the base locators, succeeds, and emits a drift warning naming the tier that actually matched.


Tests

npm test        # 65 tests, no API key, no network
npm run typecheck
Suite Covers
tests/unit.test.ts Schema validation, the invocation contract, risk classification, the policy gate, redaction, tenant overlays
tests/discovery.test.ts The real agent loop with the model swapped for a script: locator verification, relational fallback, credential and PII hygiene, record-time verification of the success condition
tests/replay.test.ts Determinism, business outcomes, dismissal, reauthentication, hard failures, slow screens, approval gating
tests/handoff.test.ts The session lock, control transfer, an operator finishing a step by hand, abort, and the approval gate on an irreversible action

The discovery and replay suites build their fixture by running real discovery against the real mock with a scripted provider standing in for the model, so the two halves of the system are tested against each other and cannot drift apart.


Evidence

evidence/ is committed.

evidence/
  capabilities/lookup-savings-balance.json   the artifact (recorded by gpt-4o)
  discovery/discovery-20260911T202612-k9vs/  the genuine LLM discovery run
  demo/                                      the five replay runs from `npm run demo`

Each run directory holds run.jsonl (structured, redacted, one event per line), screenshots/, snapshots/ (DOM captures taken on failure and at escalation), and result.json.

evidence/local/ is where your own runs go, and is gitignored.


Layout

apps/core-banking/     the hostile mock: frameset, tables, error codes, fault injection
config/
  app-profiles/        per-vendor-product config: entry point, screen-ID pattern,
                       credential refs, and the curated error taxonomy
  policy.json          allowlist, risk ceiling, what needs human approval
  tenants/             tenant overlays
src/
  agent/               the observe-decide-act loop, tool schemas, artifact compiler
  capability/          the artifact schema, locator model, validation, value binding
  replay/              the deterministic executor and the result contract
  surface/             the SurfaceAdapter seam; the Playwright web implementation
  policy/              allowlist, risk classification, credentials, redaction
  session/             the control lock and escalation
  operator/            the minimal operator console
  evidence/            structured run recorder
scripts/demo.ts        the demo path

About

LLM discovers a UI flow once; it becomes a typed capability artifact that replays deterministically with no model in the loop.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages