Test your website in plain English. Get evidence for every result.
Quick start · Write a test · Benchmarks · CLI reference · Contribute
Describe a flow and what should be true at the end. jev-e2e turns it into a test plan, uses Jev to select controls on the page, and runs the test with Playwright. Each result is PASS, FAIL, or BLOCKED, with an HTML report, JSON, and masked screenshots.
Local alpha: CLI and browser workbench for Chromium websites. Install from source; an npm release is not yet available.
Search and filter products → add two items → change quantity → remove an item → refresh and verify the cart. Three models, the same written steps, and 31 independent checks. The video shows actual browser recordings at 10.77× playback.
ebay-benchmark.mp4
Download the 10-second video · All nine attempts and methodology
This is a UI execution experiment with a shared human-authored plan and an extended benchmark observer. It does not measure natural-language planning or the unmodified CLI's reliability on eBay.
- Write cases in plain English. Use an optional planner for prose, or explicit
Goal,Step, andExpecttemplates without it. - Check the outcome. Playwright verifies expectations independently. Missing evidence or unsupported requirements produce BLOCKED.
- See what happened. Reports include expected and observed values, actions, timing, provider usage, and screenshots.
- Replay successful flows. Reuse saved controls and recheck assertions. An unchanged flow can replay with zero model calls; stale targets require Jev to repair them.
- Run locally with limits. Use your own OpenRouter key, authentication fixtures, request limits, deadlines, and cost budget. Stop execution from the workbench or with Ctrl+C.
Requires Node.js 22+ and one OpenRouter API key.
git clone https://github.com/perixtar/jev-e2e.git
cd jev-e2e
npm ci
npm run build
node dist/cli.js setup
cp .env.example .envSet this in your local .env file:
OPENROUTER_API_KEY=your_key_hereThe defaults are typesafe/jev-1.13 for control selection and optional openai/gpt-4.1-mini for interpreting prose. Both use the same OpenRouter key; a separate OpenAI key is unnecessary. Set --planner off to use explicit case templates. There is no silent model fallback.
node dist/cli.js run --url https://playwright.dev/ \
--cases examples/public-site.cases --planner on --headedThis read-only example opens the getting-started guide and checks its URL. Discovery and prose interpretation make paid model calls. Runs default to a $0.05 model budget and a 60-second deadline per case; see the limits and exit codes.
Start these in separate terminals:
npm run demonpm run uiOpen the printed workbench URL, normally http://127.0.0.1:4007. Point it at the demo on http://127.0.0.1:4177, review a plan, run it, and inspect the evidence. The demo command creates demo-only fixtures under .jev-e2e/demo. Use a fresh project name for repeated create tests; a fresh browser context does not reset application data.
Save a case in a .cases file. With --planner on, describe the flow and give a concrete expectation:
Case: Create and persist a project
Auth: @signed-in
Add a new project named "Acme", save it, refresh, and verify it persisted.
Expect: project named "Acme" exists exactly once after reload
Auth: @signed-in references your own local authentication fixture. For the included demo, run this explicit template without the prose planner:
node dist/cli.js run --url http://127.0.0.1:4177 \
--cases examples/simple.cases --planner off \
--fixtures .jev-e2e/demo/fixtures.json --headedThe case reference covers supported steps, expectations, fixtures, saved plans, and replay. Expectations are checked after the required steps; use separate cases for intermediate outcomes.
Measured September 18, 2026, through OpenRouter. We retained three attempts per model, rotated model order, and used fresh guest browser contexts. No retries, substituted models, or discarded failures.
| Model | Completed-case median | API cost / completed case, median | PASS / attempts | Correct UI choices |
|---|---|---|---|---|
| Jev 1.13 | 47.46 s | $0.006678 | 1/3 | 30/31 |
| GPT-5.6 Luna | 61.99 s | $0.027704 | 2/3 | 32/32 |
| Claude Sonnet 5 | 78.62 s | $0.406216 | 1/3 | 19/19 |
Jev was faster and cheaper among completed cases, and it missed one quantity-field choice on another attempt. Three attempts were blocked by eBay availability, and one by a detached-frame bug in the benchmark observer. The verifier stopped the incorrect choice before executing it. Completed-case sample sizes are 1 / 2 / 1; these small, unequal samples do not establish a general accuracy ranking. Site and runner blocks are separate from model mistakes.
The video uses the first registered round for all three models, selected before the batch. Timers and costs follow actual measurement events. Its final scorecard shows all three attempts per model. Read the protocol and every outcome or inspect the public result data.
For product validation, see the separate controlled-demo test results, including healthy, known-broken, and saved-flow replay runs. Those checks do not prove reliability on arbitrary websites.
| Component | Responsibility |
|---|---|
| Optional prose planner | Converts your case into semantic steps, input bindings, and explicit expectations. Saved plans need no new interpretation. |
| Jev | Selects from observed controls and navigation choices using OpenRouter's native Decisions API. |
| Playwright | Executes actions and checks expectations independently. |
| CLI + local workbench | Share the runner, budgets, cancellation, saved plans, and evidence reports. |
Jev uses POST /api/alpha/decisions; the optional planner uses POST /api/v1/chat/completions. Both use standard fetch. Direct TypeSafe transport is not implemented. Provider setup · Technical plan
The alpha supports common Chromium forms, buttons, links, native selects, checkboxes, authenticated fixtures, and async results. Native desktop/mobile apps, CAPTCHA, canvas, complex frames, payment-provider flows, and subjective visual judgments are outside this release. Hosted infrastructure is a later stage.
API keys stay in the server process. Known fixture/auth values are redacted from observations and reports; input fields are masked in screenshots. Visible application text is sent to OpenRouter, and unrelated page content can remain in reports. Inspect evidence before sharing it. Reports stay under your local .jev-e2e/runs; raw trace recording is not implemented. No telemetry is required.
Try a flow on your website and open an issue with a sanitized case, expected outcome, and what happened. Unsupported flows and reproducible failures are useful contributions. If the project helps you, give it a star.
npm run check
npm testOrdinary tests use scripted provider responses and real Chromium, with no paid calls. The contributing guide covers development, opt-in paid checks, and review requirements.
Test plan · UI design · Research · Live provider checks · MIT license