Skip to content

Repository files navigation

jev-e2e

Test your website in plain English. Get evidence for every result.

Status: alpha Node: 22+ License: MIT

Quick start · Write a test · Benchmarks · CLI reference · Contribute

Describe a flow and what should be true at the end. jev-e2e turns it into a test plan, uses Jev to select controls on the page, and runs the test with Playwright. Each result is PASS, FAIL, or BLOCKED, with an HTML report, JSON, and masked screenshots.

Local alpha: CLI and browser workbench for Chromium websites. Install from source; an npm release is not yet available.

Watch the eBay comparison

Search and filter products → add two items → change quantity → remove an item → refresh and verify the cart. Three models, the same written steps, and 31 independent checks. The video shows actual browser recordings at 10.77× playback.

ebay-benchmark.mp4

Download the 10-second video · All nine attempts and methodology

This is a UI execution experiment with a shared human-authored plan and an extended benchmark observer. It does not measure natural-language planning or the unmodified CLI's reliability on eBay.

Why jev-e2e?

  • Write cases in plain English. Use an optional planner for prose, or explicit Goal, Step, and Expect templates without it.
  • Check the outcome. Playwright verifies expectations independently. Missing evidence or unsupported requirements produce BLOCKED.
  • See what happened. Reports include expected and observed values, actions, timing, provider usage, and screenshots.
  • Replay successful flows. Reuse saved controls and recheck assertions. An unchanged flow can replay with zero model calls; stale targets require Jev to repair them.
  • Run locally with limits. Use your own OpenRouter key, authentication fixtures, request limits, deadlines, and cost budget. Stop execution from the workbench or with Ctrl+C.

Quick start

Requires Node.js 22+ and one OpenRouter API key.

1. Install from source

git clone https://github.com/perixtar/jev-e2e.git
cd jev-e2e
npm ci
npm run build
node dist/cli.js setup
cp .env.example .env

2. Add your key

Set this in your local .env file:

OPENROUTER_API_KEY=your_key_here

The defaults are typesafe/jev-1.13 for control selection and optional openai/gpt-4.1-mini for interpreting prose. Both use the same OpenRouter key; a separate OpenAI key is unnecessary. Set --planner off to use explicit case templates. There is no silent model fallback.

3. Run a website test

node dist/cli.js run --url https://playwright.dev/ \
  --cases examples/public-site.cases --planner on --headed

This read-only example opens the getting-started guide and checks its URL. Discovery and prose interpretation make paid model calls. Runs default to a $0.05 model budget and a 60-second deadline per case; see the limits and exit codes.

4. Try the workbench

Start these in separate terminals:

npm run demo
npm run ui

Open the printed workbench URL, normally http://127.0.0.1:4007. Point it at the demo on http://127.0.0.1:4177, review a plan, run it, and inspect the evidence. The demo command creates demo-only fixtures under .jev-e2e/demo. Use a fresh project name for repeated create tests; a fresh browser context does not reset application data.

Write a test

Save a case in a .cases file. With --planner on, describe the flow and give a concrete expectation:

Case: Create and persist a project
Auth: @signed-in
Add a new project named "Acme", save it, refresh, and verify it persisted.
Expect: project named "Acme" exists exactly once after reload

Auth: @signed-in references your own local authentication fixture. For the included demo, run this explicit template without the prose planner:

node dist/cli.js run --url http://127.0.0.1:4177 \
  --cases examples/simple.cases --planner off \
  --fixtures .jev-e2e/demo/fixtures.json --headed

The case reference covers supported steps, expectations, fixtures, saved plans, and replay. Expectations are checked after the required steps; use separate cases for intermediate outcomes.

Live eBay benchmark

Measured September 18, 2026, through OpenRouter. We retained three attempts per model, rotated model order, and used fresh guest browser contexts. No retries, substituted models, or discarded failures.

Model Completed-case median API cost / completed case, median PASS / attempts Correct UI choices
Jev 1.13 47.46 s $0.006678 1/3 30/31
GPT-5.6 Luna 61.99 s $0.027704 2/3 32/32
Claude Sonnet 5 78.62 s $0.406216 1/3 19/19

Jev was faster and cheaper among completed cases, and it missed one quantity-field choice on another attempt. Three attempts were blocked by eBay availability, and one by a detached-frame bug in the benchmark observer. The verifier stopped the incorrect choice before executing it. Completed-case sample sizes are 1 / 2 / 1; these small, unequal samples do not establish a general accuracy ranking. Site and runner blocks are separate from model mistakes.

The video uses the first registered round for all three models, selected before the batch. Timers and costs follow actual measurement events. Its final scorecard shows all three attempts per model. Read the protocol and every outcome or inspect the public result data.

For product validation, see the separate controlled-demo test results, including healthy, known-broken, and saved-flow replay runs. Those checks do not prove reliability on arbitrary websites.

How it works

Component Responsibility
Optional prose planner Converts your case into semantic steps, input bindings, and explicit expectations. Saved plans need no new interpretation.
Jev Selects from observed controls and navigation choices using OpenRouter's native Decisions API.
Playwright Executes actions and checks expectations independently.
CLI + local workbench Share the runner, budgets, cancellation, saved plans, and evidence reports.

Jev uses POST /api/alpha/decisions; the optional planner uses POST /api/v1/chat/completions. Both use standard fetch. Direct TypeSafe transport is not implemented. Provider setup · Technical plan

Scope and privacy

The alpha supports common Chromium forms, buttons, links, native selects, checkboxes, authenticated fixtures, and async results. Native desktop/mobile apps, CAPTCHA, canvas, complex frames, payment-provider flows, and subjective visual judgments are outside this release. Hosted infrastructure is a later stage.

API keys stay in the server process. Known fixture/auth values are redacted from observations and reports; input fields are masked in screenshots. Visible application text is sent to OpenRouter, and unrelated page content can remain in reports. Inspect evidence before sharing it. Reports stay under your local .jev-e2e/runs; raw trace recording is not implemented. No telemetry is required.

Help shape the project

Try a flow on your website and open an issue with a sanitized case, expected outcome, and what happened. Unsupported flows and reproducible failures are useful contributions. If the project helps you, give it a star.

npm run check
npm test

Ordinary tests use scripted provider responses and real Chromium, with no paid calls. The contributing guide covers development, opt-in paid checks, and review requirements.

Test plan · UI design · Research · Live provider checks · MIT license

About

Natural-language end-to-end tests for web apps, powered by Jev and Playwright.

Topics

Resources

Contributing

Stars

11 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages