Measures how well AI coding agents write Pixeltable code under different context levels (cold, with skill, with skill + MCP).
30 cells were run with real claude_code (u1 PDF RAG), then re-graded
offline on the saved artifacts after fixing harness bugs that produced
false negatives: installed skill files counted as agent output, top-level
@pxt.udf rejected under __main__ exec, TableModel answers never
materialized (no pxt schema update), check scripts calling a different
pxt than the harness interpreter, and missing env deps. Same generated
code, corrected grading:
| Context | Pass (score >= 3.0) | Functional pass | Hallucinations |
|---|---|---|---|
| cold | 90% | 50% | 0 |
| skill | 90% | 20% | 0 |
| skill_mcp | 80% | 20% | 0 |
No context lift on this story. The earlier "+33pp skill lift" figure was measured on stale graders that rewarded retired APIs and never executed code; treat it as void.
Interpretation notes:
- Pass is static-heavy. A perfect static score alone (5.0 * 0.6 = 3.0) reaches the pass bar without executing, so the functional pass rate is the stricter "works end-to-end" measure.
- Dominant real failure mode is the TableModel class-body DSL: forward
NameError(referencingChunksinside its own definition),Documents.view(not a real API),base=given an iterator call instead of a table,.data/.self_pathon Array-typed columns, and treatingQueryTemplateFunctionparameters as attributes (search_chunks.question). - Env deps the eval needs:
pixeltable[serve]for the FastAPIRouter stack (fastapi, uvicorn,python-multipartforuploadfile_inputsroutes),openai,tiktoken(token_limitsplitter),spacy+en_core_web_sm(sentence splitter),sentence-transformers(HF embeddings). Provider keys reach the sandbox via a symlinked~/.pixeltable/config.toml. - Infra flake ~10%: every cell boots a fresh embedded postgres; a couple
of
initdbcalls failed transiently under 30 back-to-back launches. Re-running a flake cell in isolation passes.
The full pipeline was exercised end-to-end: env setup, file collection,
sandbox execution, functional checks, scoring, and results.json output.
The u1 reference answer now runs ask() for real — embeddings, similarity
search and chat_completions all fire inside the sandbox ("Employees
receive 20 vacation days per calendar year").
Functional checks prove the app works, not just that the schema exists:
- u1 ingests the fixture PDFs, then runs a real
.similarity()query — it only resolves when a live embedding index exists. Provider auth failures are recorded as env issues, not app defects. - u3 runs
pxt init+pxt schema update, thenpxt service update, discovers the serving port viapxt service list --json, and POSTs to/analyze. A pass requires a running service exposing the route; the service and its daemon are stopped afterward (each sandbox gets a privatePXT_PORT).
pytest # all tests, incl. the ~15s live service test
pytest -m "not slow" # static + canary onlytests/test_verifier.py— static layer:command_evidence(prose earns no credit), everyHALLUCINATED_APISpattern must match a real snippet, known-good vs known-bad scoring, skipped-functional reweighting.tests/test_canary.py— everyevals/**/answer/must satisfy its owngrader.py; a failure means the grader drifted, not the agent.tests/test_u3_functional.py— end-to-end canary: a known-good TableModel app must boot and serve through the real sandbox.
Each result row records environment (resolved pixeltable/fastapi/spacy/
etc. versions) so scores are attributable to a dep set — pixeltable is
unbounded above. --judge-model <name> enables the cross-model LLM judge
(default off; without it the composite is static+functional only). The
summary's Func column shows how many cells ran the functional check vs
skipped it.
pip install -e ".[dev]"
python scripts/generate_fixtures.py
python -m eval run --spikeTASK.txt → Runner (Claude Code / Cursor SDK) → Generated Code → Verifier → Score
Evals live in evals/<category>/<name>/ (Convex-style). Each contains:
TASK.txt— the prompt sent to the agentanswer/— human-curated reference solution (optional)grader.py— patterns for static analysis (optional, uses defaults if missing)
Runners drive real agent runtimes — Claude Code via --print headless mode, Cursor via @cursor/sdk. These are NOT raw API calls; they include the full tool-use loop (file read/write, shell, web search, self-correction).
Environments configure context levels: cold (no hints), skill installed (npx skills add), skill + MCP server, plugin (Claude Code marketplace install).
Verifiers check static patterns (positive/negative grep) and optionally run code in a sandbox for functional correctness.
Every eval in the tree is runnable: --story accepts eval ids
(001-rag/pdf_rag) or category prefixes (004-idioms). Evals grade
statically through a generic verifier; a grader may set EXECUTE_CODE = True
to additionally require a clean sandbox exec. The curated stories below
carry bespoke functional checks.
evals/
├── 000-fundamentals/ # create_table, computed_columns, embedding_index
├── 001-rag/ # pdf_rag, semantic_search
├── 002-video/ # frame_extraction
├── 003-agents/ # tool_calling
├── 004-idioms/ # no_langchain, no_pandas_store, computed_not_loop
├── 005-hard/ # error_recovery, incremental_update, multi_view_pipeline
├── 006-negative-controls/ # raw_sql_query, simple_pandas_groupby, static_file_transform
├── 007-scaffolding/ # use_scaffolder, pxt_service, crud_service
└── external_benchmarks.json # registry of external benchmarks/prompt banks
python -m eval list # List all evals + external coverage
python -m eval run --spike # R0 spike
python -m eval run -c skill -r claude_code --reps 3
python -m eval.orchestrator --story 004-idioms --reps 3 # run an eval category
python -m eval.orchestrator --story u4 --context cold --reps 1
python -m eval.orchestrator --judge-model gpt-4o # add the LLM judge layer
python -m eval status # Show last results
python -m eval status --failed # Show failures only
python -m eval publish # Render latest run to RESULTS.md
python -m eval publish --run spike_20260919_200809 # pick a run| Axis | Values |
|---|---|
| Eval | 19 evals across 8 categories + 5 curated stories |
| Runner | Claude Code (--print), Cursor SDK |
| Context | cold, +skill, +skill+MCP, +plugin |
| Reps | 3 per cell (for variance) |
| Story | Description |
|---|---|
| u1 | PDF RAG pipeline (base table + chunk view + embedding + LLM); functional: fixture PDFs ingested + similarity query runs |
| u2 | Project scaffolding (pixeltable-new / pxt init + example); functional: skipped (static only) |
| u3 | REST API via TableModel + FastAPIRouter; functional: schema update + service actually boots and answers POST /analyze |
| u4 | CRUD routes over TableModel (dependent instructions); functional: insert -> query -> delete exercised against the live service, no provider keys needed |
| u5 | Incremental computation (two insert waves + late column); functional: backfill proven, hermetic |
| Metric | Range | What it measures |
|---|---|---|
| Pass | 0/1 | Composite >= 3.0 AND no ran functional check failed |
| Idiomaticity | 0-5 | Uses computed columns, embedding indexes, TableModel + FastAPIRouter, pxt service update |
| Hallucinations | int | Non-existent APIs called (lower = better) |
| Turns | int | How many agent turns to produce code |
| Cost | USD | total_cost_usd from the runner envelope |
A ran-and-failed functional check vetoes the pass: static coverage alone (0.6 * 5.0 = 3.0) cannot certify code that does not execute. Skipped checks (e.g. missing provider keys) reweight the composite instead of failing.
evals/external_benchmarks.json registers external coding benchmarks and
prompt banks for comparison. python -m eval list shows how each category
maps to what this repo measures: coding benchmarks (SWE-bench, LiveBench,
HumanEval) are comparison references only, prompt banks supply task patterns
for curated stories (u4+), and eval frameworks (Promptfoo, Braintrust) are
alternatives, not content.
Run: matrix_20260929_134627 | pixeltable=0.7.8, fastapi=0.141.1, openai=3.16.2, python=3.12.14
- Skill lift: cold 100% -> skill 50% pass rate (n=1/4).
- Functional lift: cold 100% -> skill 50% on executed cells (n=1/4).
- Looks-right-but-broken: 1 cell(s) passed static checks yet failed execution -- what a good skill should reduce.
- Spend: $6.07 total, median 9 turns.
- 19 trial(s) excluded as infrastructure errors.
| Story | Context | n | Pass | Func ran | Mean score |
|---|---|---|---|---|---|
| u1 | cold | 3 | all infra | - | - |
| u1 | skill | 3 | all infra | - | - |
| u3 | cold | 1 (+2 infra) | 100% [21-100%] | 1/1 | 4.3 |
| u3 | skill | 1 (+2 infra) | 100% [21-100%] | 1/1 | 5.0 |
| u4 | cold | 3 | all infra | - | - |
| u4 | skill | 1 (+2 infra) | 0% [0-79%] | 1/1 | 3.8 |
| u5 | cold | 3 | all infra | - | - |
| u5 | skill | 2 (+1 infra) | 50% [9-91%] | 2/2 | 3.8 |
Raw transcripts and extracted code stay local (results/ is gitignored):
they are arbitrary agent output and may contain environment details. What
gets committed is the per-trial results.json plus a rendered summary:
python -m eval run -s u1 u3 u4 u5 -c cold skill --reps 3 # produce a run
python -m eval publish # -> RESULTS.md
git add RESULTS.md results/*/results.json && git commit -m "results: ..."publish writes three artifacts: RESULTS.md (full table, action items
grouped by failure reason), results/chart.svg (pass-rate bars per
story x context), and the eval-results block of this README (lift,
functional-vs-static gap, hallucinations, spend). Runs recorded before
the functional veto are flagged automatically.
Run weekly, not nightly, and only against a fixed dep cohort plus one
floating-latest cell: a trend that mixes pixeltable releases, skill HEAD,
and CLI updates attributes nothing. Track functional pass, hallucinations,
and cost_usd, not composite pass. Without provider keys, similarity/LLM
probes report as skipped rather than passed, and the Func column shows
coverage. Each matrix cell is a real agent run (minutes + API spend); a
3-rep subset (u1, u3, u4, u5 x cold/skill) is the cheapest useful cohort.
Sandbox caveat before scheduling: generated code and the agent's shell
run as your user with your full environment (the runner passes
--dangerously-skip-permissions, and provider keys reach the sandbox via
a symlinked ~/.pixeltable/config.toml). A fresh PIXELTABLE_HOME isolates
the catalog, not the host — on a schedule, run inside a container with
scoped credentials.
After running the spike:
- Lift cold → skill ≥ 30pp: Premise validated → build remaining 9 stories ✓
- Lift 10-30pp: Weak → re-examine SKILL.md content
- Lift < 10pp: Thesis not supported → investigate
- Variance > 25pp: Need more reps
U1 outcome: no lift (90/90/80 at n=10) — below the 10pp floor, but the cell size is too small to distinguish skill harm from noise. Before treating "context adds nothing" as the answer, run a story where cold starts weaker (u1 may be near-saturated) and consider tightening the pass bar so static credit alone cannot carry a broken app.
- Python 3.10+
claudeCLI withANTHROPIC_API_KEYfor Claude Code runner- Node.js 18+ with
CURSOR_API_KEYfor Cursor SDK runner pip install -e .pulls the functional-check deps (pixeltable[serve]for FastAPIRouter routes incl.python-multipart, openai, tiktoken, spacy + en_core_web_sm wheel, sentence-transformers); provider keys come from~/.pixeltable/config.tomlor env vars