Biological data → predictive ML → scientific agents → independently checkable execution → regulated biopharma software.
I build life-science software where the evidence is inspectable: frozen splits, deterministic computation, machine-checked claims, replayable runs and explicit abstention when the data does not support a result.
Currently focused on trustworthy AI and human-relevant preclinical models: New Approach Methodologies (NAMs), computational toxicology, and ML tooling for drug discovery that reduces reliance on animal testing.
Related projects are grouped into monorepos — each subdirectory is a self-contained package with its own tests, config, and full commit history. The original repos are archived with pointers, so old links still resolve.
Bioinformatics and assay data — bio-qc + lab-informatics
| Repository | What it is |
|---|---|
| cytof-qc | Mass-cytometry benchmark: FCS parsing, channel/event QC, Leiden clustering on all 167k events of Levine_13dim with agreement scored on the 81.7k labeled subset (ARI 0.867). Merged and missed populations reported per gate. |
| fcs-io | Dependency-free FCS 3.0/3.1 reader/writer with explicit vendor-quirk handling (delimiter escapes, blank-header offsets, endianness). Validated bit-exact vs fcsparser on 55 real Levine files, 432k events — committed crosscheck artifact. |
| scrna-qc | Single-cell RNA-seq QC on scanpy/AnnData as a Snakemake DAG (ingest, threshold-filter, embed, report) with config-driven thresholds and per-cluster QC so a pooled pass cannot hide a depleted population. Committed PBMC 3k run: 57 high-mito cells filtered, 8 Leiden clusters, top markers on canonical PBMC families (LYZ/S100A8 monocytes, NKG7 NK, CD74 antigen-presenting). |
| spatial-qc | Visium spot QC plus a filtering-strategy benchmark (fixed vs MAD-adaptive vs tissue-only) on two public Space Ranger exports (V1 mouse brain, V1 breast cancer) — the strategy ranking reverses between them, committed. |
| nf | DSL2 Nextflow pipeline in nf-core module style (meta.yml/environment.yml/stub per module, nf-test coverage) running FCSIO_DEMO -> PARSE -> STATS over per-file parallel tasks. |
| labStackDev | RNA-seq quantification pipelines (Salmon, STAR+Salmon hybrid, pyDESeq2, METAFlux) validated at r = 0.984 against a CLC baseline, plus a SQLite LIMS with a hash-chained audit log, HMAC-signed audit rows, reason-for-change fields, and a VALIDATION.md mapping properties to IQ/OQ/PQ-style tests. |
| organoid-qc | Organoid fidelity scoring vs CELLxGENE reference centroids, reported per cluster with unmapped fractions, plus an Opentrons Flex dosing protocol validated by the official Protocol Engine. |
| cultivated-meat-multiomic | RNA + metabolic-flux clustering, 30-gene panel selection, conformal prediction, and drift monitoring on public data with cross-species checks. |
| statgen | Statistical-genetics pipeline as a config-driven Snakemake DAG: variant/sample QC (missingness, MAF, HWE, within-ancestry heterozygosity), genotype PCA, linear/logistic association with genomic-control λ, and a PC-residualized relatedness scan. Committed synthetic cohort recovers all post-QC planted causal variants at Bonferroni (β corr 0.99) and measures stratification as λ 2.65 → 0.97 with PC covariates; 1000 Genomes chr22 mode accounts for LD-proxy hits and states regional-LD limits. |
Protein engineering and ML — protein-ml
| Repository | What it is |
|---|---|
| protein-diffusion | Conditional DDPM over the GB1 fitness landscape scored against the measured oracle. The v1 model memorized (46% copied rows); v2 fitness conditioning steers measurably (80% fit vs 4% random). DDP, FastAPI service, Flax/JAX port matching torch to 3e-6. Weights on HF. |
| active-learning-loop | GP-UCB acquisition vs a 20-seed random baseline on two measured landscapes (GB1, AAV2). ~79x top-100 enrichment on GB1; AAV2 holds ~8x with overlapping AUBC bands, reported as a partial replication. ESM-2 embeddings lose to one-hot, committed as a negative finding. |
| protein-stability-uncertainty | Sequence to melting-point regression on the Meltome atlas (27,951 proteins) with a homology-separated split. Marginal conformal coverage (0.90) hides a 0.96 → 0.85 gradient across distance; Mondrian calibration recovers flat coverage. |
| protein-design-ops | ProteinMPNN generation + ESM-2 rescoring + ESMFold pLDDT screen with pinned-seed provenance. The reproducibility caveat caught upstream --seed 0 silently randomizing. |
| dti-fusion | Drug-target interaction on DAVIS (fingerprints + ESM-2), modality-ablated on held-out proteins. Fusion wins ranking/MSE; drug-only edges MAE. |
Trustworthy ML and scientific agents — trust-tools
| Repository | What it is |
|---|---|
| oncology-coscientist | TCGA survival analysis (Cox PH, RSF) where LangGraph agents draft reports and a deterministic verifier binds every number to its computation. See "What the verifier caught": fabricated cohort sizes and invented metrics for abstained models. |
| llm-posttraining | SFT + hand-rolled DPO + GRPO on SmolLM2-135M for abstention vs fabrication. DPO collapses deployed behavior at 100% train accuracy; shaped GRPO repairs it. A second reward path is programmatic, not classifier-judged: arithmetic tasks with last-integer verification, where the honest 40-step log shows format learned fast and arithmetic at chance. Checkpoints published, including the collapsed one. EVAL_REPORT.md is a published evaluation. The held-out abstention suite run on stock Apache-license models, every number bound to the committed results JSON — SmolLM2-135M-Instruct fabricates on 98.75% of unanswerable prompts, Qwen2.5-0.5B-Instruct abstains at 88.75%. |
| bioprocess-decision-runtime | Bit-exact re-execution spec of a fixed Gemma 3 270M checkpoint with predeclared held-outs, plus a preserved failure where token agreement hid 49 differing logits. |
| qms-ai-system-resume | Fail-closed deployment assessment engine and hardened architecture for AI in regulated quality systems. |
| agent-trajectory-audit | Oversight applied to my own coding-agent sessions: normalizes transcripts to an event stream and flags where actions diverge from stated plans (unbacked test/deploy claims, dropped plan items, out-of-scope writes). trajaudit redteam runs an adversarial battery against its own detectors — two evasion classes found and fixed on first run, one residual documented openly. |
| inference-receipts | Hash-bound receipts for LLM calls (weights, input, settings, output, chain link) with live replay verification. Committed example replays two real SmolLM2-135M calls bit-exact and catches a tampered receipt. |
| eval-harness | Task-spec eval runner with hash-chained result logs; sweep runs one battery across staged checkpoints — the training-run-assessment shape. inspect-export emits real Inspect dataset.jsonl + @task pairs verified under mockllm runs (oracle 100% on all three committed batteries). Committed artifact: a 3-stage SmolLM2 sweep where DPO collapsed abstention to 0.13 while answer generation degenerated at every stage. |
| agent-monitor | The audit moved upstream: a pre-execution policy gate on tool calls. Every allow/flag/block decision lands as a hash-chained event record. |
| model-serve | Serving slice: queued micro-batched inference, shadow comparison of a candidate against production, hash-bound record per response. bench measures real p50/p95/p99 under concurrent load — committed sweep on SmolLM2-135M documents batching's queue-wait vs missing fused compute honestly. |
| agent-observe | Observability layer over the stack: span/trace ingest (flat schema + OTLP-lite), live policy verdicts via agentmon, trajaudit detectors on the same event stream, rendered markdown/HTML trace reports. |
| agent-sandbox | Containment-layer benchmark: scripted tool calls route through the agentmon gate then a realpath filesystem jail, measured separately. Committed battery shows the jail catching what the gate allows (symlink, absolute-path reads) and labels the residual it found (arg aliasing — since fixed in agentmon; non-redirect writes like cp, documented open). |
| receipt-report | Audit documents generated from the other packages' chains — verifies integrity first, then recomputes every reported number from the records. |
| traj-review-ui | Static React + TypeScript review surface for the repo's real artifacts: findings panel with jump-to-event links, event timeline with call/result pairing and filters, and a span-tree view for agent-observe traces. Bundled datasets are committed trajaudit outputs; no backend, labeled as a static demo. |
Lab software and pipelines — mol-ml + lab-informatics
| Repository | What it is |
|---|---|
| comp-tox-pipeline | Reproducible Tox21 evaluation: scaffold-split LR/RF/GIN with calibration, conformal coverage, and applicability-domain conditioning. The NR-ER applicability-domain inversion only surfaced because out-of-domain metrics were measured. |
| lab-instrument-gateway | SCPI-style instrument driver against an emulated bioreactor over TCP: timeouts, reconnects, error-register polling, typed readings to SQLite, FastAPI dashboard, fault injection for failure-path tests. |
| analytics | dbt/DuckDB warehouse over the LIMS registry and lablink capture events. Staging to marts with 65 schema tests plus four domain tests (audit-sequence contiguity, lock signing, alarm-to-reading joins), seeds generated by driving the real services, docs auto-published to Pages. |
| deploy | One-image deployment for the lab-informatics stack: Dockerfile plus compose services for lablink (API on :8000), a LIMS demo run, and the dbt build with docs on :8080. compose run verify executes all three test suites inside the image. |
| dockops | Docking pipeline ops: DUD-E staging, Vina backend, AFDB-vs-PDB structure QC that correctly argues against docking the AF model for EGFR, per-run provenance manifests. |
| vector-db-mcp | Local vector DB over MCP (stdio) with flat/IVF/HNSW scans. Committed benchmark shows exact flat wins at this scale; two construction bugs were caught by recall measurement. |
- scverse/scanpy (open): PR #4383 keeps a user-supplied
hueinsc.pl.violininstead of dropping it or erroring. PR #4385 fixes multi-columngroupbycrashing on non-string observations across the BasePlot family. - OpenADMET/openadmet-models (open): PRs #607, #608, #609, #610. The last exposes
n_jobson the splito-based splitters. - chaidiscovery/chai-lab (open): PR #431 reports pTM as the aggregate score for single-chain inputs, where the ipTM-based headline understated confidence. PR #432 persists the PAE/PDE/pLDDT matrices in Python-mode
scores.npz. - UKGovernmentBEIS/inspect_evals (open): PR #2583 registers inspect-case-bench, a port of CASE-Bench (ICML 2025, context-aware safety judgment vs human majority labels, 900 samples). Full two-model
.evalrun logs included per the register's evidence requirement. - TDC fork: tested fix for silent dataset-name substitution plus an AnnData getter/split API. ProteinMPNN: effective-seed visibility fix filed as issue #154.
The repos are standalone, but a few components are shared deliberately so results cross-check each other:
- ESM-2 is the protein encoder in three places with honestly split outcomes: it drives rescoring in protein-design-ops and the target encoder in dti-fusion, and in active-learning-loop it loses to one-hot on GB1, kept as a committed negative result.
- The GB1 measured landscape is the fitness oracle in both protein-diffusion and active-learning-loop, so the generation-steering and acquisition-enrichment numbers are comparable across repos.
- AnnData/scanpy underlies scrna-qc, organoid-qc, spatial-qc and cytof-qc. The upstream scanpy PRs came from plotting bugs hit in that workflow, not from drive-by contributions.
- Claims bound to computation: the same claims verifier is vendored in three places (oncology-coscientist, statgen, and llm-posttraining's published eval report) with a shared sha256 parity pin so a reviewer can check drift in one command. The bioprocess runtime's replay certificates, protein-design-ops' pinned-seed provenance, and comp-tox's applicability-domain conditioning are the same discipline applied at different layers.
- Shared conventions, vendored not depended on: a common provenance schema (
docs/PROVENANCE.md) is adopted by oncology-coscientist, bioprocess-decision-runtime, dockops, labStackDev and comp-tox-pipeline. A canonical conformal helper with parity tests runs in comp-tox-pipeline, protein-stability-uncertainty and cultivated-meat-multiomic. One content-addressed ESM-2 cache is vendored into protein-design-ops, dti-fusion and active-learning-loop. - Cross-repo flows: lab-instrument-gateway captures feed bioprocess-decision-runtime's
capture-scenarioevaluator. job-watch's repost detection and oncology-coscientist'sembedretrieval mode can both dispatch to vector-db-mcp's E5 embedder while keeping dependency-free defaults. qms-ai-system-resume's deployment instrument has a committed worked assessment of oncology-coscientist (verdict: BLOCKING_FINDINGS, as designed for a research tool).
- Verified computation is not verified science: reproducing a calculation says nothing about whether the model or the hypothesis is right.
- Research and education use only. Nothing here is validated for GMP, clinical or manufacturing decisions.
- Rejected runs, abstentions and failures are part of the evidence record.