Skip to content
ATaylorAerospacePublic

About

Autonomous AI agents that batch analyze legal contracts for risk, missing provisions, and cross document inconsistencies. Built on AWS Bedrock + Strands Agents SDK. This app is NOT legal advice, just a really fast first pass.

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

101 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

factorai

โš–๏ธ Factor AI - Agentic AI Legal Due Diligence Platform ๐Ÿ”

License: MIT Python TypeScript AWS Dataset Contact

Autonomous AI agents that batch analyze legal contracts for missing provisions, unusual terms, and risk flags - powered by AWS Strands Agents SDK and Amazon Bedrock AgentCore.

๐Ÿšง Status: Core agents stable ยท Dashboard live ยท AgentCore deployment ready


๐Ÿค” The Problem

M&A and financing due diligence requires reviewing 10โ€“100+ contracts to identify missing clauses, non-standard terms, and cross-document inconsistencies. Manual review is slow, expensive, and error-prone.


๐Ÿ’ก The Solution

Factor AI deploys a system of autonomous AI agents that collaboratively analyze batches of legal documents:

  • Ingest PDF, DOCX, and TXT files, extracting and chunking provisions (unreadable files are skipped with a reason, never silently dropped)
  • Classify each contract (NDA, lease, loan, merger, employment, license, supply) to select the right checklist
  • Detect provision types using pattern matching and AI classification
  • Score risk levels against configurable rubrics
  • Identify missing critical clauses via gap analysis
  • Compare provisions across documents for inconsistencies
  • Generate structured risk reports with Excel and HTML export
  • Guard every reasoning step with a financial circuit breaker and Arize Phoenix telemetry

๐Ÿ›๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    FACTOR AGENT SYSTEM                   โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
โ”‚  โ”‚            Coordinator Agent                       โ”‚  โ”‚
โ”‚  โ”‚  Receives document batch, plans analysis strategy, โ”‚  โ”‚
โ”‚  โ”‚  delegates to specialist agents, assembles report  โ”‚  โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
โ”‚             โ”‚          โ”‚          โ”‚                      โ”‚
โ”‚    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”         โ”‚
โ”‚    โ”‚ Ingestion โ”‚ โ”‚ Analysis  โ”‚ โ”‚ Knowledge  โ”‚         โ”‚
โ”‚    โ”‚   Agent   โ”‚ โ”‚   Agent   โ”‚ โ”‚   Agent    โ”‚         โ”‚
โ”‚    โ”‚ โ€ข Parse   โ”‚ โ”‚ โ€ข Detect  โ”‚ โ”‚ โ€ข RAG      โ”‚         โ”‚
โ”‚    โ”‚ โ€ข Chunk   โ”‚ โ”‚ โ€ข Score   โ”‚ โ”‚ โ€ข Classify โ”‚         โ”‚
โ”‚    โ”‚ โ€ข Extract โ”‚ โ”‚ โ€ข Gaps    โ”‚ โ”‚ โ€ข Citationsโ”‚         โ”‚
โ”‚    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ”‚ โ€ข Compare โ”‚ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜         โ”‚
โ”‚                  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                          โ”‚
โ”‚    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                                      โ”‚
โ”‚    โ”‚  Reporting   โ”‚                                     โ”‚
โ”‚    โ”‚    Agent     โ”‚                                     โ”‚
โ”‚    โ”‚ โ€ข Reports   โ”‚                                     โ”‚
โ”‚    โ”‚ โ€ข Excel     โ”‚                                     โ”‚
โ”‚    โ”‚ โ€ข HTML      โ”‚                                     โ”‚
โ”‚    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                                      โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
โ”‚  โ”‚  FINANCIAL-GUARDRAIL & TELEMETRY HARNESS           โ”‚  โ”‚
โ”‚  โ”‚  Circuit Breaker โ”‚ Budget Tracker โ”‚ Loop Detector  โ”‚  โ”‚
โ”‚  โ”‚  GuardedBedrockModel โ†’ audits token I/O per step   โ”‚  โ”‚
โ”‚  โ”‚  OpenTelemetry/OTLP โ†’ Arize Phoenix (traces UI)    โ”‚  โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
โ”‚                                                         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
โ”‚  โ”‚  Bedrock AgentCore Runtime โ”‚ Memory โ”‚ Gateway      โ”‚  โ”‚
โ”‚  โ”‚  Policy โ”‚ Observability โ”‚ Identity                 โ”‚  โ”‚
โ”‚  โ”‚  Amazon Bedrock (Foundation Models)                โ”‚  โ”‚
โ”‚  โ”‚  S3 โ”‚ DynamoDB โ”‚ CloudWatch โ”‚ Cognito              โ”‚  โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿค– Agent System

Agent Role Tools Status
๐ŸŽฏ Coordinator Orchestrates pipeline, delegates tasks, assembles results ingest_documents, analyze_provisions, search_knowledge, generate_report โœ… Stable
๐Ÿ“„ Ingestion Parses PDF/DOCX, chunks into provisions parse_pdf, parse_docx, chunk_provisions โœ… Stable
๐Ÿ” Analysis Detects types, scores risk, finds gaps, compares detect_provision_type, score_risk, find_gaps, compare_across_documents โœ… Stable
๐Ÿ“š Knowledge Searches synthetic KB, classifies domains, extracts citations search_synthetic_knowledge, classify_domain, extract_citations โœ… Stable
๐Ÿ“Š Reporting Builds reports, exports Excel/HTML build_risk_report, export_excel, export_html โœ… Stable

โœจ Core Capabilities

  • ๐Ÿ” Provision Detection - 14 provision types identified via anchor patterns; section headings (GOVERNING LAW., Section 5., ARTICLE VII) stay attached to their clause
  • ๐Ÿท๏ธ Contract-Type Detection - Whole-word signals pick the NDA / lease / loan / merger / employment / license / supply checklist for each document
  • ๐Ÿ“Š Risk Scoring - Configurable rubrics with weighted signals (0โ€“10 scale)
  • โš ๏ธ Gap Analysis - Standard checklists for NDAs, leases, loans, mergers, employment, license, and supply agreements
  • ๐Ÿ”„ Cross-Document Comparison - Inconsistency detection across governing law, liability caps, termination terms, compared document-to-document (never clause-to-clause within one file)
  • ๐Ÿ—‚๏ธ Per-Document Attribution - Every risk score and gap names its source file, clause type, and a text excerpt, in the dashboard and in exports
  • ๐Ÿงฏ Resilient Batches - Corrupt, password-protected, or scanned (image-only) files are reported as not analyzed while the rest of the batch completes
  • ๐Ÿ“š RAG Knowledge Search - Synthetic legal knowledge base (Taylor658/synthetic-legal)
  • ๐Ÿ“‹ Structured Reports - Executive summary, risk assessment, gap analysis, comparison results
  • ๐Ÿ“ฅ Export - One-click download of Excel (with disclaimer tab) and HTML (with disclaimers on every page)
  • โšก SSE Streaming - Real-time, per-document progress via Server-Sent Events; parsing and scoring run off the event loop so the server stays responsive
  • ๐Ÿ’ฐ Financial Circuit Breaker - Per-session token budget enforced at every LLM reasoning step; agents are hard-halted before runaway cost
  • ๐Ÿ” Reasoning Loop Detection - Sliding-window detector halts agents stuck in repetitive, high-cost cycles
  • ๐Ÿ“ก Phoenix Telemetry - OpenTelemetry traces exported to a self-hosted Arize Phoenix instance for per-step token auditing
  • ๐Ÿ›ก๏ธ Session Isolation - Cedar policies enforce per-user data access
  • ๐Ÿ”’ Upload Validation - File type enforcement (PDF, DOCX, TXT) with size limits enforced while streaming to disk
  • ๐ŸŒ Production CORS - Explicit origin allow-list; a wildcard (*) is honoured in development only and is never applied in production, and credentials are never paired with a wildcard
  • ๐Ÿงน Automatic Cleanup - Uploads are removed after analysis; sessions, reports, and exports expire automatically or on DELETE
  • ๐Ÿš€ Non-Blocking Telemetry - Token auditing runs synchronously on each span (so the breaker can still halt the agent), while OTLP export to Phoenix is batched off the agent's call path โ€” an unreachable Phoenix never stalls analysis
  • ๐Ÿงต Thread-Safe Knowledge Base - Vector-store initialisation is lock-guarded, so concurrent first requests in FastAPI's threadpool can't open duplicate ChromaDB clients
  • ๐ŸŽ๏ธ Hot-Path Efficiency - Detection and chunking regexes are compiled once at import; per-clause logging is DEBUG-level; parsers hold the document text once (per-page detail is opt-in via include_details=True)
  • ๐Ÿ”ค Portable Exports - HTML reports are always written as UTF-8, independent of the host's locale

๐Ÿ“ Repository Layout

factor/
โ”œโ”€โ”€ src/factor/              # Python backend
โ”‚   โ”œโ”€โ”€ agents/              # Strands Agent definitions
โ”‚   โ”œโ”€โ”€ harness/             # Financial-guardrail & Phoenix telemetry harness
โ”‚   โ”œโ”€โ”€ tools/               # @tool decorated functions + contract-type detection
โ”‚   โ”œโ”€โ”€ knowledge/           # ChromaDB vector store + dataset loader
โ”‚   โ”œโ”€โ”€ models/              # Pydantic data models
โ”‚   โ”œโ”€โ”€ aws/                 # Bedrock, AgentCore, S3, Cognito
โ”‚   โ”œโ”€โ”€ reporting/           # HTML report templates
โ”‚   โ”œโ”€โ”€ db/                  # Session store (thread-safe)
โ”‚   โ”œโ”€โ”€ app.py               # FastAPI + SSE streaming
โ”‚   โ””โ”€โ”€ config.py            # pydantic-settings
โ”œโ”€โ”€ src/frontend/            # React 18 + TypeScript + Vite
โ”‚   โ”œโ”€โ”€ nginx.conf           # Production proxy: dashboard + /api โ†’ factor-api
โ”‚   โ””โ”€โ”€ src/
โ”‚       โ”œโ”€โ”€ components/      # Upload, Analysis, Report, shared
โ”‚       โ”œโ”€โ”€ hooks/           # useUpload, useAnalysis, useAgentStream
โ”‚       โ”œโ”€โ”€ api/             # API client
โ”‚       โ””โ”€โ”€ types/           # TypeScript types
โ”œโ”€โ”€ tests/                   # pytest test suite
โ”œโ”€โ”€ scripts/                 # Seed KB, generate samples, deploy, benchmark
โ”œโ”€โ”€ policies/                # Cedar policy files
โ”œโ”€โ”€ data/                    # Provision definitions, risk rubric, samples
โ”œโ”€โ”€ infra/                   # AWS CDK stacks
โ”œโ”€โ”€ docker-compose.yml       # Self-hosted Arize Phoenix (telemetry)
โ””โ”€โ”€ docker/                  # API + Frontend Dockerfiles, docker-compose

๐Ÿ’ฐ Financial-Guardrail & Telemetry Harness

When agents autonomously batch-analyze 100+ legal documents, the loop of reading, extracting, and comparing can lead to runaway token consumption. The harness wraps Bedrock AgentCore execution to audit token input/output at every discrete reasoning step and acts as a financial circuit breaker โ€” ensuring the operating cost of the AI never outpaces the value of the analysis.

How it works

Agent step โ†’ GuardedBedrockModel โ†’ CircuitBreaker.check()
                                       โ”œโ”€โ”€ SessionBudget   (token cost accounting)
                                       โ””โ”€โ”€ LoopDetector    (repetitive-cycle detection)
                                       โ†“ trip โ†’ BudgetExceededError / ReasoningLoopError

OpenTelemetry span (gen_ai.usage.*) โ”€โ”ฌโ†’ GuardrailSpanProcessor  (sync: feeds the breaker, no I/O)
                                     โ””โ†’ BatchSpanProcessor      (async: OTLP export โ†’ Arize Phoenix)

The two span processors are deliberately separate: guardrail accounting must run synchronously on span end so a tripped breaker propagates up the agent's call stack, whereas the network export is batched on a background thread so an unreachable Phoenix endpoint can never block a reasoning step.

Component Responsibility
SessionBudget Accumulates input/output token cost per session (Sonnet pricing: $3 / $15 per 1M)
LoopDetector Sliding-window detection of repeated reasoning actions
CircuitBreaker Combines budget + step-limit + loop checks; raises to hard-halt the agent
GuardedBedrockModel Proxy around BedrockModel that checks the breaker before every invocation
FinancialGuardrail Singleton registry of per-session circuit breakers
GuardrailSpanProcessor Extracts token counts from OTel spans and feeds the breaker โ€” synchronous, no network I/O
BatchSpanProcessor + OTLPSpanExporter Queues spans and ships them to Phoenix in the background

When a breaker trips, the /api/v1/analyze SSE stream emits a guardrail_halt event with the session's cost, step count, and trip reason.

๐Ÿ’ก What counts as a step: the breaker meters LLM reasoning steps (model calls and their token spend). The deterministic /api/v1/analyze pipeline โ€” parsing, regex detection, and rubric scoring โ€” makes no model calls, so it is never counted against GUARDRAIL_MAX_STEPS. A batch of 100 contracts with thousands of clauses runs to completion.

Start Phoenix (self-hosted)

# Launch the Phoenix telemetry UI + OTLP collector
docker compose up -d phoenix

# Phoenix UI:        http://localhost:6006
# OTLP gRPC:         localhost:4317

Configuration

Env Var Default Description
PHOENIX_ENABLED true Enable OTLP export to Phoenix
PHOENIX_OTLP_ENDPOINT http://localhost:6006/v1/traces Phoenix trace collector endpoint
GUARDRAIL_ENABLED true Enable the financial circuit breaker
GUARDRAIL_SESSION_BUDGET_USD 5.0 Hard cost ceiling per analysis session
GUARDRAIL_MAX_STEPS 200 Maximum LLM reasoning steps per session
GUARDRAIL_LOOP_WINDOW 10 Sliding window size for loop detection
GUARDRAIL_LOOP_THRESHOLD 5 Repeat count within window that trips a loop
GUARDRAIL_INPUT_COST_PER_1M 3.0 Input token price (USD per 1M)
GUARDRAIL_OUTPUT_COST_PER_1M 15.0 Output token price (USD per 1M)

๐Ÿ Getting Started

Prerequisites

  • โœ… Python 3.11+
  • โœ… Node.js 20+
  • โœ… AWS account with Bedrock access
  • โœ… AWS CLI configured (aws configure)

Installation

# Clone the repository
git clone https://github.com/ATaylorAerospace/Factor-AI.git
cd Factor-AI

# Create virtual environment
python -m venv .venv && source .venv/bin/activate

# Install Python dependencies
pip install -r requirements.txt

# โ€ฆor install Factor as an editable package (with dev extras)
pip install -e ".[dev]"

# Seed the knowledge base
python scripts/seed_knowledge_base.py

# Start the API server
uvicorn src.factor.app:app --reload --port 8000

Frontend

cd src/frontend
npm install
npm run dev

Run Tests

pytest tests/ -v --cov=src/factor

๐Ÿ› ๏ธ Technology Stack

Layer Technology Purpose
๐Ÿค– Agent Framework Strands Agents SDK Model-driven agents with @tool
๐Ÿง  Foundation Model Amazon Bedrock (Anthropic Sonnet) Reasoning + tool-use
โšก Agent Runtime Bedrock AgentCore Runtime Serverless execution
๐Ÿ’พ Agent Memory Bedrock AgentCore Memory Persistent context
๐Ÿ”ง Agent Gateway Bedrock AgentCore Gateway MCP tool access
๐Ÿ›ก๏ธ Agent Policy Bedrock AgentCore Policy (Cedar) Action boundaries
๐Ÿ“Š Observability Bedrock AgentCore + OTEL Tracing + dashboards
๐Ÿ“ก AI Telemetry Arize Phoenix (OTLP, self-hosted) Per-step token auditing + trace UI
๐Ÿ’ฐ Cost Guardrail Custom circuit breaker harness Budget + loop-detection hard halt
๐Ÿ” Identity Bedrock AgentCore Identity / Cognito Authentication
๐Ÿ”ข Embeddings sentence-transformers Vector embeddings
๐Ÿ“š Vector Store ChromaDB (local) / Bedrock KB (prod) Dataset indexing
โ˜๏ธ Storage Amazon S3 Document storage
๐Ÿ—„๏ธ Metadata Amazon DynamoDB Session + results
๐Ÿ“„ Doc Parsing PyMuPDF + python-docx + pdfplumber Text extraction
๐Ÿ–ฅ๏ธ Frontend React 18 + TypeScript + Vite + Tailwind Dashboard
๐Ÿ“‹ Export openpyxl + Jinja2 Reports
๐Ÿ—๏ธ IaC AWS CDK (Python) Infrastructure
๐Ÿ”„ CI/CD GitHub Actions Quality + deploy

โš ๏ธ Synthetic Dataset Disclaimer

โš ๏ธ CRITICAL: THIS IS A SYNTHETIC DATASET - ALL CONTENT IS ARTIFICIALLY GENERATED

Factor's knowledge base is powered by the Taylor658/synthetic-legal dataset on HuggingFace (140,000 rows, MIT License).

ALL text in this dataset is synthetically generated and IS NOT legally accurate. All citations, statutes, case references, legal problems, verified solutions, and pairings are synthetic constructs created through template-based randomization. No citations, statutes, or case references in this dataset are real.

This dataset exists for research, experimentation, and model training only.


๐Ÿ“Š API Endpoints

Method Endpoint Description
POST /api/v1/analyze Upload documents (PDF, DOCX, TXT) + stream agentic analysis
GET /api/v1/sessions/{id} Session status (processing ยท completed ยท halted ยท failed) + results
DELETE /api/v1/sessions/{id} Delete a session, its report, and exported files
GET /api/v1/sessions/{id}/trace Agent reasoning trace
GET /api/v1/reports/{session_id} Structured report
GET /api/v1/reports/{session_id}/export?format=excel|html Download Excel/HTML as a file attachment
GET /api/v1/sessions/{id}/budget Real-time guardrail budget + token status
GET /api/v1/guardrail/status Guardrail config + all active sessions
GET /api/v1/knowledge/search Search synthetic KB
GET /api/v1/knowledge/domains List legal domains
GET /api/v1/health Health check

๐Ÿ“ก Analysis Event Stream

POST /api/v1/analyze answers with a Server-Sent Events stream. Events are separated by a blank line (\r\n\r\n); a single event โ€” especially the report โ€” can span many network reads, so clients must buffer until the blank line before parsing.

Event When Key fields
session First event session_id
status Stage changes stage: ingestion โ†’ analysis โ†’ reporting
progress After each document in each stage stage, document, provisions_found / provisions_scored, gaps_found
document_skipped A file could not be read (corrupt, encrypted, or scanned with no text) document, reason
guardrail Breaker initialized / completed budget and step status
report Analysis finished full report incl. documents and skipped_documents
done Stream complete session_id
guardrail_halt Circuit breaker tripped reason, steps, total_cost_usd
error Unexpected failure (session marked failed) message, detail

โš™๏ธ Session Settings

Env Var Default Description
FACTOR_SESSION_TTL_HOURS 24 Sessions, reports, and exports older than this are removed
FACTOR_MAX_SESSIONS 500 Oldest sessions are evicted beyond this count
FACTOR_MAX_UPLOAD_MB 50 Per-file upload limit
FACTOR_MAX_BATCH_SIZE 100 Files per analysis batch
FACTOR_ENV development development ยท staging ยท production โ€” controls CORS strictness
FACTOR_ALLOWED_ORIGINS * Comma-separated CORS origins. * is accepted only outside production; a production deployment must list explicit origins or no cross-origin requests are allowed

๐Ÿงช Testing

# Run all tests with coverage
pytest tests/ -v --cov=src/factor

# Run specific test modules
pytest tests/test_tools/ -v          # Tool tests
pytest tests/test_agents/ -v         # Agent tests
pytest tests/test_knowledge/ -v      # Knowledge base tests

Tests cover:

  • โœ… Each @tool function independently with assertions
  • โœ… Agent creation with mocked Bedrock responses
  • โœ… Synthetic dataset loading and metadata
  • โœ… ChromaDB vector store operations
  • โœ… Provision detection, scoring, and gap analysis
  • โœ… Cross-document comparison and inconsistency detection
  • โœ… Domain classification across 13 legal domains
  • โœ… Citation extraction (cases, statutes, regulations)
  • โœ… Report building, Excel export, and HTML export
  • โœ… Financial guardrail: budget accounting, loop detection, circuit breaker trips
  • โœ… Batch pipeline regressions: 250-clause batches, same-name uploads, unreadable and scanned files, unexpected-error events, file download, and session deletion
  • โœ… Chunking of all-caps, Section N., and ARTICLE headings; whole-word contract-type detection
  • โœ… Session expiry and eviction
  • โœ… CORS settings: wildcard allowed in development, dropped in production, credentials never paired with *
  • โœ… HTML export is byte-for-byte UTF-8; parsers return per-page detail only when asked
  • โœ… Guardrail span processor feeds the breaker without performing any export
  • โœ… Citation regex: party names never swallow the preceding sentence; linear-time on long capitalised prose
  • โœ… All outputs label synthetic content

๐Ÿงฎ 142 tests run in CI on Python 3.11 and 3.12.


๐Ÿš€ Deployment

Docker

# Start both API and frontend services
docker compose -f docker/docker-compose.yml up --build

# Dashboard: http://localhost:3000   (nginx proxies /api/* to the API container)
# API:       http://localhost:8000

๐Ÿ”Œ The frontend container runs nginx, which serves the built dashboard and proxies /api/* to factor-api with response buffering disabled, so analysis progress streams live.

AWS CDK (AgentCore)

# Deploy infrastructure
cd infra && cdk deploy --all

# Deploy agent configuration
python scripts/deploy_agentcore.py --env production

๐Ÿฉน Reliability & Accuracy Fixes

The latest release hardens the batch pipeline end to end. Each fix below is covered by a regression test.

# Area Before After
1 ๐Ÿ’ฐ Guardrail Every clause scored counted as a reasoning step, so batches of ~10 contracts hit GUARDRAIL_MAX_STEPS and halted with no report; the halt message always said steps=0 Only LLM steps are metered; the halt message reports the real step count
2 ๐Ÿ–ฅ๏ธ Dashboard streaming Reports larger than one network read were dropped, leaving a blank screen; progress vanished on the first event; halts and errors were never shown Events are buffered until complete; live per-stage progress; clear error panel with Start over
3 ๐Ÿงฏ Unreadable files One corrupt PDF or .doc aborted the whole batch; scanned PDFs were reported as low risk while "missing" every clause Bad files are listed under Not analyzed with a reason and the batch continues; .doc is rejected up front; an all-unreadable batch reports unknown risk
4 ๐Ÿ” Clause detection Section 1. headings were never split and all-caps headings were cut off, so real clauses showed up as gaps; "indemnity" wasn't recognised Headings stay with their clause; Section/Article split in any case; indemnity detected
5 ๐Ÿ—‚๏ธ Document attribution Same-name uploads overwrote each other; results showed random IDs, with no document column Every upload is kept (Agreement.txt (2)); results name their file, clause type, and excerpt
6 โšก Server responsiveness Parsing and scoring ran on the event loop โ€” a /health check waited 2.4 s during a 100-contract batch Work runs in worker threads โ€” worst /health latency 0.28 s during the same batch
7 ๐Ÿ”„ Comparison Two clauses in one document were reported as a cross-document inconsistency; comparisons never appeared in the dashboard Documents are compared to each other by name; a new Cross-Document Inconsistencies table
8 ๐Ÿ“ฅ Export Wrote to the server's disk and showed the server path in a popup Real Excel/HTML downloads, available from the dashboard
9 ๐Ÿณ Docker The dashboard container couldn't reach the API nginx serves the UI and proxies /api
10 ๐Ÿท๏ธ Contract types The API always used the generic checklist; the agent classified a lease mentioning "calendar" as an NDA Whole-word type detection drives the NDA/lease/loan/... checklists
11 ๐Ÿงน Retention Sessions and reports were kept in memory and on disk forever TTL + size-capped eviction and DELETE /api/v1/sessions/{id}

๐Ÿ”ง Round 2 โ€” Packaging, Security & Efficiency

A follow-up code review swept the whole backend. Every item below is covered by a regression test, and the suite grew from 105 to 142 tests.

๐Ÿ› Bugs fixed

# Area Before After
12 ๐Ÿ“ฆ Packaging pyproject.toml carried both a PEP 639 license = "MIT" expression and the legacy license classifier, so modern setuptools refused to build โ€” pip install -e . failed outright Classifier removed; the package builds and installs as pip install -e ".[dev]"
13 ๐Ÿงช Orphaned tests tests/test_tools/test_citations had no .py extension, so its six tests were never collected Renamed; all citation tests run in CI
14 ๐ŸŒ CORS The production guard was dead code โ€” with the default FACTOR_ALLOWED_ORIGINS=*, production still answered with Access-Control-Allow-Origin: *; credentials were paired with the wildcard in dev (invalid per spec) * is honoured only outside production; allow_credentials is derived from the origin list and never combined with *
15 ๐Ÿ”ค HTML export Written with the host's locale encoding while declaring UTF-8 and containing โš ๏ธ โ€” UnicodeEncodeError on Windows Always written as UTF-8
16 ๐Ÿ“„ PDF parsing doc.close() was not in a finally; a malformed page leaked the PyMuPDF handle Parsed inside a with fitz.open(...) context manager
17 ๐Ÿ“š Vector store get_collection() had no lock; two first requests in the FastAPI threadpool could open two PersistentClients on the same directory Creation is lock-guarded
18 ๐Ÿ“‘ Citation regex Open-ended party-name pattern swallowed the preceding sentence into the plaintiff and backtracked quadratically on capitalised prose Party names are bounded runs of capitalised tokens; reporter matching is token-based

โšก Efficiency

# Area Before After
19 ๐Ÿ“ก Telemetry GuardrailSpanProcessor extended SimpleSpanProcessor, performing a blocking HTTP export per span on the agent's hot path โ€” with Phoenix enabled by default and unreachable, every span stalled Guardrail accounting is a pure synchronous processor; export moved to a BatchSpanProcessor on a background thread
20 ๐Ÿ” Detection & chunking ~50 regexes re-parsed via re.findall / re.search for every clause; anchor labels rebuilt per call All patterns compiled once at import
21 ๐Ÿ“ Logging Two INFO lines per clause โ€” thousands of lines on a 100-contract batch Per-clause logs are DEBUG; per-document summaries stay at INFO
22 ๐Ÿง  Parser memory page_details / paragraph_details duplicated the entire document text even though the pipeline only reads text Detail is opt-in (include_details=True), halving peak memory per upload

๐Ÿงน Cleanups

  • FastAPI startup migrated from the deprecated @app.on_event("startup") to a lifespan handler (no more deprecation warning in test runs)
  • Deleted the empty stray docs/test.py
  • Removed the Coordinator's pass-through _infer_doc_type wrapper and the guardrail's no-op try/except โ€ฆ raise
  • search_synthetic_knowledge now delegates to vectorstore.query instead of re-implementing hit construction
  • init_phoenix_tracing keeps the guardrail processor across repeat calls instead of returning None

๐Ÿ™ Contributing

Contributions are welcome! Please see the issue templates for bug reports and feature requests.

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'feat: add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

๐Ÿ‘ค Author

A Taylor ยท 2026 Contact


๐Ÿ“„ License

MIT ยฉ 2026 A Taylor See LICENSE for details.

About

Autonomous AI agents that batch analyze legal contracts for risk, missing provisions, and cross document inconsistencies. Built on AWS Bedrock + Strands Agents SDK. This app is NOT legal advice, just a really fast first pass.

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages