Pravda is a Python library for durable web evidence capture. It drives a remote Playwright browser to preserve rendered HTML, plaintext, full-page screenshots, metadata, and HAR recordings with response bodies. Snapshots are recorded in a SQL database (PostgreSQL or SQLite) and on any fsspec-compatible backend for later inspection or comparison.
Pravda is a library, not a service: it connects directly from the caller's process to the browser, database, and storage backend. Applications own that infrastructure (see Infrastructure).
- Python 3.12+
- Browser: a remote Playwright Chromium WebSocket endpoint (headed Chrome under xvfb). The browser is a client connection; Pravda does not launch one.
- Database: PostgreSQL or SQLite, upgraded to Pravda's schema with the migration helper.
- Storage: any fsspec URL (local path,
s3://,gs://, …) for content-addressed artifacts.
pip install opensanctions-pravdaThe application owns the SQLAlchemy engine and an async_sessionmaker
configured with expire_on_commit=False; construct a
PravdaConfig for the remaining settings and build a
long-lived Pravda. Reuse a single instance across captures — each capture
opens its own browser connection and database session — and dispose the engine
on shutdown.
from sqlalchemy.ext.asyncio import async_sessionmaker, create_async_engine
from pravda import Pravda, PravdaConfig
engine = create_async_engine("postgresql+psycopg://pravda:pravda@localhost:5432/pravda")
sessionmaker = async_sessionmaker(engine, expire_on_commit=False)
config = PravdaConfig(
browser_ws_url="ws://localhost:3000",
storage_base_path="./data",
)
async def capture_example():
pravda = Pravda(config, sessionmaker)
snapshot = await pravda.snapshot("https://example.com")
print(snapshot.id, snapshot.http_status, snapshot.rendered_html)PravdaConfig takes two settings, supplied explicitly per instance, plus an
application-owned async_sessionmaker passed to Pravda:
browser_ws_url— remote Playwright WebSocket URL.storage_base_path— fsspec storage URL, such as./data,s3://bucket, orgs://bucket.
The session factory must use expire_on_commit=False, because Pravda reads
persisted snapshots after commit.
By default, Pravda navigates to the URL, waits for the normal load state,
captures the evidence, and persists the result. Capture and HAR finalization
share one wall-clock deadline; persistence is bounded separately:
async def capture_example():
pravda = Pravda(config, sessionmaker)
snapshot = await pravda.snapshot("https://example.com")For custom navigation or interaction, pass an async drive(page, url)
callback. The callback owns the initial navigation; Pravda owns capture,
persistence, and cleanup, and records the state it leaves behind.
async def drive(page, url):
await page.goto(url, wait_until="commit")
await page.wait_for_selector(".results")
async def capture_results():
pravda = Pravda(config, sessionmaker)
snapshot = await pravda.snapshot("https://example.com", drive=drive)page is a real playwright.async_api.Page, so selectors, clicks, form
fills, and other Playwright operations are available. Playwright errors and
timeouts from drive are persisted as failed snapshots. A recording-context
close failure also persists a failed snapshot without artifacts because its HAR
could not be finalized. Other callback exceptions propagate and persist nothing.
Chrome is configured to download PDFs instead of opening its viewer. In custom
callbacks, page.goto() may therefore raise Download is starting; catch
that Playwright error if the download is expected. Pravda recovers the
downloaded body into the HAR.
The configured instance returns all snapshots for an exact URL, newest first:
async def print_history():
pravda = Pravda(config, sessionmaker)
history = await pravda.snapshots("https://example.com")
for snapshot in history:
print(snapshot.captured_at, snapshot.http_status)Alembic owns the Pravda schema; the migration scripts ship inside the
distribution. There is no migration API: consumers apply the packaged
revisions with their own Alembic configuration, addressing the packaged
scripts through the package-resource location pravda:migrations. A
consumer alembic.ini alongside its own migrations:
[DEFAULT]
prepend_sys_path = %(here)s
[myapp]
script_location = %(here)s/migrations
[pravda]
script_location = pravda:migrations$ alembic -n myapp upgrade head
$ alembic -n pravda upgrade headEach environment tracks its own version table (pravda_alembic_version),
so Pravda's history stays independent of the consumer's and the two
upgrade commands work in either order. The consumer supplies the database
connection: sqlalchemy.url in its configuration or PRAVDA_DATABASE_URI in
the environment (see pravda/migrations/env.py). Upgrades run the packaged
revisions through Alembic (not metadata.create_all) and are safe to
repeat (a database already at head is a no-op). There is no downgrade or
automatic-startup behavior: run the upgrade where and when you want the
schema applied.
Artifacts are content-addressed files organized under the captured URL's
hostname. The public Snapshot resolves its artifact fields to full storage
paths: plaintext, rendered_html, and screenshot point at stored files,
and each HAR response.content._file resolves to its stored response body.
Persisted database fields and HAR values keep their relative, content-addressed
names; consumers read the resolved paths directly from the shared fsspec
backend. Storage writes are bounded; timeouts and other storage failures
propagate, and no snapshot record is persisted.
Applications own the external infrastructure Pravda talks to; Pravda does not launch or manage it:
- Browser — a remote Playwright Chromium server exposed over WebSocket. Applications provide their own, such as the Playwright Docker image running headed Chrome under xvfb, or a hosted browser service.
- Database — a PostgreSQL or SQLite database the application provisions, opens an async engine for, and upgrades with the packaged migrations. The application owns the engine and session factory.
- Storage — an fsspec backend the application points at via
storage_base_path.
Requires uv and a remote Playwright browser
endpoint. Tests run against an in-memory SQLite database, so no database
service is needed. Environment variables live in .env; uv does not read
it automatically:
# Install dependencies
uv sync
# Local environment (once)
cp .env.example .env
# Validate
uv run --env-file .env pytest
uv run ruff check .
uv run ruff format --check .The migration scripts live inside the package at pravda/migrations. After
changing models in pravda/db.py, the developer alembic command reads
PRAVDA_DATABASE_URI from .env and points at the packaged scripts via
alembic.ini:
uv run --env-file .env alembic upgrade head
uv run --env-file .env alembic revision --autogenerate -m "describe the change"MIT — see LICENSE.