Skip to content

Repository files navigation

Pravda

Pravda is a Python library for durable web evidence capture. It drives a remote Playwright browser to preserve rendered HTML, plaintext, full-page screenshots, metadata, and HAR recordings with response bodies. Snapshots are recorded in a SQL database (PostgreSQL or SQLite) and on any fsspec-compatible backend for later inspection or comparison.

Pravda is a library, not a service: it connects directly from the caller's process to the browser, database, and storage backend. Applications own that infrastructure (see Infrastructure).

  • Python 3.12+
  • Browser: a remote Playwright Chromium WebSocket endpoint (headed Chrome under xvfb). The browser is a client connection; Pravda does not launch one.
  • Database: PostgreSQL or SQLite, upgraded to Pravda's schema with the migration helper.
  • Storage: any fsspec URL (local path, s3://, gs://, …) for content-addressed artifacts.

Installation

pip install opensanctions-pravda

Quick start

The application owns the SQLAlchemy engine and an async_sessionmaker configured with expire_on_commit=False; construct a PravdaConfig for the remaining settings and build a long-lived Pravda. Reuse a single instance across captures — each capture opens its own browser connection and database session — and dispose the engine on shutdown.

from sqlalchemy.ext.asyncio import async_sessionmaker, create_async_engine

from pravda import Pravda, PravdaConfig

engine = create_async_engine("postgresql+psycopg://pravda:pravda@localhost:5432/pravda")
sessionmaker = async_sessionmaker(engine, expire_on_commit=False)
config = PravdaConfig(
    browser_ws_url="ws://localhost:3000",
    storage_base_path="./data",
)


async def capture_example():
    pravda = Pravda(config, sessionmaker)
    snapshot = await pravda.snapshot("https://example.com")
    print(snapshot.id, snapshot.http_status, snapshot.rendered_html)

Configuration

PravdaConfig takes two settings, supplied explicitly per instance, plus an application-owned async_sessionmaker passed to Pravda:

  • browser_ws_url — remote Playwright WebSocket URL.
  • storage_base_path — fsspec storage URL, such as ./data, s3://bucket, or gs://bucket.

The session factory must use expire_on_commit=False, because Pravda reads persisted snapshots after commit.

Usage

Capture a page

By default, Pravda navigates to the URL, waits for the normal load state, captures the evidence, and persists the result. Capture and HAR finalization share one wall-clock deadline; persistence is bounded separately:

async def capture_example():
    pravda = Pravda(config, sessionmaker)
    snapshot = await pravda.snapshot("https://example.com")

For custom navigation or interaction, pass an async drive(page, url) callback. The callback owns the initial navigation; Pravda owns capture, persistence, and cleanup, and records the state it leaves behind.

async def drive(page, url):
    await page.goto(url, wait_until="commit")
    await page.wait_for_selector(".results")


async def capture_results():
    pravda = Pravda(config, sessionmaker)
    snapshot = await pravda.snapshot("https://example.com", drive=drive)

page is a real playwright.async_api.Page, so selectors, clicks, form fills, and other Playwright operations are available. Playwright errors and timeouts from drive are persisted as failed snapshots. A recording-context close failure also persists a failed snapshot without artifacts because its HAR could not be finalized. Other callback exceptions propagate and persist nothing.

Chrome is configured to download PDFs instead of opening its viewer. In custom callbacks, page.goto() may therefore raise Download is starting; catch that Playwright error if the download is expected. Pravda recovers the downloaded body into the HAR.

Query history

The configured instance returns all snapshots for an exact URL, newest first:

async def print_history():
    pravda = Pravda(config, sessionmaker)
    history = await pravda.snapshots("https://example.com")
    for snapshot in history:
        print(snapshot.captured_at, snapshot.http_status)

Database migrations

Alembic owns the Pravda schema; the migration scripts ship inside the distribution. There is no migration API: consumers apply the packaged revisions with their own Alembic configuration, addressing the packaged scripts through the package-resource location pravda:migrations. A consumer alembic.ini alongside its own migrations:

[DEFAULT]
prepend_sys_path = %(here)s

[myapp]
script_location = %(here)s/migrations

[pravda]
script_location = pravda:migrations
$ alembic -n myapp upgrade head
$ alembic -n pravda upgrade head

Each environment tracks its own version table (pravda_alembic_version), so Pravda's history stays independent of the consumer's and the two upgrade commands work in either order. The consumer supplies the database connection: sqlalchemy.url in its configuration or PRAVDA_DATABASE_URI in the environment (see pravda/migrations/env.py). Upgrades run the packaged revisions through Alembic (not metadata.create_all) and are safe to repeat (a database already at head is a no-op). There is no downgrade or automatic-startup behavior: run the upgrade where and when you want the schema applied.

Storage

Artifacts are content-addressed files organized under the captured URL's hostname. The public Snapshot resolves its artifact fields to full storage paths: plaintext, rendered_html, and screenshot point at stored files, and each HAR response.content._file resolves to its stored response body. Persisted database fields and HAR values keep their relative, content-addressed names; consumers read the resolved paths directly from the shared fsspec backend. Storage writes are bounded; timeouts and other storage failures propagate, and no snapshot record is persisted.

Infrastructure

Applications own the external infrastructure Pravda talks to; Pravda does not launch or manage it:

  • Browser — a remote Playwright Chromium server exposed over WebSocket. Applications provide their own, such as the Playwright Docker image running headed Chrome under xvfb, or a hosted browser service.
  • Database — a PostgreSQL or SQLite database the application provisions, opens an async engine for, and upgrades with the packaged migrations. The application owns the engine and session factory.
  • Storage — an fsspec backend the application points at via storage_base_path.

Development

Requires uv and a remote Playwright browser endpoint. Tests run against an in-memory SQLite database, so no database service is needed. Environment variables live in .env; uv does not read it automatically:

# Install dependencies
uv sync

# Local environment (once)
cp .env.example .env

# Validate
uv run --env-file .env pytest
uv run ruff check .
uv run ruff format --check .

The migration scripts live inside the package at pravda/migrations. After changing models in pravda/db.py, the developer alembic command reads PRAVDA_DATABASE_URI from .env and points at the packaged scripts via alembic.ini:

uv run --env-file .env alembic upgrade head
uv run --env-file .env alembic revision --autogenerate -m "describe the change"

License

MIT — see LICENSE.

About

The evidence layer

Resources

Stars

6 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages