Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
72 commits
Select commit Hold shift + click to select a range
7939bb6
feat: full GTFS history backfill with partitioned silver layer (#19)
VictorF13 Aug 1, 2026
0bf97fb
fix: tolerate trailing t suffix in AFC dump filenames (#20)
VictorF13 Aug 1, 2026
0429ac6
fix: harden AVL ingestion against raw data quirks (#22)
VictorF13 Aug 2, 2026
4aa8c7e
fix: allow null validator_id in AFC ingestion (#23)
VictorF13 Aug 2, 2026
0afce5a
feat: support 2018's whole-month AVL raw file layout (#24)
VictorF13 Aug 2, 2026
adf00a0
feat: support 2015-2019 legacy GTFS export naming schemes (#25)
VictorF13 Aug 2, 2026
2b09502
feat: ingest device and legacy vehicle dictionary sources into bronze…
VictorF13 Aug 3, 2026
58f3588
feat: load vehicle dictionary family bronze sources into silver (#27)
VictorF13 Aug 3, 2026
96abc08
feat: rename vehicle dictionary silver tables to dictionary_* prefix …
VictorF13 Aug 3, 2026
a66c4bc
feat: treat vehicle_dictionary as a static source, same as the rest (…
VictorF13 Aug 3, 2026
4c91a88
chore: remove gold dbt layer (#30)
VictorF13 Aug 7, 2026
50c36bf
refactor: rename SILVER_* environment variables to DB_*
VictorF13 Aug 14, 2026
f5a9ac6
chore: add jupyter and ipykernel as dev dependencies
VictorF13 Aug 14, 2026
a91ce2f
chore: ignore print statements in notebook lint rules
VictorF13 Aug 14, 2026
d908b5d
docs: require running prek before considering a change done
VictorF13 Aug 14, 2026
2a98411
feat: add Trip Validity model dataset build notebooks
VictorF13 Aug 14, 2026
7f7ecb5
chore: gitignore pytest cache, coverage artifacts, notebook checkpoin…
VictorF13 Aug 15, 2026
a8ff387
docs: fix stale opa credential example in remote-access guide
VictorF13 Aug 15, 2026
26cf757
fix: restrict Postgres/Adminer to bind only on Tailscale, not 0.0.0.0
VictorF13 Aug 15, 2026
d8f8117
fix: stop requiring raw_data_root to exist at Settings construction time
VictorF13 Aug 15, 2026
dbd31a5
feat: add covering composite indexes for AVL vehicle/device lookups
VictorF13 Aug 15, 2026
d083513
feat: match AFC buses to AVL vehicles and materialize trip positions
VictorF13 Aug 15, 2026
7d84c3e
feat: add AVL match info and trip timestamps to the final dataset
VictorF13 Aug 15, 2026
581ca34
docs: add a usage example for joining GTFS shape + AVL positions
VictorF13 Aug 15, 2026
5279de0
feat: add active learning labeling UI for the Trip Validity model
VictorF13 Aug 15, 2026
350a435
feat: cap active-learning split at exactly 250/125/125 of the 500-lab…
VictorF13 Aug 15, 2026
74db004
feat: add weekday_number and is_weekend to the Trip Validity dataset
VictorF13 Aug 15, 2026
5db7a54
feat: session-only skip tracking; fix: artifact pickling breaks on ho…
VictorF13 Aug 15, 2026
1c1e4f5
feat: raise uncertain-draw probability once calibration/test are full
VictorF13 Aug 16, 2026
bf8ca88
feat: run full optimization on every retrain once eval sets are full
VictorF13 Aug 16, 2026
f57cf95
feat: retrain every 5 labels instead of 15 once eval sets are full
VictorF13 Aug 16, 2026
2894339
fix: add NOT NULL on company_id and refresh the stale table comment
VictorF13 Aug 16, 2026
3928434
feat: train the model on company_id and garage-distance features
VictorF13 Aug 16, 2026
287dc01
feat: add route straight-line distance features to the final dataset
VictorF13 Aug 16, 2026
935bb96
feat: train the model on route straight-line distance features
VictorF13 Aug 16, 2026
b443343
feat: add terminal-distance features to the final dataset
VictorF13 Aug 16, 2026
8c76adc
feat: train the model on the 6 terminal-distance features
VictorF13 Aug 16, 2026
c9133b8
feat: add notebook 06, a full model-family sweep over the labeled dat…
VictorF13 Aug 16, 2026
470a90d
fix: penalize CV variance in winner selection; rank features per family
VictorF13 Aug 16, 2026
8131292
feat: run final model sweep
VictorF13 Aug 16, 2026
5cada38
revert: remove notebook 06, the model sweep
VictorF13 Aug 17, 2026
16f7716
feat: track the active learning app's model artifacts in git
VictorF13 Aug 17, 2026
eef99cc
feat: add notebook 06, SFFS + nested repeated CV final model sweep
VictorF13 Aug 18, 2026
0c93fe7
feat: add ml.trip_validity_route_stops, ordered stops per route+direc…
VictorF13 Aug 22, 2026
7f72d5d
feat: add ml.trip_validity_final, the final per-trip validity table
VictorF13 Aug 22, 2026
6f5e342
feat: add ml.trip_validity_fares_final, the final fare-collection table
VictorF13 Aug 22, 2026
7735868
chore: retrigger CI after retargeting PR to develop
VictorF13 Aug 22, 2026
315df31
fix: install the ml dependency group in CI
VictorF13 Aug 22, 2026
79ae40d
Merge pull request #31 from VictorF13/feat/train-trip-validity-model
VictorF13 Aug 22, 2026
ae1e250
feat: build bus_matching_candidate_pairs and avl_positions view
VictorF13 Aug 22, 2026
ec0dbd8
fix: clamp out-of-range AVL headings to 0 in bus_matching_avl_positions
VictorF13 Aug 22, 2026
b49a829
chore: ignore regenerable bus matching feature checkpoints
VictorF13 Aug 23, 2026
a2c8601
feat: add contestedness, signature, and blocking-candidate notebooks
VictorF13 Aug 23, 2026
3953d05
feat: build bus-to-device matching active-learning labeler app
VictorF13 Aug 23, 2026
bfd631e
chore: commit bus matching model training checkpoints
VictorF13 Aug 23, 2026
a5e3ebf
feat: add per-date global assignment (plan Section 9)
VictorF13 Aug 23, 2026
8858a0c
feat: add per-device temporal smoothing (plan Section 10)
VictorF13 Aug 23, 2026
ffd36db
feat: exclude no-AVL companies, weight dictionary evidence, extend sa…
VictorF13 Aug 23, 2026
d42839d
fix: gate dictionary boost on being competitive with the actual evidence
VictorF13 Aug 23, 2026
6c75531
chore: commit bus matching model training checkpoints
VictorF13 Aug 23, 2026
9a0b40b
feat: add heading, AFC-direction, fare-timing and day-continuity feat…
VictorF13 Aug 23, 2026
87c3d3c
feat: add month-level pair features, labels, and pair model
VictorF13 Aug 23, 2026
e5c3cff
feat: add pair-validation UI and centralize bus exclusions
VictorF13 Aug 23, 2026
3084268
feat: add stopping signal to the pair-validation UI
VictorF13 Aug 23, 2026
004cfa6
feat: add audit sampling so precision is measured, not assumed
VictorF13 Aug 23, 2026
a77ba87
feat: stratify audit sampling, add k-fold CV, and persist the pair model
VictorF13 Aug 23, 2026
7b04303
feat: build Section 13's final output table, with device-swap split d…
VictorF13 Aug 24, 2026
f54ccf5
feat: add device-side coverage table, the other half of Section 13
VictorF13 Aug 24, 2026
df5074c
fix: correct 3 mutual-best matches the pair model wrongly zeroed out
VictorF13 Aug 24, 2026
4ca87ab
fix: rebuild final tables with the bus-21517 correction applied
VictorF13 Sep 11, 2026
b7fcc3a
Merge pull request #32 from VictorF13/feat/resolve-bus-device-pairs
VictorF13 Sep 11, 2026
1845b43
feat: add gold layer, a normalized trip-centric warehouse for Nov 2023
VictorF13 Sep 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 14 additions & 5 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,17 @@ RAW_DATA_ROOT=/srv/lab-data/dados-tp
# This is where the actual bronze layer parquets will be stored, below is the default
BRONZE_ROOT=/data/bronze

# Postgres+PostGIS silver layer, provisioned via `docker compose up -d`
SILVER_DB_USER=opa
SILVER_DB_PASSWORD=opa
SILVER_DB_NAME=opa
SILVER_DSN=postgresql://opa:opa@localhost:5432/opa
# Postgres+PostGIS database, provisioned via `docker compose up -d`.
# Shared by every schema (silver, ml, ...), not just silver.
DB_USER=opa
DB_PASSWORD=opa
DB_NAME=opa
DB_DSN=postgresql://opa:opa@localhost:5432/opa

# Which network interface Postgres/Adminer bind to. Defaults to
# 127.0.0.1 (localhost-only) if left unset - only set this if you need
# remote access, e.g. to your Tailscale IP (`tailscale ip -4`) to reach
# them over Tailscale without exposing them to the wider network. Must
# be a literal IP, not a hostname - Docker doesn't resolve DNS names for
# port binding.
# BIND_HOST=127.0.0.1
4 changes: 2 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ jobs:
with:
version: ${{ env.UV_VERSION }}
enable-cache: true
- run: uv sync
- run: uv sync --group ml
- run: uv run ruff format --check .
- run: uv run ruff check .
- run: uv run ty check
Expand All @@ -51,7 +51,7 @@ jobs:
with:
version: ${{ env.UV_VERSION }}
enable-cache: true
- run: uv sync
- run: uv sync --group ml
- name: Run tests (tolerate "no tests collected")
run: |
set +e
Expand Down
19 changes: 13 additions & 6 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -6,21 +6,28 @@ dist/
wheels/
*.egg-info
.ruff_cache
.pytest_cache
.coverage
htmlcov/

# Jupyter
.ipynb_checkpoints/

# Virtual environments
.venv

# Environment variables
.env

# Local bronze/gold layer output
# Local bronze/silver layer output
data/

# dbt-generated artifacts
gold/target/
gold/logs/
gold/dbt_packages/
gold/.user.yml
# Regenerable per-date feature checkpoints (ml/bus_matching_model)
ml/bus_matching_model/artifacts/features/
ml/bus_matching_model/artifacts/features_v2/

# Claude Code local state
.claude/

# macOS
.DS_Store
11 changes: 11 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,17 @@ repos:
language: python
types: [python]
pass_filenames: true
# ml/bus_matching_model/app and ml/trip_validity_model/app both
# have same-named sibling modules (features.py, db.py, ...). A
# single shared pyproject.toml `environment.extra-paths` list
# makes bare `from features import X` ambiguous between the two
# -- confirmed live, ty silently resolves it to whichever tree is
# listed first, breaking the other tree's own sibling imports.
# Excluded here rather than added to that shared list; matches
# the standing call to leave this app's lint/format debt for
# later too. `scripts/` is excluded for the same reason -- it
# imports those same sibling modules by bare name.
exclude: ^ml/bus_matching_model/(app|scripts)/
- repo: https://github.com/astral-sh/uv-pre-commit
rev: 0.11.22
hooks:
Expand Down
67 changes: 17 additions & 50 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,16 +19,26 @@ uv run opa-database ingest-reference vehicle_dictionary
uv run opa-database load-silver {avl,afc,gtfs} --year Y --month M
uv run opa-database load-silver-reference vehicle_dictionary

# Gold layer (dbt project lives in gold/, not under src/)
uv run dbt run --project-dir gold --profiles-dir gold
uv run dbt test --project-dir gold --profiles-dir gold

uv run prek run --all-files # dry-run every pre-commit hook before committing
```

There's no single-test invocation documented yet since the pytest suite is
empty; use standard `pytest path::test_name` once tests exist.

**Before considering any code change done** (this includes notebooks),
run `uv run prek run --all-files` and fix everything it reports — don't
stop at a clean `ruff check` in isolation; `prek` also runs `ruff
format`, `ty check`, and the `requirements*.txt`/`uv.lock` sync hooks.
Two gotchas learned the hard way:

- `prek run --all-files` only checks files **git already tracks**. A
freshly created, still-untracked file (e.g. a new notebook) is
silently skipped — `git add` it first, or run `ruff
check`/`ruff format --check`/`ty check` directly against the new
path(s), before trusting a clean `prek` result.
- The `ruff-check`/`ruff-format` hooks cover `.ipynb` files, not just
`.py` — don't assume notebooks are exempt from linting.

Always invoke Python through `uv run` (`uv run <script.py>`, `uv run
pytest`, ...) rather than calling `python`/`python3` directly, so it runs
against this project's synced environment and pinned Python version.
Expand All @@ -44,21 +54,13 @@ needed). See `src/opa_database/config.py` or
convention to match.

**Python version is pinned to 3.12** (`pyproject.toml`, `.python-version`).
This is a hard requirement, not a preference. `dbt-core`'s dependency
`mashumaro` fails to import under Python 3.14 (`UnserializableField` on a
plain `Optional[str]`, verified across the entire mashumaro version range
dbt allows) and hits a different incompatibility under 3.13. If dbt starts
throwing `mashumaro`/`typing_extensions` errors, check whether 3.13/3.14
crept back in before assuming it's a config problem. Don't try pinning
mashumaro/typing_extensions instead, that was already tried and failed.

## Architecture

This is a three-layer ("medallion") pipeline turning Fortaleza, Brazil's
This is a two-layer ("medallion") pipeline turning Fortaleza, Brazil's
raw public transit data into a queryable PostgreSQL+PostGIS database:
**bronze** (raw files -> typed Parquet) -> **silver** (Parquet -> per-source
normalized Postgres tables) -> **gold** (dbt models joining across sources).
Full rationale in `docs/architecture.md` and `docs/gold-layer.md`; the
normalized Postgres tables). Full rationale in `docs/architecture.md`; the
essential cross-file structure is:

### Bronze (`src/opa_database/{adapters,contracts}/`, `loaders/bronze.py`)
Expand Down Expand Up @@ -93,8 +95,7 @@ its bronze partition key.
Silver is **strictly per-source and stays flat/denormalized** (e.g.
`silver.afc_boardings` repeats every trip/line/vehicle/company column on
every boarding row, ~16.4x redundancy, measured). This is deliberate, not
unfinished. Fact/dimension splitting and any cross-source join belongs in
gold, not silver.
unfinished.

Timezones: AFC's raw timestamps are naive Fortaleza local time (UTC-3, no
DST since 2008) and are converted to UTC during the silver load; AVL/GPS is
Expand All @@ -105,40 +106,6 @@ PostGIS `geometry(Point, 4326)` columns are Postgres `GENERATED ALWAYS AS
... STORED` columns computed from lat/lon by Postgres itself, not written
directly; bulk `COPY` only ever carries the plain lat/lon columns.

### Gold (`gold/`, a self-contained dbt-core project)

Not under `src/opa_database/`, a separate SQL toolchain. Builds on
`silver` via `{{ source(...) }}`, materializes into schema `gold`. Full
model inventory and testing conventions in `docs/gold-layer.md`; key
patterns that apply to any new gold model:

- **GTFS/AFC ids repeat across snapshots**, so every model's real key is a
composite `(feed_version_date, entity_id)` (or the AFC equivalent), never
a bare id. dbt's built-in single-column `relationships`/`unique` tests
would silently pass broken references in this shape, so referential
integrity and uniqueness are enforced with hand-written singular tests in
`gold/tests/*_exists.sql` / `*_unique_per_*.sql` that join/group on the
full composite key.
- **Surrogate keys** for entities with no natural id (AFC has no trip id)
use `md5(concat_ws('|', ...))` over the full natural-key column set,
factored into a shared macro (`gold/macros/afc_trip_key.sql`) rather than
repeated inline.
- **Conformed dimension pattern**: `dim_vehicle_master` reconciles AFC
vehicle identity (`vehicle_number`) and GPS vehicle identity
(`vehicle_id`), which are independently assigned by each source. It's a
deterministic function of the latest `vehicle_dictionary` snapshot (not
a persisted/stateful registry). Every vehicle from every source gets
exactly one `master_vehicle_id`, synthesizing a source-prefixed id where
no confirmed match exists, tagged via `match_status` rather than dropped.
- **Don't duplicate an id reachable through an existing relationship**:
fact tables carry `master_vehicle_id` only at the grain it canonically
belongs to (e.g. `dim_afc_trip`, not `fact_afc_boarding`), and never keep
a raw per-source id alongside the conformed one once it's recoverable via
a dimension's crosswalk column.
- Referential integrity in gold is dbt tests, not enforced Postgres
`FOREIGN KEY`/`PRIMARY KEY` constraints (dbt model contracts aren't
turned on yet).

## CI / release

`.github/workflows/ci.yml`: PR titles are enforced as Conventional Commits;
Expand Down
7 changes: 3 additions & 4 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,8 @@ uv run prek install # or: pre-commit install
```

See the [README](README.md) for environment variables and running the
pipeline end to end, and [`docs/architecture.md`](docs/architecture.md) /
[`docs/gold-layer.md`](docs/gold-layer.md) for how the codebase is
organized.
pipeline end to end, and [`docs/architecture.md`](docs/architecture.md)
for how the codebase is organized.

## Dependencies

Expand Down Expand Up @@ -100,7 +99,7 @@ The end-to-end flow:
[Conventional Commits](https://www.conventionalcommits.org/)
(`feat:`, `fix:`, `chore:`, `docs:`, `refactor:`, `perf:`, `test:`,
`ci:`, `build:`, `style:`, `revert:`, optionally scoped like
`feat(gold): ...`), `validate` runs formatting/lint/type checks, and
`feat(silver): ...`), `validate` runs formatting/lint/type checks, and
`test` runs the pytest suite. All three must pass before merging.
2. Once merged into `develop`, that push triggers `prep-release-pr`:
CI automatically opens (or updates, if one is already open) a
Expand Down
44 changes: 14 additions & 30 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,14 +6,14 @@ queryable, analysis-ready PostgreSQL+PostGIS database.

## Overview

The pipeline follows a three-layer ("medallion") architecture:
The pipeline follows a two-layer ("medallion") architecture:

```bash
raw files (CSV/XML/zip) bronze silver gold
------------------------ --> -------------------- --> -------------------- --> --------------------
On disk, agency-supplied Typed, validated, Per-source, normalized Cross-source,
formats, one file per day/ Hive-partitioned Parquet PostgreSQL+PostGIS dimensional models
export/snapshot (data/bronze/...) tables (schema `silver`) (dbt, schema `gold`)
raw files (CSV/XML/zip) bronze silver
------------------------ --> -------------------- --> --------------------
On disk, agency-supplied Typed, validated, Per-source, normalized
formats, one file per day/ Hive-partitioned Parquet PostgreSQL+PostGIS
export/snapshot (data/bronze/...) tables (schema `silver`)
```

- **Bronze**: raw agency exports (CSV, nested XML, zipped GTFS feeds) are
Expand All @@ -25,16 +25,9 @@ export/snapshot (data/bronze/...) tables (schema `silver
PostgreSQL+PostGIS (schema `silver`). Still strictly per-source (no
AFC-to-GPS vehicle reconciliation, no GTFS-to-ridership joins here) but
now typed, deduplicated, indexed, and queryable with SQL/PostGIS.
- **Gold**: a [dbt](https://docs.getdbt.com/) project (schema `gold`) that
builds normalized fact/dimension models on top of silver: a proper
star-schema split of the flat AFC feed, a conformed vehicle dimension
that reconciles AFC and GPS vehicle identities, and dimensional models
for every GTFS table. This is where cross-source joins and analytical
modeling happen.

See [`docs/architecture.md`](docs/architecture.md) for the full design
rationale, and [`docs/gold-layer.md`](docs/gold-layer.md) for the dbt
project specifically.
rationale.

## Data sources

Expand All @@ -53,7 +46,7 @@ conventions, known data-quality issues) are in

```text
src/opa_database/
config.py Settings (env-driven: raw_data_root, bronze_root, silver_dsn)
config.py Settings (env-driven: raw_data_root, bronze_root, db_dsn)
cli.py Click CLI: ingest, ingest-reference, load-silver, load-silver-reference
contracts/ Pandera schemas for each raw source (bronze validation)
adapters/ Raw file -> validated bronze Parquet, one module per source
Expand All @@ -62,23 +55,17 @@ src/opa_database/
silver.py Generic idempotent "replace a period" loader for silver tables
silver/ Bronze Parquet -> silver PostgreSQL+PostGIS, one module per source

gold/ dbt project: fact/dimension models on top of silver (schema `gold`)
models/{afc,gtfs,vehicle}/
macros/
tests/ Hand-written composite-key uniqueness/referential-integrity tests

docs/ Architecture, gold-layer reference, remote access
docs/ Architecture reference, remote access
docker-compose.yml Postgres+PostGIS and Adminer (web SQL UI)
```

## Getting started

### Prerequisites

- Python 3.12 (see [`docs/gold-layer.md`](docs/gold-layer.md) for why not
3.13/3.14)
- Python 3.12
- [uv](https://docs.astral.sh/uv/)
- Docker + Docker Compose (for the Postgres+PostGIS silver/gold database)
- Docker + Docker Compose (for the Postgres+PostGIS silver database)

### Setup

Expand All @@ -94,8 +81,9 @@ docker compose up -d # starts Postgres+PostGIS on :5432 and Adminer on :8080
| --- | --- |
| `RAW_DATA_ROOT` | Path to the raw agency data on disk |
| `BRONZE_ROOT` | Where bronze Parquet files are written |
| `SILVER_DB_USER` / `SILVER_DB_PASSWORD` / `SILVER_DB_NAME` | Postgres credentials, used both by `docker compose` and by the app |
| `SILVER_DSN` | Full connection string the pipeline uses to reach Postgres |
| `DB_USER` / `DB_PASSWORD` / `DB_NAME` | Postgres credentials, used both by `docker compose` and by the app |
| `DB_DSN` | Full connection string the pipeline uses to reach Postgres. Shared by every schema (`silver`, `ml`, ...), not silver-specific |
| `BIND_HOST` | Optional. Network interface Postgres/Adminer bind to, defaults to `127.0.0.1` (localhost-only). See [`docs/remote-access.md`](docs/remote-access.md) to expose them over Tailscale instead |

### Running the pipeline

Expand All @@ -111,10 +99,6 @@ uv run opa-database load-silver avl --year 2023 --month 11
uv run opa-database load-silver afc --year 2023 --month 11
uv run opa-database load-silver gtfs --year 2023 --month 11
uv run opa-database load-silver-reference vehicle_dictionary

# Gold: build the dbt models on top of silver
uv run dbt run --project-dir gold --profiles-dir gold
uv run dbt test --project-dir gold --profiles-dir gold
```

Each `load-silver`/`load-silver-reference` run is idempotent: re-running it
Expand Down
10 changes: 5 additions & 5 deletions docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2,18 +2,18 @@ services:
postgres:
image: postgis/postgis:16-3.4
environment:
POSTGRES_USER: ${SILVER_DB_USER:-opa}
POSTGRES_PASSWORD: ${SILVER_DB_PASSWORD:-opa}
POSTGRES_DB: ${SILVER_DB_NAME:-opa}
POSTGRES_USER: ${DB_USER:-opa}
POSTGRES_PASSWORD: ${DB_PASSWORD:-opa}
POSTGRES_DB: ${DB_NAME:-opa}
ports:
- "5432:5432"
- "${BIND_HOST:-127.0.0.1}:5432:5432"
volumes:
- opa_postgres_data:/var/lib/postgresql/data

adminer:
image: adminer
ports:
- "8080:8080"
- "${BIND_HOST:-127.0.0.1}:8080:8080"
depends_on:
- postgres

Expand Down
Loading
Loading