Lightweight multimodal classifier for scientific poster quality control in open repositories.
Part of the quality control pipeline for posters.science, a platform for making scientific conference posters Findable, Accessible, Interoperable, and Reusable (FAIR).
Developed by the FAIR Data Innovations Hub at the California Medical Innovations Institute (CalMI2).
Open repositories like Zenodo and Figshare host tens of thousands of records labeled as scientific posters. However, approximately 20% of these records are mislabeled — containing multi-page papers, conference proceedings, abstract booklets, slide decks, or other non-poster documents. This label noise is a significant barrier to automated poster processing at scale.
PosterSentry classifies PDFs using three complementary feature channels summarized into a compact 31-feature vector (a stage-one text score plus 15 visual and 15 structural features):
| Channel | Features | Dimensions | Signal |
|---|---|---|---|
| Text | model2vec (potion-base-32M) embedding | 512 | Semantic content |
| Visual | Color stats, edge density, FFT spatial complexity, whitespace | 15 | Visual layout |
| Structural | Page count, area, font diversity, text blocks, density | 15 | PDF geometry |
PosterSentry uses a two-stage (stacked) design: stage one scores the 512-dimensional text embedding alone, and its poster probability becomes a single text_score feature for a stage-two logistic regression over 31 features (text score plus the 15 visual and 15 structural features). Each stage has its own StandardScaler fit on the training split only, and stage two is trained on inner five-fold out-of-fold text scores so it never sees an in-sample-optimistic text score.
Both stages live in one numpy .npz head (20 KB). Inference is pure numpy — no GPU or deep learning framework required.
Validated on the human-validated, license-cleared corpus (3,298 documents, zero synthetic data):
| Metric | Value |
|---|---|
| Held-out accuracy | 93.1% (95% CI 90.6 to 95.0) |
| Nested out-of-fold accuracy | 93.9% |
| F1 (poster) | 93.3% |
| F1 (non-poster) | 93.0% |
| Precision / Recall (poster) | 91.8% / 94.8% |
| Inference speed | < 1 sec/PDF (CPU) |
Applied to 30,139 readable PDFs from Zenodo and Figshare, PosterSentry classified 80.5% as posters and 19.5% as non-posters: roughly one in five records labeled as posters is something else.
| Feature | Coefficient | Signal |
|---|---|---|
page_count |
-3.26 | More pages pushes away from poster |
size_per_page_kb |
+2.42 | Dense, high-resolution single pages |
line_count |
+1.94 | Posters pack many short text lines |
file_size_kb |
-1.70 | Multi-page documents are bigger overall |
mean_g |
+1.10 | Colorful, non-white pages |
is_landscape |
+1.00 | Many posters are landscape |
pip install poster-sentry# Classify a single PDF
poster-sentry classify document.pdf
# Classify multiple PDFs
poster-sentry classify *.pdf --output results.tsv
# Print model info
poster-sentry infofrom poster_sentry import PosterSentry
sentry = PosterSentry()
sentry.initialize()
# Classify a PDF (uses text + visual + structural features)
result = sentry.classify("document.pdf")
print(f"Is poster: {result['is_poster']}, Confidence: {result['confidence']:.2f}")
# Batch classification
results = sentry.classify_batch(["poster1.pdf", "paper.pdf", "newsletter.pdf"])
# Text-only classification (no PDF needed)
result = sentry.classify_text("Title: My Poster\nAuthors: ...")PosterSentry sits at the front of the posters.science pipeline — it screens incoming PDFs before expensive LLM-based extraction:
PDF Input
|
v
PosterSentry --> poster2json --> FAIR output
(classify: poster?) (Llama 3.1 8B structured extraction) (poster-json-schema)
| Requirement | Value |
|---|---|
| CPU | Any modern CPU (no GPU needed) |
| RAM | 4 GB+ |
| Python | 3.10+ |
| Model size | 20 KB head + ~60 MB embeddings (downloaded once) |
| Resource | Description |
|---|---|
| poster-sentry (HuggingFace) | Model weights and config |
| poster-sentry-training-data (HuggingFace) | Training dataset (3,298 samples) |
| poster-sentry-training (GitHub) | Training code and replication |
| poster2json | Poster to structured JSON extraction |
| posters.science | Platform |
git clone https://github.com/fairdataihub/poster-sentry.git
cd poster-sentry
pip install -e ".[dev]"
pytest@software{poster_sentry_2026,
title = {PosterSentry: Multimodal Scientific Poster Classifier},
author = {O'Neill, James and Soundarajan, Sanjay and Portillo, Dorian and Patel, Bhavesh},
year = {2026},
url = {https://github.com/fairdataihub/poster-sentry},
note = {Part of the posters.science initiative at FAIR Data Innovations Hub}
}MIT License. See LICENSE for details.
- FAIR Data Innovations Hub at California Medical Innovations Institute (CalMI2)
- posters.science platform
- MinishLab for the model2vec embedding backbone
- Funded by The Navigation Fund — "Poster Sharing and Discovery Made Easy"
