Skip to content

Repository files navigation

PosterSentry

Lightweight multimodal classifier for scientific poster quality control in open repositories.

License: MIT Python 3.10+ HuggingFace

PosterSentry

Part of the quality control pipeline for posters.science, a platform for making scientific conference posters Findable, Accessible, Interoperable, and Reusable (FAIR).

Developed by the FAIR Data Innovations Hub at the California Medical Innovations Institute (CalMI2).

The Problem

Open repositories like Zenodo and Figshare host tens of thousands of records labeled as scientific posters. However, approximately 20% of these records are mislabeled — containing multi-page papers, conference proceedings, abstract booklets, slide decks, or other non-poster documents. This label noise is a significant barrier to automated poster processing at scale.

Architecture

PosterSentry classifies PDFs using three complementary feature channels summarized into a compact 31-feature vector (a stage-one text score plus 15 visual and 15 structural features):

Channel Features Dimensions Signal
Text model2vec (potion-base-32M) embedding 512 Semantic content
Visual Color stats, edge density, FFT spatial complexity, whitespace 15 Visual layout
Structural Page count, area, font diversity, text blocks, density 15 PDF geometry

PosterSentry uses a two-stage (stacked) design: stage one scores the 512-dimensional text embedding alone, and its poster probability becomes a single text_score feature for a stage-two logistic regression over 31 features (text score plus the 15 visual and 15 structural features). Each stage has its own StandardScaler fit on the training split only, and stage two is trained on inner five-fold out-of-fold text scores so it never sees an in-sample-optimistic text score.

Both stages live in one numpy .npz head (20 KB). Inference is pure numpy — no GPU or deep learning framework required.

Performance

Validated on the human-validated, license-cleared corpus (3,298 documents, zero synthetic data):

Metric Value
Held-out accuracy 93.1% (95% CI 90.6 to 95.0)
Nested out-of-fold accuracy 93.9%
F1 (poster) 93.3%
F1 (non-poster) 93.0%
Precision / Recall (poster) 91.8% / 94.8%
Inference speed < 1 sec/PDF (CPU)

Applied to 30,139 readable PDFs from Zenodo and Figshare, PosterSentry classified 80.5% as posters and 19.5% as non-posters: roughly one in five records labeled as posters is something else.

Top Discriminative Features

Feature Coefficient Signal
page_count -3.26 More pages pushes away from poster
size_per_page_kb +2.42 Dense, high-resolution single pages
line_count +1.94 Posters pack many short text lines
file_size_kb -1.70 Multi-page documents are bigger overall
mean_g +1.10 Colorful, non-white pages
is_landscape +1.00 Many posters are landscape

Quick Start

Installation

pip install poster-sentry

CLI Usage

# Classify a single PDF
poster-sentry classify document.pdf

# Classify multiple PDFs
poster-sentry classify *.pdf --output results.tsv

# Print model info
poster-sentry info

Python API

from poster_sentry import PosterSentry

sentry = PosterSentry()
sentry.initialize()

# Classify a PDF (uses text + visual + structural features)
result = sentry.classify("document.pdf")
print(f"Is poster: {result['is_poster']}, Confidence: {result['confidence']:.2f}")

# Batch classification
results = sentry.classify_batch(["poster1.pdf", "paper.pdf", "newsletter.pdf"])

# Text-only classification (no PDF needed)
result = sentry.classify_text("Title: My Poster\nAuthors: ...")

Pipeline Position

PosterSentry sits at the front of the posters.science pipeline — it screens incoming PDFs before expensive LLM-based extraction:

PDF Input
   |
   v
PosterSentry          -->  poster2json                     -->  FAIR output
(classify: poster?)        (Llama 3.1 8B structured extraction)  (poster-json-schema)

System Requirements

Requirement Value
CPU Any modern CPU (no GPU needed)
RAM 4 GB+
Python 3.10+
Model size 20 KB head + ~60 MB embeddings (downloaded once)

Related Resources

Resource Description
poster-sentry (HuggingFace) Model weights and config
poster-sentry-training-data (HuggingFace) Training dataset (3,298 samples)
poster-sentry-training (GitHub) Training code and replication
poster2json Poster to structured JSON extraction
posters.science Platform

Development

git clone https://github.com/fairdataihub/poster-sentry.git
cd poster-sentry
pip install -e ".[dev]"
pytest

Citation

@software{poster_sentry_2026,
  title = {PosterSentry: Multimodal Scientific Poster Classifier},
  author = {O'Neill, James and Soundarajan, Sanjay and Portillo, Dorian and Patel, Bhavesh},
  year = {2026},
  url = {https://github.com/fairdataihub/poster-sentry},
  note = {Part of the posters.science initiative at FAIR Data Innovations Hub}
}

License

MIT License. See LICENSE for details.

Acknowledgments

About

Lightweight multimodal scientific poster classifier — text + visual + structural features. Part of posters.science.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages