Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Semantic Document Intelligence

Evidence-Grounded Question Answering over Documents
A Django-based semantic document intelligence system that retrieves document evidence, extracts concept-specific answer spans, validates grounding, and returns concise answers with paragraph-level citations.

Python Django Qdrant OpenAI Docker Git Markdown Status

Overview • Architecture • Features • Installation • Demo • Evaluation • Status


Overview

Semantic Document Intelligence is an evidence-grounded question-answering system for extracting reliable answers from ingested documents.

The system is designed around a simple principle:

An answer should be supported by evidence from the selected document.

Instead of returning a large retrieval chunk, the final question-answering pipeline progressively narrows evidence through retrieval, concept-specific extraction, evidence windowing, grounding validation, answer compression, and final answer-contract enforcement.

The project reached its final implementation milestone at Phase 4.8.

What the system does

User Question
      │
      ▼
Question Analysis
      │
      ▼
Semantic / Graph Retrieval
      │
      ▼
Candidate Evidence
      │
      ▼
Concept-Specific Span Extraction
      │
      ▼
Concept-Scoped Evidence Window
      │
      ▼
Grounding + Confidence Validation
      │
      ├── LLM Answer Path
      │
      └── Deterministic Extractive Fallback
      │
      ▼
Answer Compression
      │
      ▼
Final Answer Contract
      │
      ▼
Concise Answer + Citations

Key Features

  • 📄 Document-scoped question answering
  • 🔎 Semantic and graph-aware evidence retrieval
  • 🎯 Concept-specific answer span extraction
  • 🪟 Concept-scoped evidence windowing
  • 🧠 LLM-assisted answer generation
  • 🛡️ Deterministic extractive fallback
  • 🔗 Paragraph-level citation tracking
  • 📊 Evidence quality and grounding metadata
  • 🎚️ Confidence-aware answer selection
  • ✂️ Answer compression to remove irrelevant context
  • 🚫 No-evidence behavior instead of unsupported answers
  • 🧪 Regression-oriented demonstration questions
  • 🐳 Docker-compatible development environment

Final Question-Answering Pipeline

The final pipeline contains the following logical stages.

1. Retrieval

Relevant document paragraphs are retrieved using the project's semantic/graph retrieval layer.

2. Concept-Specific Answer Span Extraction

The system identifies the portion of retrieved evidence that is most directly associated with the requested concept.

For example:

Question:
What is a power set?

Retrieved paragraph:
Power Set: Power set = set of all subsets.
Power set: P(S) = {...}

Selected evidence:
Power Set: Power set = set of all subsets.

This prevents unrelated material from dominating the final response.

3. Concept-Scoped Evidence Windowing

Evidence is narrowed around the concept-bearing portion instead of passing the entire paragraph blindly into later stages.

4. Grounding and Confidence Validation

Candidate answers are checked against the available evidence.

The system tracks information such as:

evidence_quality
grounding_level
citation_id
paragraph_id
page

5. LLM Answer Path

When the available evidence satisfies the configured conditions, the system can use the LLM answer path to produce a grounded response.

6. Deterministic Extractive Fallback

If the LLM path cannot safely produce an answer, the system falls back to deterministic evidence extraction.

This is particularly important for document QA because unsupported generation is preferable to a confident hallucination.

7. Answer Compression

Long evidence spans are compressed so that the final response focuses on the requested concept.

8. Final Answer Contract

The final stage ensures the public service response maintains the expected contract:

{
    "method": ...,
    "answer": ...,
    "citations": [...]
}

Architecture

semantic-document-intelligence/
│
├── apps/
│   └── question_answering/
│       ├── answer_service.py
│       └── ...
│
├── manage.py
│
├── README.md
├── semantic_document_intelligence_advanced_project_term_paper_Subhankar_Biswas_30008967.pdf
├── DEMONSTRATION.md
│
├── requirements.txt
├── docker-compose.yml
│
└── ...

🏗️ System Architecture

Semantic Document Intelligence Architecture

The architecture combines application services, asynchronous processing, object storage, relational persistence, graph-based knowledge representation, and semantic RDF/SPARQL capabilities.

Neo4j Graphical Example

Neo4j Sample Graph

A Neo4j Graphical Representation of a processed document.

Core application flow

                    ┌──────────────────────┐
                    │      User Question   │
                    └──────────┬───────────┘
                               │
                               ▼
                    ┌──────────────────────┐
                    │  AnswerService       │
                    └──────────┬───────────┘
                               │
             ┌─────────────────┴─────────────────┐
             ▼                                   ▼
   ┌────────────────────┐             ┌────────────────────┐
   │ Retrieval Service  │             │ Question Analysis  │
   └─────────┬──────────┘             └─────────┬──────────┘
             │                                  │
             └────────────────┬─────────────────┘
                              ▼
                   ┌─────────────────────┐
                   │ Evidence Selection  │
                   └──────────┬──────────┘
                              ▼
                   ┌─────────────────────┐
                   │ Span / Windowing    │
                   └──────────┬──────────┘
                              ▼
                   ┌─────────────────────┐
                   │ Grounding Validator │
                   └──────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ LLM / Extractive   │
                    │ Answer Generation  │
                    └─────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ Compression +      │
                    │ Final Contract     │
                    └─────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ Answer + Citations │
                    └────────────────────┘

Tech Stack

Core Technologies

Technology Role Official Documentation
Python Application and NLP/data-processing language python.org
Django Web/application framework djangoproject.com
Qdrant Vector similarity search / retrieval infrastructure qdrant.tech
OpenAI API LLM-assisted answer generation platform.openai.com
Docker Containerized development and services docker.com
Git Version control git-scm.com
Markdown Project documentation markdownguide.org

GitHub Tech Stack Links

The following badges are intentionally clickable so visitors can directly explore the technologies used by the project:

Python Django Qdrant OpenAI Docker Git Markdown

The badges use standard Shields.io badge URLs, which are suitable for GitHub README files. citeturn0search0turn0search2


Installation

Prerequisites

Recommended environment:

  • Python 3.12+
  • Git
  • Docker Desktop
  • Docker Compose
  • Access to the required LLM/API credentials
  • A configured Qdrant instance or the project's Docker-based Qdrant service

Clone the repository

git clone https://github.com/Subiswas36218/semantic-document-intelligence.git
cd semantic-document-intelligence

Create a virtual environment

python3 -m venv venv
source venv/bin/activate

On Windows:

python -m venv venv
venv\Scripts\activate

Install dependencies

pip install -r requirements.txt

Configure environment variables

Create the project's environment file according to the variables expected by the application.

Example:

OPENAI_API_KEY=your_api_key_here

Do not commit secrets to GitHub.

Start supporting services

If the project uses Docker Compose:

docker compose up -d

Check running services:

docker compose ps

Run Django checks

python manage.py check

Ontology-grounded document intelligence with Django, Apache Jena/Fuseki, Neo4j/Cypher, Celery, PostgreSQL and MinIO.

Pipeline:

PDF/DOCX -> structured JSON -> LLM semantic extraction -> validation -> ontology mapping -> RDF/Jena + Neo4j projections -> graph retrieval -> source grounding -> answer.

The original document is preserved, JSON remains the exact source-recovery layer, and graph stores provide semantic retrieval.

Quick start

cp .env.example .env
docker compose up --build
docker compose exec web python manage.py migrate

Services:

Run checks:

make check

Demonstration

The final regression demonstration uses the following five questions:

questions = [
    "What are De Morgan’s laws?",
    "What is a power set?",
    "Explain measurable function composition.",
    "What is a Cartesian product?",
    "What are sigma fields?",
]

Run:

python manage.py shell -c "
from apps.question_answering.answer_service import create_answer_service

service = create_answer_service()

questions = [
    'What are De Morgan’s laws?',
    'What is a power set?',
    'Explain measurable function composition.',
    'What is a Cartesian product?',
    'What are sigma fields?',
]

try:
    for question in questions:
        result = service.answer(
            question=question,
            document_id='doc_fdd043b9648d',
        )

        print('=' * 100)
        print('QUESTION:', question)
        print('METHOD:', result['method'])
        print('ANSWER:', result['answer'])
        print('CITATIONS:')

        for citation in result['citations']:
            print(
                ' ->',
                citation['citation_id'],
                'paragraph=', citation['paragraph_id'],
                'page=', citation['page'],
                'quality=', citation['evidence_quality'],
                'grounding=', citation['grounding_level'],
            )
finally:
    service.retrieval_service.close()
"

Demonstration Results

The final regression run successfully produced answer-ready evidence for all five demonstration questions.

Question Method Result
What are De Morgan’s laws? extractive_fallback Answer extracted from page 1
What is a power set? extractive_fallback Answer extracted from page 3
Explain measurable function composition. extractive_fallback Answer extracted from page 6
What is a Cartesian product? extractive_fallback Answer extracted from page 3
What are sigma fields? extractive_fallback Evidence extracted from pages 4–5

Representative outputs include:

QUESTION: What is a power set?

METHOD: extractive_fallback

ANSWER:
Power Set: Power set = set of all subsets.

CITATIONS:
 -> 1 paragraph=doc_fdd043b9648d_p003_001
    page=3
    quality=medium
    grounding=dual_graph

and:

QUESTION: Explain measurable function composition.

METHOD: extractive_fallback

ANSWER:
Measurable Function Composition Given measurable spaces,
(S₁,F₁), (S₂,F₂), (S₃,F₃) ...

CITATIONS:
 -> 1 paragraph=doc_fdd043b9648d_p006_001
    page=6
    quality=high
    grounding=dual_graph

Citation Model

Each returned citation contains structured evidence metadata.

Example:

{
    "citation_id": 1,
    "paragraph_id": "doc_fdd043b9648d_p003_001",
    "page": 3,
    "evidence_quality": "medium",
    "grounding_level": "dual_graph",
}

This provides traceability from:

Question
   ↓
Answer
   ↓
Citation
   ↓
Paragraph
   ↓
Document Page

Evidence-Grounded Behavior

The system is deliberately designed not to answer every question from general model knowledge.

If the document does not contain sufficient answer-ready evidence, the service can return:

METHOD: no_evidence

ANSWER:
I could not find enough answer-ready evidence in the document
to answer this question reliably.

This behavior is an important part of the system's reliability design.


Evaluation

The final implementation was evaluated using a five-question regression suite covering:

  1. Set-theoretic laws
  2. Power sets
  3. Measurable functions
  4. Cartesian products
  5. Sigma fields

Evaluation dimensions

Dimension Objective
Retrieval relevance Retrieve paragraphs related to the question
Concept precision Isolate the requested concept
Evidence grounding Keep the answer supported by retrieved evidence
Citation traceability Connect answers to document paragraphs/pages
Conciseness Avoid returning unnecessary surrounding context
Fallback safety Prefer evidence extraction over unsupported generation
Contract stability Preserve the expected answer/citation response structure

Final assessment

The final regression output demonstrates that the system can:

  • locate concept-relevant evidence;
  • distinguish between answer-ready and insufficient evidence;
  • extract compact answer spans;
  • preserve document-level grounding;
  • return paragraph/page citations;
  • use deterministic extraction as a reliable fallback;
  • reduce retrieval noise before final answer construction.

Project Development Phases

The final Question Answering implementation evolved through the following milestones:

Phase 4.4.3
    ↓
Concept-Specific Answer Span Extraction

Phase 4.4.4
    ↓
Improved answer-span selection and extraction

Phase 4.4.5
    ↓
Concept-Scoped Evidence Windowing

Phase 4.5
    ↓
Grounding / evidence-aware answer generation

Phase 4.6
    ↓
Confidence-aware answer selection and fallback behavior

Phase 4.7
    ↓
Answer compression and final answer-quality refinement

Phase 4.8
    ↓
Final answer contract / production-ready response path

Project Status

Status: ✅ COMPLETE

The project is considered complete at the planned Phase 4.8 endpoint.

No additional Question Answering development phase is required for the current project scope.

Remaining activities are limited to:

  • repository cleanup;
  • committing the final implementation;
  • adding project documentation;
  • recording/presenting the demonstration;
  • optional GitHub Actions/CI configuration;
  • final submission.

Documentation

Document Purpose
README.md Project overview and technical documentation
FINAL_REPORT.md Final evaluation and reporting
DEMONSTRATION.md Step-by-step demonstration guide

Recommended GitHub Repository Checklist

Before publishing:

[ ] README.md added
[ ] FINAL_REPORT.md added
[ ] DEMONSTRATION.md added
[ ] requirements.txt verified
[ ] docker-compose.yml verified
[ ] .env excluded from Git
[ ] API keys/secrets removed
[ ] temporary/debug files removed
[ ] final tests executed
[ ] Git status reviewed
[ ] GitHub repository pushed

Check for accidentally tracked secrets:

git status
git diff --cached

Example Git Workflow

git add README.md semantic_document_intelligence_advanced_project_term_paper_Subhankar_Biswas_30008967.pdf DEMONSTRATION.md
git add apps/
git add requirements.txt docker-compose.yml

git status

git commit -m "Finalize semantic document intelligence QA pipeline"

git push origin main

Future Extensions

The current project is complete, but potential future work could include:

  • multi-document question answering;
  • conversational follow-up questions;
  • answer highlighting in the source document;
  • richer citation previews;
  • automated evaluation datasets;
  • retrieval-quality benchmarking;
  • confidence calibration dashboards;
  • streaming answers;
  • asynchronous document ingestion;
  • web-based document exploration UI;
  • automated CI regression testing.

These are future extensions, not required for the current project completion.


License

Add the project's chosen license here, for example:

MIT License

If a license has not yet been selected, do not claim a license until one is actually added to the repository.


Author

Subhankar Biswas // M.Sc. Data Engineering // Constructor University, Bremen //

Semantic Document Intelligence — Evidence-Grounded Question Answering System


Acknowledgements

This project combines document retrieval, semantic search, evidence extraction, grounding validation, deterministic fallback strategies, and LLM-assisted question answering into a single document intelligence workflow.

Semantic Document Intelligence
Evidence first. Answers second.

About

A production-oriented semantic document intelligence platform combining Django, Apache Jena Fuseki, Neo4j, PostgreSQL, Celery, and semantic question answering with grounded evidence and citations. (MDE-DIS-02_Advance Project-1 / Constructor University, Bremen / Subhankar Biswas

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages