Evidence-Grounded Question Answering over Documents
A Django-based semantic document intelligence system that retrieves document evidence, extracts concept-specific answer spans, validates grounding, and returns concise answers with paragraph-level citations.
Overview • Architecture • Features • Installation • Demo • Evaluation • Status
Semantic Document Intelligence is an evidence-grounded question-answering system for extracting reliable answers from ingested documents.
The system is designed around a simple principle:
An answer should be supported by evidence from the selected document.
Instead of returning a large retrieval chunk, the final question-answering pipeline progressively narrows evidence through retrieval, concept-specific extraction, evidence windowing, grounding validation, answer compression, and final answer-contract enforcement.
The project reached its final implementation milestone at Phase 4.8.
User Question
│
▼
Question Analysis
│
▼
Semantic / Graph Retrieval
│
▼
Candidate Evidence
│
▼
Concept-Specific Span Extraction
│
▼
Concept-Scoped Evidence Window
│
▼
Grounding + Confidence Validation
│
├── LLM Answer Path
│
└── Deterministic Extractive Fallback
│
▼
Answer Compression
│
▼
Final Answer Contract
│
▼
Concise Answer + Citations
- 📄 Document-scoped question answering
- 🔎 Semantic and graph-aware evidence retrieval
- 🎯 Concept-specific answer span extraction
- 🪟 Concept-scoped evidence windowing
- 🧠 LLM-assisted answer generation
- 🛡️ Deterministic extractive fallback
- 🔗 Paragraph-level citation tracking
- 📊 Evidence quality and grounding metadata
- 🎚️ Confidence-aware answer selection
- ✂️ Answer compression to remove irrelevant context
- 🚫 No-evidence behavior instead of unsupported answers
- 🧪 Regression-oriented demonstration questions
- 🐳 Docker-compatible development environment
The final pipeline contains the following logical stages.
Relevant document paragraphs are retrieved using the project's semantic/graph retrieval layer.
The system identifies the portion of retrieved evidence that is most directly associated with the requested concept.
For example:
Question:
What is a power set?
Retrieved paragraph:
Power Set: Power set = set of all subsets.
Power set: P(S) = {...}
Selected evidence:
Power Set: Power set = set of all subsets.
This prevents unrelated material from dominating the final response.
Evidence is narrowed around the concept-bearing portion instead of passing the entire paragraph blindly into later stages.
Candidate answers are checked against the available evidence.
The system tracks information such as:
evidence_quality
grounding_level
citation_id
paragraph_id
page
When the available evidence satisfies the configured conditions, the system can use the LLM answer path to produce a grounded response.
If the LLM path cannot safely produce an answer, the system falls back to deterministic evidence extraction.
This is particularly important for document QA because unsupported generation is preferable to a confident hallucination.
Long evidence spans are compressed so that the final response focuses on the requested concept.
The final stage ensures the public service response maintains the expected contract:
{
"method": ...,
"answer": ...,
"citations": [...]
}semantic-document-intelligence/
│
├── apps/
│ └── question_answering/
│ ├── answer_service.py
│ └── ...
│
├── manage.py
│
├── README.md
├── semantic_document_intelligence_advanced_project_term_paper_Subhankar_Biswas_30008967.pdf
├── DEMONSTRATION.md
│
├── requirements.txt
├── docker-compose.yml
│
└── ...
The architecture combines application services, asynchronous processing, object storage, relational persistence, graph-based knowledge representation, and semantic RDF/SPARQL capabilities.
A Neo4j Graphical Representation of a processed document.
┌──────────────────────┐
│ User Question │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ AnswerService │
└──────────┬───────────┘
│
┌─────────────────┴─────────────────┐
▼ ▼
┌────────────────────┐ ┌────────────────────┐
│ Retrieval Service │ │ Question Analysis │
└─────────┬──────────┘ └─────────┬──────────┘
│ │
└────────────────┬─────────────────┘
▼
┌─────────────────────┐
│ Evidence Selection │
└──────────┬──────────┘
▼
┌─────────────────────┐
│ Span / Windowing │
└──────────┬──────────┘
▼
┌─────────────────────┐
│ Grounding Validator │
└──────────┬──────────┘
▼
┌────────────────────┐
│ LLM / Extractive │
│ Answer Generation │
└─────────┬──────────┘
▼
┌────────────────────┐
│ Compression + │
│ Final Contract │
└─────────┬──────────┘
▼
┌────────────────────┐
│ Answer + Citations │
└────────────────────┘
| Technology | Role | Official Documentation |
|---|---|---|
| Python | Application and NLP/data-processing language | python.org |
| Django | Web/application framework | djangoproject.com |
| Qdrant | Vector similarity search / retrieval infrastructure | qdrant.tech |
| OpenAI API | LLM-assisted answer generation | platform.openai.com |
| Docker | Containerized development and services | docker.com |
| Git | Version control | git-scm.com |
| Markdown | Project documentation | markdownguide.org |
The following badges are intentionally clickable so visitors can directly explore the technologies used by the project:
The badges use standard Shields.io badge URLs, which are suitable for GitHub README files. citeturn0search0turn0search2
Recommended environment:
- Python 3.12+
- Git
- Docker Desktop
- Docker Compose
- Access to the required LLM/API credentials
- A configured Qdrant instance or the project's Docker-based Qdrant service
git clone https://github.com/Subiswas36218/semantic-document-intelligence.git
cd semantic-document-intelligencepython3 -m venv venv
source venv/bin/activateOn Windows:
python -m venv venv
venv\Scripts\activatepip install -r requirements.txtCreate the project's environment file according to the variables expected by the application.
Example:
OPENAI_API_KEY=your_api_key_hereDo not commit secrets to GitHub.
If the project uses Docker Compose:
docker compose up -dCheck running services:
docker compose pspython manage.py checkOntology-grounded document intelligence with Django, Apache Jena/Fuseki, Neo4j/Cypher, Celery, PostgreSQL and MinIO.
Pipeline:
PDF/DOCX -> structured JSON -> LLM semantic extraction -> validation -> ontology mapping -> RDF/Jena + Neo4j projections -> graph retrieval -> source grounding -> answer.
The original document is preserved, JSON remains the exact source-recovery layer, and graph stores provide semantic retrieval.
cp .env.example .env
docker compose up --build
docker compose exec web python manage.py migrateServices:
- Django: http://localhost:8000
- Jena Fuseki: http://localhost:3030
- Neo4j Browser: http://localhost:7474
- MinIO: http://localhost:9001
Run checks:
make checkThe final regression demonstration uses the following five questions:
questions = [
"What are De Morgan’s laws?",
"What is a power set?",
"Explain measurable function composition.",
"What is a Cartesian product?",
"What are sigma fields?",
]Run:
python manage.py shell -c "
from apps.question_answering.answer_service import create_answer_service
service = create_answer_service()
questions = [
'What are De Morgan’s laws?',
'What is a power set?',
'Explain measurable function composition.',
'What is a Cartesian product?',
'What are sigma fields?',
]
try:
for question in questions:
result = service.answer(
question=question,
document_id='doc_fdd043b9648d',
)
print('=' * 100)
print('QUESTION:', question)
print('METHOD:', result['method'])
print('ANSWER:', result['answer'])
print('CITATIONS:')
for citation in result['citations']:
print(
' ->',
citation['citation_id'],
'paragraph=', citation['paragraph_id'],
'page=', citation['page'],
'quality=', citation['evidence_quality'],
'grounding=', citation['grounding_level'],
)
finally:
service.retrieval_service.close()
"The final regression run successfully produced answer-ready evidence for all five demonstration questions.
| Question | Method | Result |
|---|---|---|
| What are De Morgan’s laws? | extractive_fallback |
Answer extracted from page 1 |
| What is a power set? | extractive_fallback |
Answer extracted from page 3 |
| Explain measurable function composition. | extractive_fallback |
Answer extracted from page 6 |
| What is a Cartesian product? | extractive_fallback |
Answer extracted from page 3 |
| What are sigma fields? | extractive_fallback |
Evidence extracted from pages 4–5 |
Representative outputs include:
QUESTION: What is a power set?
METHOD: extractive_fallback
ANSWER:
Power Set: Power set = set of all subsets.
CITATIONS:
-> 1 paragraph=doc_fdd043b9648d_p003_001
page=3
quality=medium
grounding=dual_graph
and:
QUESTION: Explain measurable function composition.
METHOD: extractive_fallback
ANSWER:
Measurable Function Composition Given measurable spaces,
(S₁,F₁), (S₂,F₂), (S₃,F₃) ...
CITATIONS:
-> 1 paragraph=doc_fdd043b9648d_p006_001
page=6
quality=high
grounding=dual_graph
Each returned citation contains structured evidence metadata.
Example:
{
"citation_id": 1,
"paragraph_id": "doc_fdd043b9648d_p003_001",
"page": 3,
"evidence_quality": "medium",
"grounding_level": "dual_graph",
}This provides traceability from:
Question
↓
Answer
↓
Citation
↓
Paragraph
↓
Document Page
The system is deliberately designed not to answer every question from general model knowledge.
If the document does not contain sufficient answer-ready evidence, the service can return:
METHOD: no_evidence
ANSWER:
I could not find enough answer-ready evidence in the document
to answer this question reliably.
This behavior is an important part of the system's reliability design.
The final implementation was evaluated using a five-question regression suite covering:
- Set-theoretic laws
- Power sets
- Measurable functions
- Cartesian products
- Sigma fields
| Dimension | Objective |
|---|---|
| Retrieval relevance | Retrieve paragraphs related to the question |
| Concept precision | Isolate the requested concept |
| Evidence grounding | Keep the answer supported by retrieved evidence |
| Citation traceability | Connect answers to document paragraphs/pages |
| Conciseness | Avoid returning unnecessary surrounding context |
| Fallback safety | Prefer evidence extraction over unsupported generation |
| Contract stability | Preserve the expected answer/citation response structure |
The final regression output demonstrates that the system can:
- locate concept-relevant evidence;
- distinguish between answer-ready and insufficient evidence;
- extract compact answer spans;
- preserve document-level grounding;
- return paragraph/page citations;
- use deterministic extraction as a reliable fallback;
- reduce retrieval noise before final answer construction.
The final Question Answering implementation evolved through the following milestones:
Phase 4.4.3
↓
Concept-Specific Answer Span Extraction
Phase 4.4.4
↓
Improved answer-span selection and extraction
Phase 4.4.5
↓
Concept-Scoped Evidence Windowing
Phase 4.5
↓
Grounding / evidence-aware answer generation
Phase 4.6
↓
Confidence-aware answer selection and fallback behavior
Phase 4.7
↓
Answer compression and final answer-quality refinement
Phase 4.8
↓
Final answer contract / production-ready response path
Status: ✅ COMPLETE
The project is considered complete at the planned Phase 4.8 endpoint.
No additional Question Answering development phase is required for the current project scope.
Remaining activities are limited to:
- repository cleanup;
- committing the final implementation;
- adding project documentation;
- recording/presenting the demonstration;
- optional GitHub Actions/CI configuration;
- final submission.
| Document | Purpose |
|---|---|
README.md |
Project overview and technical documentation |
FINAL_REPORT.md |
Final evaluation and reporting |
DEMONSTRATION.md |
Step-by-step demonstration guide |
Before publishing:
[ ] README.md added
[ ] FINAL_REPORT.md added
[ ] DEMONSTRATION.md added
[ ] requirements.txt verified
[ ] docker-compose.yml verified
[ ] .env excluded from Git
[ ] API keys/secrets removed
[ ] temporary/debug files removed
[ ] final tests executed
[ ] Git status reviewed
[ ] GitHub repository pushed
Check for accidentally tracked secrets:
git status
git diff --cachedgit add README.md semantic_document_intelligence_advanced_project_term_paper_Subhankar_Biswas_30008967.pdf DEMONSTRATION.md
git add apps/
git add requirements.txt docker-compose.yml
git status
git commit -m "Finalize semantic document intelligence QA pipeline"
git push origin mainThe current project is complete, but potential future work could include:
- multi-document question answering;
- conversational follow-up questions;
- answer highlighting in the source document;
- richer citation previews;
- automated evaluation datasets;
- retrieval-quality benchmarking;
- confidence calibration dashboards;
- streaming answers;
- asynchronous document ingestion;
- web-based document exploration UI;
- automated CI regression testing.
These are future extensions, not required for the current project completion.
Add the project's chosen license here, for example:
MIT License
If a license has not yet been selected, do not claim a license until one is actually added to the repository.
Subhankar Biswas // M.Sc. Data Engineering // Constructor University, Bremen //
Semantic Document Intelligence — Evidence-Grounded Question Answering System
This project combines document retrieval, semantic search, evidence extraction, grounding validation, deterministic fallback strategies, and LLM-assisted question answering into a single document intelligence workflow.
Semantic Document Intelligence
Evidence first. Answers second.