Critical Reading Exam: Archival Foundations & Doctoral Portfolio
π Examination: PhD Critical Reading Exam (Doctoral Commission, June 2026)
π€ Candidate: Stergios Konstantinidis β’ Department of Information Systems (DESI), HEC Lausanne, University of Lausanne (UNIL)
π Dissertation: Large Language Models for Historical Document Intelligence: Retrieval-Augmented Generation, Automated Evaluation, and Agentic Architectures
π Repositories: GitHub CRE Repository β’ Overleaf Exam Workspace
π₯ Primary Document: Download Critical Reading Exam Paper (PDF, 8 Pages)
1. The 5 Literature Review Articles (Critical Reading Exam)
The exam selection examines five external cornerstone publications spanning four interconnected theoretical pillars:
| Article | Primary Contribution | Addressed Architectural Vulnerability | Direct PhD Synthesis Link |
|---|---|---|---|
| Tran et al. (JCDL 2024) RAG for Historical Newspapers |
First end-to-end RAG system over multilingual historical newspapers (NewsEye); proposed hybrid dense retrieval and LLM-based ground-truth-free evaluation. | Vocabulary mismatch in OCR noise; absence of labeled QA benchmarks. | Serves as our primary baseline; exposed the fragility of keyword search under OCR degradation and motivated our VP-tree metric pruning. |
| Guan et al. (EMNLP 2024) Effective Synthetic Data & Test-Time Adaptation |
Principled synthetic OCR degradation generation and test-time adaptation targeting proper nouns and out-of-vocabulary entities. | Severe training data scarcity for historical periods; catastrophic forgetting of rare historical entities. | Informs our DocEng feature modeling; inspired our post-correction safeguard to prevent modernizing archaic nomenclature. |
| Gu et al. (The Innovation 2026) A Survey on LLM-as-a-Judge |
Systematic taxonomy of LLM judges: position bias, verbosity bias, self-enhancement bias, and calibration protocols. | Unreliable evaluation in ground-truth-free RAG evaluation pipelines. | Establishes the methodological foundation for validating automated evaluation in our IP&M journal pipeline. |
| Sun et al. (EMNLP 2025) DocAgent: Multi-Modal Long-Context Framework |
Agentic architecture featuring selective retrieval, multimodal document inspection tools, and iterative answer verification agents. | Context window limits; inability of pure text LLMs to resolve visual layout artifacts. | Directly anticipates our Apple Vision Pro spatial agent and multi-turn historical inquiry workflows. |
| Lewis et al. (NeurIPS 2020) Foundational Retrieval-Augmented Generation |
The foundational dual-encoder dense retrieval and seq2seq generator architecture trained end-to-end with latent documents. | Hallucination in closed-book parametric memory; inability to update knowledge. | The theoretical origin of RAG; reveals how modern assumptions (clean English Wikipedia) break down when deployed over 300 years of degraded archives. |
2. Theoretical & Architectural Synthesis
The Foundational Tension in Historical Document Intelligence
Across all reviewed articles and our doctoral systems, a persistent architectural tension emerges: the fundamental assumptions that make LLM-based document intelligence effective are systematically violated by the historical archives that most urgently need it:
- Assumption of Orthographic Cleanliness: Standard RAG models assume high-quality input text. In 18th-to-20th-century historical newspapers, character error rates (CER) frequently exceed 15β25% due to font fading, broken typefaces, ink bleed, and complex multi-column typography. Our DocEng and Guan et al. research demonstrates that dense vector spaces must be explicitly regularized against OCR degradation.
- Assumption of Monolingual Modernity: Foundational retrievers are optimized for contemporary English. Regional archives (e.g., Swiss historical press) encompass archaic French, Swiss German dialectal shifts, and evolving syntactic conventions across three centuries. As shown in Tran et al. and our IP&M pipeline, multilingual embeddings combined with metric tree indexing are essential to bridge temporal linguistic drift.
- Assumption of Annotation Abundance: Modern benchmarks rely on millions of human QA annotations. Archival collections possess zero labeled question-answering pairs. Reconciling this gap requires robust ground-truth-free evaluation frameworks (Tran et al., Gu et al.) carefully calibrated against prompt bias and verbosity artifacts.
- Assumption of Flat Screen Interfaces: Traditional document retrieval presents flat lists of snippets. Navigating centuries of interrelated historical narratives demands agentic reasoning (DocAgent) and immersive spatial computing (AppleVision Cultural Archives) that turn passive archival retrieval into active sensemaking.
3. Research Milestone & Examination Context
- Doctoral Commission: Faculty of Business and Economics (HEC Lausanne), University of Lausanne.
- Examination Date: June 2026.
- Outcome: Passed
4. The 5 Doctoral Papers in this Research Portfolio ("Our Papers")
The doctoral dissertation synthesizes five authored publications addressing the architectural failure modes identified in the literature, evaluated directly on the multi-century collections of the Bibliothèque Cantonale et Universitaire de Lausanne (BCUL):
| Venue / Track | Authored Publication Title | Addressed Architectural Limitation | Primary Artifact |
|---|---|---|---|
| ACM IUI 2026 Intelligent User Interfaces |
AppleVision Cultural Archives: Immersive Spatial Exploration | Replaces flat retrieval lists with an immersive 3D spatial canvas (visionOS), multi-document spatial clustering, and multimodal gaze-and-pinch interaction. | π₯ Download PDF (10 MB) |
| Elsevier IP&M Information Processing & Management |
BCUL Historical Newspaper Processing Pipeline | Provides complete archival processing: semantic layout segmentation, dual-stage OCR confidence scoring, hybrid dense retrieval, and temporal cross-encoder re-ranking. | π₯ Download PDF (5.1 MB) |
| ECML PKDD 2026 Machine Learning & Data Mining |
Historical Document Retrieval & Metric Learning | Applies metric space learning and VP-tree spatial partitioning over Swiss French and German press, pruning non-viable retrieval candidates and cutting latency by >45%. | π₯ Download PDF (5.4 MB) |
| EDBT 2026 Database Technology Demo |
Interactive System Demonstration: 300 Years at Scale | Live interactive demonstration of end-to-end historical question-answering over 12 TB of high-resolution scan imagery and full-text transcripts. | π₯ Download PDF (3.2 MB) |
| ACM DocEng 2026 Document Engineering |
Cost-Aware Collaborative Human-LLM Post-OCR Correction | Dynamic three-tier degradation router allocating text segments between bypass, LLM correction, and expert human review (<5% budget) with post-correction safeguards. | π₯ Download PDF (2.5 MB) |