Skip to main content

Critical Reading Exam

🎓 Examination: PhD Critical Reading Exam (Doctoral Commission, June 2026)

👤 Candidate: Stergios Konstantinidis • Department of Information Systems (DESI), HEC Lausanne, University of Lausanne (UNIL)

📖 Dissertation:Topic: Large Language Models for Historical Document Intelligence: Retrieval-Augmented Generation, Automated Evaluation, and Agentic Architectures

🔗 Repositories: GitHub CRE Repository • Overleaf Exam Workspace

📥 Primary Examination Documents: Exam Paper (PDF, 8 Pages) • Presentation Deck (PDF, 13 Slides) • Presentation Deck (PPTX, 11 MB) • View Presentation Slides →

1. The 5 Literature Review Articles (Critical Reading Exam)

The exam selection systematically examines five external cornerstone publications spanning four interconnected theoretical pillars:pillars in historical document intelligence:

Article Primary Contribution Addressed Architectural Vulnerability DirectArchival PhD& SynthesisMethodological LinkRelevance
Tran et al. (JCDL 2024)
RAG for Historical Newspapers
First end-to-end RAG system over multilingual historical newspapers (NewsEye); proposed hybrid dense retrieval and LLM-basedNER-enhanced re-ranking with ground-truth-free evaluation.LLM metrics. VocabularyDegraded mismatchOCR query drift; severe annotation scarcity in OCRhistorical noise; absence of labeled QA benchmarks.corpora. Serves as our primary baseline; exposedEstablishes the fragilityproblem ofsetting: keywordmultilingual, searchOCR-corrupted underarchival OCRtext degradationrequires hybrid dense retrieval and motivatedrobust ourreference-free VP-tree metric pruning.evaluation.
Guan et al. (EMNLPAAAI 2024)
Effective Synthetic DataOCR Degradation & Test-TimeRobust AdaptationRetrieval
Principled synthetic OCR degradation generation and test-time adaptation targeting proper nouns and out-of-vocabulary entities. Severe training data scarcity for historical periods; catastrophic forgetting of rare historical entities. InformsProvides ourmethodology DocEngfor featuresynthetic modeling;degradation inspiredmodeling ourand post-correctiontest-time safeguard mechanisms to prevent modernizinghallucinated modernisation of archaic nomenclature.
Gu et al. (The Innovation 2026)
A Survey on LLM-as-a-Judge
Systematic taxonomy of LLM judges: position bias, verbosity bias, self-enhancement bias, and calibration protocols. Unreliable evaluation in ground-truth-free RAG evaluation pipelines. Establishes the theoretical and methodological foundation for validating automated evaluation metrics in ourground-truth-free IP&Marchival journal pipeline.pipelines.
Sun et al. (EMNLP 2025)
DocAgent: Multi-Modal Long-Context Framework
Agentic architecture featuring selective retrieval, multimodal document inspection tools, and iterative answer verification agents. Context window limits; inability of pure text LLMs to resolve visual layout artifacts. DirectlyDemonstrates anticipateshow ouragentic Appledecomposition, Visionmultimodal Protool spatial agentuse, and multi-turn historicalverification inquiryovercome workflows.the limitations of flat passage retrieval.
Lewis et al. (NeurIPS 2020)
Foundational Retrieval-Augmented Generation
The foundational dual-encoder dense retrieval and seq2seq generator architecture trained end-to-end with latent documents. Hallucination in closed-book parametric memory; inability to update knowledge. The theoretical origin of RAG; reveals how modern assumptions (clean English Wikipedia) break down when deployed over 300 yearscenturies of degraded archives.

2. Theoretical & Architectural Synthesis

The Foundational Tension in Historical Document Intelligence

Across all reviewed articles and our doctoral systems,articles, a persistent architectural tension emerges: the fundamental assumptions that make LLM-based document intelligence effective are systematically violated by the historical archives that most urgently need it:

  1. Assumption of Orthographic Cleanliness: Standard RAG models assume high-quality input text. In 18th-to-20th-century historical newspapers, character error rates (CER) frequently exceed 15–25% due to font fading, broken typefaces, ink bleed, and complex multi-column typography. Our DocEng and Guan et al. research demonstrates that denseDense vector spaces must therefore be explicitly regularized against OCR degradation.
  2. Assumption of Monolingual Modernity: Foundational retrievers are optimized for contemporary English. RegionalArchival archives (e.g., Swiss historical press)collections encompass archaic French,dialects, Swisshistorical Germanspelling dialectal shifts,variants, and evolving syntactic conventions across threemultiple centuries. As shown in Tran et al. and our IP&M pipeline, multilingualMultilingual embeddings combined with metricrobust treespatial indexing are essential to bridge temporal linguistic drift.
  3. Assumption of Annotation Abundance: Modern benchmarks rely on millions of human QA annotations. Archival collections possess virtually zero labeled question-answering pairs. Reconciling this gap requires robust ground-truth-free evaluation frameworks (Tran et al., Gu et al.) carefully calibrated against prompt biasbias, verbosity artifacts, and verbositypositional artifacts.skew.
  4. Assumption of Flat Screen Interfaces: Traditional document retrieval presents flat lists of snippets. Navigating centuries of interrelated historical narratives demands agentic reasoningreasoning, (DocAgent)multimodal layout inspection, and immersiveinteractive spatial computing (AppleVision Cultural Archives)exploration that turn passive archival retrieval into active sensemaking.

3. Research Milestone & Examination Context

  • Doctoral Commission: Faculty of Business and Economics (HEC Lausanne), University of Lausanne.
  • Examination Date: June 2026.
  • Outcome: Passed

4. The 5 Doctoral Papers in this Research Portfolio ("Our Papers")

The doctoral dissertation synthesizes five authored publications addressing the architectural failure modes identified in the literature, evaluated directly on the multi-century collections of the Bibliothèque Cantonale et Universitaire de Lausanne (BCUL):

Venue / Track Authored Publication Title Addressed Architectural Limitation Primary Artifact ACM IUI 2026
Intelligent User Interfaces AppleVision Cultural Archives: Immersive Spatial Exploration Replaces flat retrieval lists with an immersive 3D spatial canvas (visionOS), multi-document spatial clustering, and multimodal gaze-and-pinch interaction. 📥 Download PDF (10 MB) Elsevier IP&M
Information Processing & Management BCUL Historical Newspaper Processing Pipeline Provides complete archival processing: semantic layout segmentation, dual-stage OCR confidence scoring, hybrid dense retrieval, and temporal cross-encoder re-ranking. 📥 Download PDF (5.1 MB) ECML PKDD 2026
Machine Learning & Data Mining Historical Document Retrieval & Metric Learning Applies metric space learning and VP-tree spatial partitioning over Swiss French and German press, pruning non-viable retrieval candidates and cutting latency by >45%. 📥 Download PDF (5.4 MB) EDBT 2026
Database Technology Demo Interactive System Demonstration: 300 Years at Scale Live interactive demonstration of end-to-end historical question-answering over 12 TB of high-resolution scan imagery and full-text transcripts. 📥 Download PDF (3.2 MB) ACM DocEng 2026
Document Engineering Cost-Aware Collaborative Human-LLM Post-OCR Correction Dynamic three-tier degradation router allocating text segments between bypass, LLM correction, and expert human review (<5% budget) with post-correction safeguards. 📥 Download PDF (2.5 MB)