Skip to main content

Critical Reading Exam: Archival Foundations & Doctoral Portfolio

🎓 Examination: PhD Critical Reading Exam (Doctoral Commission, June 2026)

👤 Candidate: Stergios Konstantinidis • Department of Information Systems (DESI), HEC Lausanne, University of Lausanne (UNIL)

📖 Dissertation: Large Language Models for Historical Document Intelligence: Retrieval-Augmented Generation, Automated Evaluation, and Agentic Architectures

🔗 Repositories: GitHub CRE Repository • Overleaf Exam Workspace

 

1. The 5 Core Authored Doctoral Papers ("Our Papers")

The doctoral research builds directly upon five peer-reviewed publications and conference submissions establishing end-to-end historical document intelligence across the 300-year archives of the Bibliothèque Cantonale et Universitaire de Lausanne (BCUL):

Venue / Track Publication Title Core Contribution Download ACM IUI 2026
Intelligent User Interfaces AppleVision Cultural Archives: Immersive Spatial Exploration First mixed-reality spatial computing environment (Apple visionOS) for historical archives, supporting 3D timeline navigation, multimodal gaze-and-pinch querying, and multi-document spatial synthesis. 📥 Download PDF (10 MB) Elsevier IP&M
Information Processing & Management BCUL Historical Newspaper Processing Pipeline Comprehensive archival processing journal: layout semantic segmentation, dual-stage OCR confidence estimation, dense hybrid retrieval, and temporal cross-encoder re-ranking across centuries of Swiss press. 📥 Download PDF (5.1 MB) ECML PKDD 2026
Machine Learning & Data Mining Historical Document Retrieval & Metric Learning Metric space learning and VP-tree spatial partitioning over multi-century Swiss French and German newspapers, pruning non-viable retrieval candidates and reducing retrieval latency by >45%. 📥 Download PDF (5.4 MB) EDBT 2026
Database Technology Demo Interactive System Demonstration: 300 Years at Scale Live interactive search and historical question-answering demonstration over 12 TB of high-resolution scan imagery and full-text transcripts from BCUL collections. 📥 Download PDF (3.2 MB) ACM DocEng 2026
Document Engineering Cost-Aware Collaborative Human-LLM Post-OCR Correction Adaptive three-tier degradation router allocating text segments between bypass, LLM correction, and expert human verification under a strict <5% human budget with post-correction safeguards. 📥 Download PDF (2.5 MB)

2. The 5 Literature Review Articles (Critical Reading Exam)

The exam selection examines five external cornerstone publications spanning four interconnected theoretical pillars:

Article Primary Contribution Addressed Architectural Vulnerability Direct PhD Synthesis Link
Tran et al. (JCDL 2024)
RAG for Historical Newspapers
First end-to-end RAG system over multilingual historical newspapers (NewsEye); proposed hybrid dense retrieval and LLM-based ground-truth-free evaluation. Vocabulary mismatch in OCR noise; absence of labeled QA benchmarks. Serves as our primary baseline; exposed the fragility of keyword search under OCR degradation and motivated our VP-tree metric pruning.
Guan et al. (EMNLP 2024)
Effective Synthetic Data & Test-Time Adaptation
Principled synthetic OCR degradation generation and test-time adaptation targeting proper nouns and out-of-vocabulary entities. Severe training data scarcity for historical periods; catastrophic forgetting of rare historical entities. Informs our DocEng feature modeling; inspired our post-correction safeguard to prevent modernizing archaic nomenclature.
Gu et al. (The Innovation 2026)
A Survey on LLM-as-a-Judge
Systematic taxonomy of LLM judges: position bias, verbosity bias, self-enhancement bias, and calibration protocols. Unreliable evaluation in ground-truth-free RAG evaluation pipelines. Establishes the methodological foundation for validating automated evaluation in our IP&M journal pipeline.
Sun et al. (EMNLP 2025)
DocAgent: Multi-Modal Long-Context Framework
Agentic architecture featuring selective retrieval, multimodal document inspection tools, and iterative answer verification agents. Context window limits; inability of pure text LLMs to resolve visual layout artifacts. Directly anticipates our Apple Vision Pro spatial agent and multi-turn historical inquiry workflows.
Lewis et al. (NeurIPS 2020)
Foundational Retrieval-Augmented Generation
The foundational dual-encoder dense retrieval and seq2seq generator architecture trained end-to-end with latent documents. Hallucination in closed-book parametric memory; inability to update knowledge. The theoretical origin of RAG; reveals how modern assumptions (clean English Wikipedia) break down when deployed over 300 years of degraded archives.

3.2. Theoretical & Architectural Synthesis

The Foundational Tension in Historical Document Intelligence

Across all reviewed articles and our doctoral systems, a persistent architectural tension emerges: the fundamental assumptions that make LLM-based document intelligence effective are systematically violated by the historical archives that most urgently need it:

  1. Assumption of Orthographic Cleanliness: Standard RAG models assume high-quality input text. In 18th-to-20th-century historical newspapers, character error rates (CER) frequently exceed 15–25% due to font fading, broken typefaces, ink bleed, and complex multi-column typography. Our DocEng and Guan et al. research demonstrates that dense vector spaces must be explicitly regularized against OCR degradation.
  2. Assumption of Monolingual Modernity: Foundational retrievers are optimized for contemporary English. Regional archives (e.g., Swiss historical press) encompass archaic French, Swiss German dialectal shifts, and evolving syntactic conventions across three centuries. As shown in Tran et al. and our IP&M pipeline, multilingual embeddings combined with metric tree indexing are essential to bridge temporal linguistic drift.
  3. Assumption of Annotation Abundance: Modern benchmarks rely on millions of human QA annotations. Archival collections possess zero labeled question-answering pairs. Reconciling this gap requires robust ground-truth-free evaluation frameworks (Tran et al., Gu et al.) carefully calibrated against prompt bias and verbosity artifacts.
  4. Assumption of Flat Screen Interfaces: Traditional document retrieval presents flat lists of snippets. Navigating centuries of interrelated historical narratives demands agentic reasoning (DocAgent) and immersive spatial computing (AppleVision Cultural Archives) that turn passive archival retrieval into active sensemaking.

4.3. Research Milestone & Examination Context

  • Doctoral Commission: Faculty of Business and Economics (HEC Lausanne), University of Lausanne.
  • Examination Date: June 2026.
  • Outcome: Validation of the theoretical, empirical, and architectural foundations for the completion and defense of the PhD dissertation.Passed