Overview & Executive Summary
Paper: DocEng Short Paper - Cost-Aware Human-LLM Collaboration for Post-OCR Corrections
Authors: Stergios Konstantinidis, Hayman Lotfy, Michalis Vlachos (University of Lausanne)
Venue: ACM Symposium on Document Engineering (DocEng) - Short Paper
GitHub: https://github.com/Stergios-Konstantinidis/Cost-Aware-Human-LLM-Collaboration-for-post-OCR-Corrections
Overleaf: https://git.overleaf.com/6a2cb93de1bec86186ee8613
Authors: Stergios Konstantinidis, Hayman Lotfy, Michalis Vlachos (University of Lausanne)
Venue: ACM Symposium on Document Engineering (DocEng) - Short Paper
GitHub: https://github.com/Stergios-Konstantinidis/Cost-Aware-Human-LLM-Collaboration-for-post-OCR-Corrections
Overleaf: https://git.overleaf.com/6a2cb93de1bec86186ee8613
1. Research Motivation & Core Problem
Historical document archives, such as digitized centuries-old Swiss newspapers, suffer from substantial OCR degradation caused by aged paper, bleed-through, complex multi-column typography, and antique fonts.
While modern Large Language Models (LLMs) excel at OCR correction, naive all-to-LLM pipelines suffer from three severe pitfalls:
- Financial Cost: Processing millions of archive pages with commercial LLM APIs is economically non-viable.
- Degradation of Clean Text: LLMs frequently "hallucinate" or modernize archaic French spellings on segments that were already correctly transcribed by the OCR engine.
- Severe OCR Failures: On heavily damaged regions, LLMs fabricate plausible-sounding text instead of preserving historical fidelity.
This paper formulates OCR post-correction as an optimal budget allocation and routing problem uniting human domain experts and LLMs.
2. Key Contributions
- Three-Tier Human-LLM Collaboration: Explicit routing of each segment to either *(1) No Correction*, *(2) Automated LLM Correction*, or *(3) Human Review*.
- Regression-Guided Prioritization: Learned Lasso/Ridge router predicting expected Character Error Rate (CER) reduction \hat{Δ}_i before sending tokens to an LLM.
- Post-LLM Safeguard Layer: Classification mechanism detecting when LLM correction introduced regression or hallucination, diverting compromised outputs to human reviewers.
- Superior Budget Efficiency: Achieves 14% relative CER reduction over the All-LLM baseline with <5% human review budget.
No comments to display
No comments to display