Post-OCR correction using large language models with constrained decoding

Sastre, Ignacio - Etcheverry, Lorena - Rey, Guillermo - Moncecchi, Guillermo - Rosá, Aiala

Resumen:

This article addresses the problem of correcting noisy Optical Character Recognition (OCR) outputs from digitized historical documents, specifically those from the Berrutti Archive related to Uruguay’s civic-military dictatorship. These documents—produced with typewriters, diverse layouts, and overlaid annotations—pose significant hallenges for standard OCR tools, resulting in highly errorprone text. We present a novel post-OCR correction method that leverages fine-tuned open-source Large Language Models (LLMs) combined with a constrained decoding strategy. This strategy incorporates character-level similarity between the OCR input and the generated output at decoding time, steering the model toward corrections that closely preserve the original text structure. We evaluate our method on a gold-standard dataset of over 2000 annotated lines and show that it outperforms prompting and standard fine-tuning approaches, reducing both character error rate (CER) and word error rate (WER). The corrected outputs provide more accurate input for downstream tasks, such as named entity recognition, relation and event extraction, and knowledge graph construction, thereby supporting the broader goal of extracting knowledge from historically significant and sensitive archives.

Detalles Bibliográficos
2025
Proyecto ANII IA_1_2022_1_173863
Natural Language Processing
Optical Character Recognition (OCR)
Post-OCR Correction
LLMs with Constrained Decoding
Inglés
Universidad de la República
COLIBRI
https://hdl.handle.net/20.500.12008/56413
Acceso abierto
Licencia Creative Commons Atribución - No Comercial - Sin Derivadas (CC - By-NC-ND 4.0)