Aller au contenu principal
Accès ouvert déclaré 2026 article

Reframing Historical Text Extraction: A Cross-Pathway Validation of OCR, LLM-Assisted Correction, and Direct Multimodal Transcription

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Historical document collections are increasingly available as digitised images and PDFs, but their conversion into reliable text remains affected by optical character recognition (OCR) errors, degraded pages, heterogeneous layouts, and domain-specific terminology. This study proposes a pathway-level framework for documenting and comparing conventional OCR, OCR followed by large language model (LLM)-assisted correction, and direct multimodal transcription. The framework is demonstrated using the Portuguese Agricultural and Forestry Surveys (1950–1958). A stratified validation sample of 45 pages was selected by visual quality, page type, and geographic coverage. Outputs were evaluated against manually verified reference transcriptions using content-normalised character error rate (CER) and word error rate (WER), document-condition analysis, paired statistical tests, and an entity-level semantic preservation assessment focused on place names, agricultural terms, and measurement expressions. Under the evaluated model and interface conditions, both LLM-based pathways produced lower mean CER and WER than the conventional OCR baseline, with the lowest values observed for direct multimodal transcription. Semantic preservation was also higher for the LLM-based pathways, although measurement expressions remained the most persistent risk, particularly in table-based pages. Downstream tasks were not directly evaluated. The findings support the framework as a method for validating text-extraction pathways before reuse, rather than establishing a universal ranking of tools.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Reframing Historical Text Extraction: A Cross-Pathway Validation of OCR, LLM-Assisted Correction, and Direct Multimodal Transcription
Date Crossref
25/07/2026
Éditeur
MDPI AG
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Handwritten Text Recognition TechniquesImage Retrieval and Classification TechniquesGeographic Information Systems Studies

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.