Probabilistic Record Linkage for Arabic-Script Archival Material: A Fellegi–Sunter Framework with Patronymic-Chain Features
Rattachement africain : ca, in. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Author's original version (preprint). Submitted to Digital Scholarship in the Humanities (Oxford University Press), 25 August 2026, manuscript DSH-2026-1071. Cross-source record linkage in Arabic-script archival material poses challenges that mainstream genealogical and prosopographical infrastructure does not address: multi-token patronymic naming (ism, nasab, nisba) interacts with heterogeneous transliteration practice, so the same individual appears under many inconsistent surface forms across colonial-era and pre-modern sources. We present a reference implementation of the Fellegi-Sunter probabilistic framework with eight Arabic-aware comparison features, each assigned a log-likelihood-ratio weight from domain-tuned initial probabilities, with Beta(1,1)-smoothed Bayesian re-estimation from archivist review decisions. On a two-corpus synthetic test set the pipeline achieves 100% mention recall on narrative material and 87.5% cross-source merge recall with 100% precision on observed merges; per-fact provenance coverage, an invariant of the architecture, is verified at 100% over the consolidated graph.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.