Opportunities and challenges in automated coding of electronic health records: a pilot study for rare disease registries
Rattachement africain : it, cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Accurate medical coding is essential for disease registries, particularly in the context of rare conditions. Manually transforming electronic health records data into standardized codes is time-consuming and resource intensive. This study evaluates an automated coding system using synthetic health records data and explores the potential benefits and challenges of introducing this tool into rare diseases registries activities. We developed a hybrid architecture combining a symbolic component with medical knowledge graphs and an ensemble of three widely used Large Language Models with a critical review mechanism. Ninety-nine synthetic Italian-language clinical reports were coded by the system. Subsequently, a multidisciplinary expert panel performed a double-coding validation of extracted terms, categorizing automated results into four groups: correct, incorrect, inaccurate, or missing codes. The system extracted a total of 479 terms (264 diagnosis codes and 215 procedure codes) mapped to ICD-9-CM classification. The expert panel, considered as the gold standard, identified 500 terms (302 diagnosis codes and 198 procedure codes). Chi-square analysis highlighted statistically significant differences between diagnosis and procedure coding in at least one of the four groups of results ( p =0.001). The system achieved an accuracy of 70.53% for diagnoses, compared to 78.28% for procedures. Additionally, the relative frequency of the various incorrect codes is generally consistent and uniform, except for two incorrect procedure codes that are particularly prevalent. Considering all the findings, we critically point out the potential contribution and impact of an automated coding system into rare diseases registries process, examining the benefits and barriers that could facilitate or hamper progress in this specialized field. The automated coding system demonstrated reasonable accuracy with health records synthetic data. Key challenges include limited ICD-9-CM codes, particularly for rare diseases, overreliance on nonspecific residual codes, and tendency to generate details not present in reports. Opportunities for future improvement may include adopting the ICD-10/ICD-11 classification, implementing reliability metrics and a multi-ontology approach, thus promoting data interoperability according to FAIR principles. An automated coding system, properly improved, may have an essential impact for rare disease registries and are welcomed by several initiatives such as the EHDS.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Opportunities and challenges in automated coding of electronic health records: a pilot study for rare disease registries
- Date Crossref
- 08/07/2026
- Éditeur
- Frontiers Media SA
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.