Aller au contenu principal
2026 preprint

Large language and rule-based models for diabetes diagnosis date extraction from electronic medical records : comparative study (Preprint)

0Citations signalées — pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Résumé fourni par la source

BACKGROUND Diabetes duration is a major determinant of complications and an essential variable for clinical research, yet the diagnosis date is frequently unavailable in structured electronic health records (EHRs). Although large language models (LLMs) have shown promise for clinical information extraction, their added value over established rule-based approaches for this task remains unclear. OBJECTIVE This study aimed to compare the performance and computational efficiency of prompting-based LLMs and a rule-based extraction method for identifying diabetes diagnosis dates from French clinical notes. METHODS Clinical notes of patients with diabetes (600 type 1 and 100 type 2) were extracted from the Assistance Publique-Hôpitaux de Paris Clinical Data Warehouse and manually annotated. Notes of type 1 diabetes were divided into training (n=200), validation (n=200) and test (n=200) sets; notes of type 2 were only used as test set. We compared a rule-based ContextualMatcher with five open-weight LLMs (Qwen3-8B, LLaMA-3.1-8B-Instruct, Ministral-8B-Reasoning, Ministral-14B-Reasoning, and MedGemma-27B) using zero-shot and few-shot prompting, with reasoning enabled when supported. Predictions within ±1 year of the reference annotation were considered correct. Performance was evaluated using F1-score, precision, recall, balanced accuracy, specificity, hallucination rate, no-prediction rate, mean absolute error (MAE), and inference time. RESULTS On the type 1 diabetes test set, MedGemma-27B achieved the highest F1-score (0.94, 95% CI 0.91–0.96), followed by Qwen3-8B with reasoning (0.93, 95% CI 0.90–0.96), compared with 0.86 (95% CI 0.82–0.90) for the best ContextualMatcher. LLMs showed the greatest advantage for relative diagnosis dates (best F1 1.00 vs 0.82). Performance generalized to type 2 diabetes despite model development being conducted exclusively on type 1 notes, with MedGemma-27B achieving an F1-score of 0.96 and the best ContextualMatcher 0.89. Few-shot prompting did not consistently improve performance, and dedicated reasoning models performed inconsistently. The rule-based approach processed notes in 0.01–0.02 seconds using CPUs only, whereas LLM inference required 0.2–2.0 seconds per note and two NVIDIA A100 GPUs. Qwen3-8B with reasoning required approximately 50-fold longer inference time than the best ContextualMatcher. CONCLUSIONS Prompting-based LLMs achieved the highest accuracy for diabetes diagnosis date extraction from French clinical notes but provided only modest performance gains over a well-tuned rule-based approach while requiring substantially greater computational resources. For large-scale EHR research, rule-based methods remain an efficient and competitive option, whereas LLMs may be most valuable when maximizing extraction accuracy for linguistically heterogeneous expressions justifies the additional computational cost.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Large language and rule-based models for diabetes diagnosis date extraction from electronic medical records : comparative study (Preprint)
Date Crossref
14/08/2026
Éditeur
JMIR Publications Inc.
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Sujets associés

Machine Learning in HealthcareArtificial Intelligence in HealthcareGenomics and Rare Diseases

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.