Aller au contenu principal
Accès ouvert déclaré 2025 article

Extending CARDIO:DE: Additional annotation guidelines and evaluation of NLP approaches for clinical applications

2Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : de. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

BACKGROUND: Cardiovascular diseases are a major cause of morbidity and mortality, and the management of these conditions generates extensive clinical data. The CARDIO:DE dataset, a German-language corpus of cardiovascular clinical routine letters, has been developed to support natural language processing research. This study seeks to enhance the dataset by introducing refined annotation guidelines and expanding the annotation schema. OBJECTIVE: The objective of this study was to extend the CARDIO:DE dataset with additional annotation categories, and evaluate state-of-the-art NLP models to enhance the utility of the dataset for clinical applications. METHODS: The annotation schema was expanded to include categories such as diagnostic procedures, medical finding, and therapeutic interventions (Diagnostic, Diagnosis, Drug, Medical_Finding, Therapy). The iterative annotation process involved expert annotators, ensuring high-quality, consistent annotations. Four models-GBERT, medBERT.de, XLM-RoBERTa, and TinyLlama-were fine-tuned and evaluated on the dataset. Model performance was assessed using entity-wise precision, recall, and F1 scores. RESULTS: The extended dataset includes 304,582 token-based annotations, with the highest concentration in medical finding. The inter-annotator agreement scores improved during the iterative process, reaching up to 0.98 for certain subsets. Among the evaluated models, TinyLlama outperformed the other models in entity recognition, achieving a macro-average F1 score of 0.845, highlighting its potential for clinical NLP tasks. CONCLUSIONS: The extended CARDIO:DE dataset, with its refined annotation guidelines provides a robust foundation for natural language processing applications in the clinical domain. The performance of the TinyLlama model demonstrates the potential of fine-tuning non-domain-specific models for clinical text processing. This work paves the way for more accurate NLP solutions in healthcare, particularly for information extraction and decision support in cardiology.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Extending CARDIO:DE: Additional annotation guidelines and evaluation of NLP approaches for clinical applications
Date Crossref
01/11/2025
Éditeur
Elsevier BV
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Topic ModelingBiomedical Text Mining and OntologiesMachine Learning in Healthcare

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.