Aller au contenu principal
Accès ouvert déclaré 2026 dataset

Benchmarking LLM-based Information Extraction Tools for Medical Documents

0Citations signalées — pas une note de qualité
3Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Background: Medical documents are a crucial resource for medical research around the world. While troves of valuable health data exist, they are largely computationally inaccessible as hard copies of unstructured text. Moreover, the persistent prevalence of fax machines in medical settings contributes to further degradation of document quality. Manual data extraction from these documents is time-consuming and resource intensive. However, large language models (LLMs) have recently shown great promise for automated digitization and information extraction (IE), greatly improving upon previous technologies in terms of accuracy. Objective: Here we aim to survey the landscape of LLM-based IE tools with respect to their applicability to biomedical documents and to systematically benchmark them on synthetic documents with paired reference extractions. We put special focus on locally executable tools and models that suit privacy- and resource-constrained settings. Methods: We reviewed LLM-based IE tools from the literature and assessed them with respect to their suitability for use in biomedical research. We found only one of these tools (NuExtract2) to satisfy our selection criteria and compared it to LLM foundation models prompted to perform extractions. We created 200 mock molecular test reports with paired reference data using a bespoke procedural generation approach and evaluated the tools’ performance across different prompting strategies, input modalities, and document qualities. Results: We found model performance to be very sensitive to input modalities and quality. The best overall performance was observed for the proprietary GPT 4.1-mini, with an F1 score of 79.4% averaged across quality levels and modalities. The best performing open-weights model was Mistral-small 3.1 with an average F1 score of 76.5%. The NuExtract2 model was the only one able to run on regular laptop computers, but did not perform well. Surprisingly, we found the choice of one-shot over zero-shot prompts to only have a small effect size on extraction performance in most cases. Conclusions: The unique constraints of processing highly complex medical documents in resource-constrained and privacy-sensitive settings still hamper their effectiveness and necessitate human oversight. Availability: Source code available on Github at https://github.com/courtotlab/extraction-benchmark

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.