Aller au contenu principal
Accès ouvert déclaré 2026 article

Evaluating large language model performance in US FDA regulatory science

0Citations signalées — pas une note de qualité
7Institutions déclarées
3Pays d’affiliation déclarés

Résumé fourni par la source

Abstract Background Clinical and population decision-making relies on the systematic evaluation of extensive regulatory evidence. The FDA drug reviews provide detailed information on clinical trial design, enrollment criteria, sample size, randomization, comparators, endpoints, and indications. However, extracting these data is resource-intensive and time-consuming. Generative Artificial Intelligence large language models (LLMs) may accelerate the extraction and synthesis of such information. This study compares the performance of three LLMs, ChatGPT-4o, Gemini 2.5 Pro, and DeepSeek R1, in extracting, analyzing, and synthesizing regulatory and clinical information from FDA drug reviews, guidance for the industry, and drug labels as accessed through their standard user interfaces, using antibiotics approved for complicated urinary tract infections (cUTIs) between 2010 and 2025. Methods LLMs were evaluated using general (short, direct) and detailed (structured, guidance-referencing) prompts across five domains including accuracy (precision and recall), explanation quality, error type (hallucination rate, misclassification, and omission), operational efficiency (response time, correct answers per second, and seconds per correct answer), and consistency with responses generated in duplicate runs. Two investigators independently reviewed outputs against FDA guidance, resolving discrepancies by consensus. Statistical analyses included χ 2 , Wilcoxon, and Kruskal–Wallis tests with false discovery rate correction. Mixed-effects logistic regression was conducted to account for clustering by drug and question. Results Among 324 responses, accuracy differed significantly across models (χ 2 , p < 0.001) with Gemini 2.5 Pro achieving the highest accuracy (66.7%), followed by ChatGPT-4o (51.9%) and DeepSeek R1 (37.0%). General prompts were associated with higher accuracy than detailed prompts (59.3% vs 44.4%; p = 0.011). Gemini 2.5 Pro showed the highest explanation quality, while Gemini 2.5 Pro and ChatGPT-4o showed comparable consistency and DeepSeek R1 was less consistent. Hallucination was the most frequent error type across models. Conclusions LLMs showed variable capability in extracting regulatory and clinical information. Gemini 2.5 Pro showed the strongest overall performance, while ChatGPT-4o was faster but less accurate, and DeepSeek R1 underperformed across most domains. These findings highlight both the promise and limitations of LLMs in regulatory science and support their complementary use with human review to support evidence extraction and synthesis for regulatory review.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Evaluating large language model performance in US FDA regulatory science
Date Crossref
08/08/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Artificial Intelligence in Healthcare and EducationMachine Learning in HealthcarePharmacovigilance and Adverse Drug Reactions

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.