Aller au contenu principal
Accès ouvert déclaré 2026 article

Evaluating the Performance of Large Language Models GPT-4, Claude 3 Sonnet, and Gemini Pro in Recurrent Pregnancy Loss

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Background:The utility of large language models (LLMs) in recurrent pregnancy loss (RPL) consultation and patient education has not yet been systematically investigated. This study evaluated the performance of 3 LLMs (GPT-4, Claude 3 Sonnet, and Gemini Pro) in the field of RPL by assessing accuracy, comprehensiveness, and readability.Methods:Two experienced obstetricians and gynecologists developed medical questions based on the 2022 guidelines of the European Society of Human Reproduction and Embryology (ESHRE). The questionnaire included multiple formats, including choice questions (single-answer and multiple-answer) and short-answer questions. Short-answer questions were further categorized as common questions or clinical cases based on the question type, and as prevention, diagnosis, or treatment based on the content. Subsequently, the LLMs-generated answers were graded for accuracy, comprehensiveness, and readability. Choice questions were evaluated for accuracy only, whereas short-answer questions were evaluated for accuracy, comprehensiveness, and readability. Accuracy and comprehensiveness were evaluated using a 5-point Likert scale. Readability was evaluated using the Flesch Reading Ease (FRE) score and the Flesch–Kincaid Grade Level (FKGL).Results:Responses to 47 questions generated by LLMs showed that the best-performing model, Claude 3 Sonnet, achieved higher scores in short-answer questions for both accuracy (median score 5.00 [interquartile range (IQR), 4.00–5.00]) and comprehensiveness (median score 5.00 [IQR, 4.13–5.00]). No differences were observed between LLMs in accuracy scores for all choice questions, including single-choice and multiple-choice questions (p > 0.05). Regarding readability, the FRE and FKGL scores indicated difficult readability, ranging from college-level to professional-level reading skill. For single LLM, the median accuracy scores did not differ significantly across question types. After targeted, grounded prompting based on specific categories (prevention, diagnosis, and treatment), the accuracy scores of all 3 LLMs were improved (GPT-4, median score 4.00 [IQR, 3.00–5.00] vs. 5.00 [IQR, 5.00–5.00], p < 0.001; Claude 3 Sonnet, median score 5.00 [IQR, 4.00–5.00] vs. 5.00 [IQR, 5.00–5.00], p = 0.001; Gemini Pro, median score 3.50 [IQR, 2.00–5.00] vs. 5.00 [IQR, 5.00–5.00], p < 0.001). The comprehensiveness scores for GPT-4 improved significantly after grounded prompting, whereas Claude 3 Sonnet and Gemini Pro performed worse than baseline, although these differences were not statistically significant.Conclusions:Within the field of RPL consultation, Claude 3 Sonnet outperformed GPT-4 and Gemini Pro in terms of accuracy and comprehensiveness of short-answer questions. After targeted, grounded prompting across specific categories (prevention, diagnosis, and treatment), the accuracy scores of all 3 LLMs improved. These findings suggest the potential of LLMs as an important supplementary tool for the current medical system in the field of RPL, supporting improvements in patient management.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Evaluating the Performance of Large Language Models GPT-4, Claude 3 Sonnet, and Gemini Pro in Recurrent Pregnancy Loss
Date Crossref
17/06/2026
Éditeur
IMR Press
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Reproductive System and PregnancyArtificial Intelligence in Healthcare and EducationPregnancy and Medication Impact

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.