L26/O-131 Can free-access general-purpose AI platforms reliably predict IVF outcomes? A benchmarking study of image-based embryo interpretation against live-birth results
Résumé fourni par la source
Abstract Study question Are free-access, general-purpose AI platforms capable of accurately discriminating live-birth outcomes when predictions are based solely on blastocyst morphology images? Summary answer When restricted to embryo-images alone, free-access AI-platforms showed high-prognostic discordance and frequent false-negative predictions, highlighting the limitations of image-only AI and the need for clinician-oversight. What is known already AI systems trained on large, curated IVF datasets can assist embryo selection, particularly when integrating embryo images with clinical, demographic, and cycle-specific variables. However, such validated tools are largely proprietary and not universally accessible. In contrast, free-access general-purpose AI platforms are increasingly used by patients—and occasionally explored informally by clinicians—to interpret embryo images and estimate IVF prognosis. While these platforms are not designed as clinical decision tools, their real-world use raises concerns regarding how image-only prognostic outputs are generated and interpreted. Unsupervised AI-derived interpretations may contribute to pessimistic bias, patient anxiety, and misrepresentation of true embryo potential. Study design, size, duration This retrospective, single-center benchmarking study evaluated 151 single embryo transfer cycles performed in 2024, including 100 cycles resulting in live birth and 51 with negative outcomes. All cycles were non-PGT, self-gamete IVF cycles in women aged 25–40 years. Inclusion required availability of standardized blastocyst images and complete outcome data. This mixed-outcome design enabled assessment of discrimination, concordance, and misclassification patterns of free-access AI platforms under blinded, image-only conditions. Participants/materials, setting, methods Transferred-blastocysts were graded by embryologists using the Gardner-system. Images of the transferred embryo were independently uploaded to four free-access-AI-platforms(ChatGPT, Grok, Gemini, and Claude), selected for their widespread public availability and common real-world use for health-related information. Platforms were fully masked to clinical-characteristics, treatment-details, and pregnancy-outcomes, and were asked to predict outcome based solely on blastocyst-morphology. Predictions were dichotomized as live-birth or no live-birth. Analyses included concordance/discordance-rates, Fisher’s exact test, ROC–AUC analysis, and Cohen’s-kappa-statistics. Main results and the role of chance Across 151 embryo transfer cycles, free-access AI platforms demonstrated limited ability to discriminate live-birth outcomes when predictions were based on embryo images alone. Prognostic discordance was high, with Grok showing the highest discordance (60.9%), followed by Claude (56.3%) and ChatGPT (54.3%). Gemini demonstrated the highest concordance (64.9%), yet still misclassified over one-third of outcomes. Misclassification was predominantly driven by false-negative predictions, reflecting a consistent pessimistic bias across platforms. While Fisher’s exact testing showed nominal associations for some models, overall discriminatory performance remained limited, with ROC–AUC values ranging from 0.52 to 0.64—substantially lower than those reported for validated IVF-trained AI systems incorporating multimodal clinical and embryological data. Agreement between AI-derived assessments and embryologist Gardner grading was poor, indicating that AI interpretations of blastocyst morphology often differed from professional embryologist evaluations. Weighted Cohen’s κ for blastocyst expansion ranged from –0.12 to 0.10, while κ values for inner cell mass and trophectoderm grading clustered near zero. Collectively, these findings indicate that general-purpose AI platforms do not replicate professional morphological reasoning, resulting in unreliable prognostic outputs when applied in an image-only context. Limitations, reasons for caution The single-center design limits generalizability, and binary outcome classification oversimplifies embryo prognosis. Image-only prediction does not reflect contemporary AI approaches integrating clinical and cycle-level data. Performance of free-access AI platforms may evolve with model updates; however, this study reflects current real-world unsupervised use. Wider implications of the findings Free-access AI platforms should not be used independently for embryo prognosis. Clinician-guided interpretation remains essential, and validated IVF-trained AI systems integrating multimodal data should be preferred. These findings emphasize the importance of AI literacy, professional guidance, and responsible integration of AI into ART practice. Trial registration number No
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- L26/O-131 Can free-access general-purpose AI platforms reliably predict IVF outcomes? A benchmarking study of image-based embryo interpretation against live-birth results
- Date Crossref
- 01/07/2026
- Éditeur
- Oxford University Press (OUP)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.