Aller au contenu principal
Accès ouvert déclaré 2026 article

Select large language models outperform hip preservation experts on consensus‐based hip preservation questionnaire

0Citations signalées — pas une note de qualité
6Institutions déclarées
3Pays d’affiliation déclarés

Résumé fourni par la source

PURPOSE: Artificial intelligence (AI) is increasingly utilized in medical education and clinical contexts, yet few studies compare the performance of large language models (LLMs) to subspecialized experts in providing guideline-based medical information on hip preservation. The purpose of this study was to evaluate the performance of three LLMs compared to a panel of international hip preservation experts in answering guideline-based questions related to femoroacetabular impingement syndrome, hip dysplasia and microinstability of the hip. METHODS: A 21-item questionnaire was developed based on published consensus guidelines. The survey was distributed to a panel of hip preservation specialists identified through the professional network of the senior author. Ten experts responded and were included in the analysis. Three LLMs (ChatGPT 5.2, Gemini 3 and Claude 4.5 Sonnet) were prompted using the same questionnaire, with each LLM performing three runs per item. Outcomes included overall accuracy, percent agreement, Fleiss' κ, generalized linear mixed-effects modelling and qualitative assessment of AI answer justifications. RESULTS: Expert accuracy was 90.5% (95% confidence interval [CI] 87.1-93.9), compared to 100% (p = 0.004) for Gemini, 98.4% (p = 0.016) for ChatGPT and 96.8% (p = 0.053) for Claude. Expert percent agreement was 42.9% and Fleiss' κ was 0.769; alternatively, AI intra-item percent agreement was 100% (Gemini) and 95.2% (ChatGPT and Claude). ChatGPT and Claude provided thorough justifications for even incorrect responses, and Gemini demonstrated formatting deviations despite 100% accuracy. CONCLUSION: The three LLMs demonstrated high accuracy and consistency when answering the hip preservation questionnaire, with two of the LLMs statistically outperforming the expert panel. In structured, verifiable question sets, the ability of newer LLMs to accurately and consistently respond to consensus-based questions is improving compared to prior reports. LLMs are likely to serve as an adjunct in orthopaedic education and practice, and limitations to AI's implementation into practice should be continuously and rigorously explored. LEVEL OF EVIDENCE: Level V.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Select large language models outperform hip preservation experts on consensus‐based hip preservation questionnaire
Date Crossref
07/09/2026
Éditeur
Wiley
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Artificial Intelligence in Healthcare and EducationHip disorders and treatmentsClinical Reasoning and Diagnostic Skills

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.