Aller au contenu principal
Accès ouvert déclaré 2026 article

Classification Performance of General-Purpose Multimodal Large Language Models Across Orthodontic Radiographic Tasks: A Comparative Study of ChatGPT, Gemini, and Claude

0Citations signalées — pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Background and Objectives: General-purpose multimodal large language models (MLLMs) can interpret radiographic images, but their classification performance across orthodontic tasks remains uncertain. This study compared the classification performance of ChatGPT, Gemini, and Claude on lateral cephalometric, hand–wrist, and panoramic radiographs. Materials and Methods: This retrospective diagnostic accuracy study included 250 individuals, each contributing one lateral cephalometric, hand–wrist, and panoramic pretreatment radiograph (750 total). Reference classifications were established by two experienced orthodontists, with disagreements resolved by consensus. Lateral cephalometric radiographs were classified as skeletal Class I, II, or III based on the ANB angle according to Steiner analysis; hand–wrist radiographs as prepubertal, pubertal, or postpubertal; and panoramic radiographs as early mixed, late mixed, or permanent dentition. Each image was evaluated once by each AI platform using identical Turkish prompts in separate chat sessions. Classification accuracy, balanced accuracy, macro-F1, class-specific metrics, and reference agreement were assessed. Generalized estimating equations (GEE) assessed platform, radiograph type, and interaction effects on correct classification. Results: ChatGPT had the highest hand–wrist accuracy (81.6%; 95% CI, 76.3–85.9), whereas Gemini had the highest panoramic accuracy (92.8%; 95% CI, 88.9–95.4). Lateral cephalometric accuracies were 70.0%, 64.4%, and 64.8% for ChatGPT, Gemini, and Claude, respectively, with no significant interplatform difference (p = 0.336). The platform × radiograph type interaction was significant (Wald χ2 = 42.52; df = 4; p < 0.001). Agreement with the reference standard was highest for ChatGPT on hand–wrist radiographs (κw = 0.758) and Gemini on panoramic radiographs (κw = 0.883). Conclusions: Classification performance was task- and platform-dependent, with no model consistently achieving the highest performance. For the predefined classification tasks, these models should not be used as standalone tools for these classification tasks. Their potential as decision-support tools requires prospective evaluation of AI-assisted clinician performance and external validation.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Classification Performance of General-Purpose Multimodal Large Language Models Across Orthodontic Radiographic Tasks: A Comparative Study of ChatGPT, Gemini, and Claude
Date Crossref
06/09/2026
Éditeur
MDPI AG
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Dental Radiography and ImagingDental Research and COVID-19Orthodontics and Dentofacial Orthopedics

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.