Aller au contenu principal
Accès ouvert déclaré 2026 article

Beyond black‐box AI: Comparing ChatGPT‐4 interpretability and accuracy to CNNs in melanocytic lesions diagnosis

1Citations signalées — pas une note de qualité
11Institutions déclarées
4Pays d’affiliation déclarés

Résumé fourni par la source

BACKGROUND: Artificial intelligence (AI) algorithms have advanced and recently shown high accuracy in diagnosing skin cancer from dermoscopic images. This study compared the diagnostic performance of the large language model ChatGPT-4 with that of specialized convolutional neural network (CNN)-based models in analyzing melanocytic lesions. PATIENTS AND METHODS: A cross-sectional comparative study was conducted using 117 dermoscopic images. The performance of ChatGPT-4 was assessed under two conditions: diagnosing lesions directly without annotations and diagnosing after annotating dermoscopic features. Results were compared with CNN-based models (YPSONO and ResNet) and human expert evaluations. The confusion matrices of all the models were calculated in addition to the diagnostic accuracy, sensitivity, specificity, and interobserver agreement (Cohen's Kappa). RESULTS: ChatGPT-4 achieved 92 % sensitivity, 89 % specificity, and an accuracy of 89.7 % in direct diagnosis. When annotations were required, sensitivity and specificity dropped to 68 % and 64 %, respectively. Agreement with experts on dermoscopic patterns was minimal (Cohen's Kappa = 0.13). ChatGPT-4 outperformed CNN models in direct diagnosis but exhibited notable limitations in describing dermoscopic features. CONCLUSIONS: ChatGPT-4 demonstrated promising potential for accurate melanoma versus nevus classification without annotations, surpassing CNN-based models. However, its limited ability to describe dermoscopic features accurately highlights the need for further research and training.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Beyond black‐box AI: Comparing ChatGPT‐4 interpretability and accuracy to CNNs in melanocytic lesions diagnosis
Date Crossref
02/04/2026
Éditeur
Wiley
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Cutaneous Melanoma Detection and ManagementArtificial Intelligence in Healthcare and EducationAI in cancer detection

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.