Variability of outcomes of AI software for breast cancer screening classification by screening units
Rattachement africain : gb, de, cz, nl. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Purpose: Training datasets for artificial intelligence (AI) software often differ from real world use cases, whichcan affect model performance. The aim is to measure the effect of deploying AI without additional training orfine tuning across different clinical environments. Methods: Mammography images and associated metadatawere selected from the OPTIMAM database for six different screening centres covering three mammographicmanufacturers for images acquired during breast screening. The images were processed by the open-accessbreast classification convolutional neural network (CNN) model from New York University (NYU) to give ascore for malignancy. The area under the curve (AUC) was calculated for the receiver operating characteristics(ROC) for the predicted malignancy for the OPTIMAM dataset. Results: The performance of the AI tool waslower (AUC=0.744) with the OPTIMAM data than the reported AUC of 0.830 for side-wise reading in theoriginal paper. The key part of this work is how the model works when introduced into the six differentscreening centres, there was a large spread of AUC of between 0.639 and 0.750. The lowest AUC score wasassociated with a screening unit primarily using GE scanners, which was not included in the training data forthe NYU software. Testing by manufacturer and model showed GE systems (AUC=0.569 ± 0.037) which waslower than Hologic systems (AUC=0.750 ± 0.008) and Siemens systems (AUC=0.750 ± 0.012). Conclusion: TheNYU AI model was sensitive to different screening environments, in particular the manufacturer and model ofthe mammographic equipment. AI models may underperform when deployed in new clinical environments,potentially requiring additional training or fine-tuning, even for need to tune for different models ofmammography units. Overall, there is a need for independent onsite evaluation of AI software and postmarket surveillance.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.