Aller au contenu principal
2026 preprint

Vision Language Models Fail to Reliably Detect Acute Myeloid Leukemia in Bone Marrow Smears

0Citations signalées — pas une note de qualité
10Institutions déclarées
2Pays d’affiliation déclarés

Résumé fourni par la source

Abstract Hematologic diagnostics and especially cytomorphologic assessment are time-intensive and require high levels of expertise. Vision Language Models (VLM) show promise in medical image analysis in radiology and histopathology, while an evaluation on detecting acute myeloid leukemia (AML) is lacking. Our goal was to evaluate three Vision Language Models regarding their diagnostic accuracy and safety in clinical decision support in detecting AML from digitized bone marrow smears (BMS). Whole slide images were obtained from bone marrow smears of 50 AML patients and 50 bone marrow donors. Ten representative fields of view per sample were extracted manually. Three VLMs were used, two of which are considered generalist models (Qwen3.5-397B-A17B-FP8, GLM-4.6V-FP8), while the other one is a medically adapted model (Medgemma-27b-it). All models performed zero-shot analysis using two prompting strategies: First, a context-rich prompt requesting reporting of WHO/FAB diagnostic criteria in a structured manner, and secondly a minimal prompt without specific hematologic context. Overall diagnostic accuracy was poor for all models as they exhibited the overwhelming tendency to classify most samples as leukemic: With context-rich prompts, GLM4.6 identified 90% of leukemic samples while also labeling 92% of bone marrow donors as AML. The medical specialist model MedGemma-27b showed similar failure, misclassifying 86% of healthy donors and correctly detecting AML in only 66% of cases. Qwen3.5 performed best under detailed prompting, achieving a specificity of 0.26 and accuracy of 0.51. Accuracy of all models improved with context-free prompts (accuracies range 0.47-0.79), yet they still lacked the ability to correctly distinguish between leukemia and healthy bone marrow. Qwen3.5 was the only model to maintain meaningful specificity (0.64) and correctly identified 94% of AML, yielding an overall accuracy of 0.79. Morphologic feature-level agreement with human expert reports was poor across all models, indicating poor recognition of cell-level morphologies. This failure is likely driven by the fact that pathology imaging archives are vastly scraped during model training while hematological samples are not as widely available and therefore, hematology is an out-of-bounds use-case for these models, rendering them currently unsuitable for clinical decision support in hematology.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Vision Language Models Fail to Reliably Detect Acute Myeloid Leukemia in Bone Marrow Smears
Date Crossref
22/08/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Digital Imaging for Blood DiseasesAI in cancer detectionMultimodal Machine Learning Applications

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.