---
title: "Vision Language Models underperform for detecting Acute Myeloid Leukemia in bone marrow smears"
id: "medrxiv-17-vision-language-models-fail-to-reliably-detect-acute-myeloid-leukemia-in-bone"
canonical_url: "https://medichelpline.com/clinical-feed/medrxiv-17-vision-language-models-fail-to-reliably-detect-acute-myeloid-leukemia-in-bone"
content_type: "clinical_feed_article"
specialty: "Hematology"
source_name: "medRxiv (Clinical Preprints)"
source_url: "https://www.medrxiv.org/content/10.64898/2026.08.19.26359329v1?rss=1"
published_at: "2026-08-22T12:00:00.000Z"
evidence_level: "Verified Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Vision Language Models underperform for detecting Acute Myeloid Leukemia in bone marrow smears
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/medrxiv-17-vision-language-models-fail-to-reliably-detect-acute-myeloid-leukemia-in-bone
- **Specialty:** [Hematology](https://medichelpline.com/clinical-feed/hematology.md)
- **Primary Source:** medRxiv (Clinical Preprints)
- **Source URL:** [Original Journal Publication](https://www.medrxiv.org/content/10.64898/2026.08.19.26359329v1?rss=1)
- **Published At:** 2026-08-22T12:00:00.000Z
- **Evidence Rating:** Verified Feed
## Executive GIST (TL;DR)
- The study evaluated three **Vision Language Models (VLMs)** for zero-shot detection of **Acute Myeloid Leukemia (AML)** on digitized bone marrow smears (BMS). Data comprised whole-slide images from 50 AML patients and 50 bone marrow donors, with ten representative fields of view extracted per sample. - Models tested were two generalist VLMs (Qwen3.5-397B-A17B-FP8 and GLM-4.6V-FP8) and one medically adapted model (Medgemma-27b-it). Two prompting strategies were used: a context-rich prompt asking for WHO/FAB criteria and a minimal context-free prompt. - All models tended to overcall leukemia. With context-rich prompts, GLM-4.6V labeled 90% of leukemic samples as AML but also misclassified 92% of healthy donors as AML. - The medically adapted model MedGemma-27b-it misclassified 86% of healthy donors and correctly detected AML in only 66% of cases under context-rich prompting. - Qwen3.5 performed best with a detailed prompt but still had low specificity (0.26) and overall accuracy 0.51. Under context-free prompting, accuracies improved across models (range 0.47–0.79). - Qwen3.5 maintained the highest specificity (0.64) with context-free prompting, correctly identifying 94% of AML and yielding an overall accuracy of 0.79 in that configuration. - Agreement between model outputs and human expert morphological feature reports was poor for all models, indicating inadequate recognition of cell-level morphologies. - Authors attribute failures to limited hematology image representation in training data; pathology imaging archives and histopathology are more widely scraped than hematologic slides, making hematology an out-of-distribution use case for these VLMs. - Conclusion: Current generalist and the tested medically adapted VLM are unreliable for clinical decision support in hematology and unsuitable for AML detection from BMS in their present form. Data are available on reasonable request; study ethics approvals and informed consent were obtained.
## Clinical Analysis & Structured Key Points
Hematologic diagnostics and especially cytomorphologic assessment are time-intensive and require high levels of expertise. Vision Language Models (VLM) show promise in medical image analysis in radiology and histopathology, while an evaluation on detecting acute myeloid leukemia (AML) is lacking. Our goal was to evaluate three Vision Language Models regarding their diagnostic accuracy and safety in clinical decision support in detecting AML from digitized bone marrow smears (BMS). Whole slide images were obtained from bone marrow smears of 50 AML patients and 50 bone marrow donors. Ten representative fields of view per sample were extracted manually. Three VLMs were used, two of which are considered generalist models (Qwen3.5-397B-A17B-FP8, GLM-4.6V-FP8), while the other one is a medically adapted model (Medgemma-27b-it). All models performed zero-shot analysis using two prompting strategies: First, a context-rich prompt requesting reporting of WHO/FAB diagnostic criteria in a structured manner, and secondly a minimal prompt without specific hematologic context. Overall diagnostic accuracy was poor for all models as they exhibited the overwhelming tendency to classify most samples as leukemic: With context-rich prompts, GLM4.6 identified 90% of leukemic samples while also labeling 92% of bone marrow donors as AML. The medical specialist model MedGemma-27b showed similar failure, misclassifying 86% of healthy donors and correctly detecting AML in only 66% of cases. Qwen3.5 performed best under detailed prompting, achieving a specificity of 0.26 and accuracy of 0.51. Accuracy of all models improved with context-free prompts (accuracies range 0.47-0.79), yet they still lacked the ability to correctly distinguish between leukemia and healthy bone marrow. Qwen3.5 was the only model to maintain meaningful specificity (0.64) and correctly identified 94% of AML, yielding an overall accuracy of 0.79. Morphologic feature-level agreement with human expert reports was poor across all models, indicating poor recognition of cell-level morphologies. This failure is likely driven by the fact that pathology imaging archives are vastly scraped during model training while hematological samples are not as widely available and therefore, hematology is an out-of-bounds use-case for these models, rendering them currently unsuitable for clinical decision support in hematology.
## Related Clinical Research

- [T lymphocytes and natural killer cells in myelodysplastic syndromes: review overview and access no](https://medichelpline.com/clinical-feed/frontiers-in-immunology-12-t-lymphocytes-and-natural-killer-cells-in-myelodysplastic-syndromes-function.md)
- [Molecular landscape and ELN risk stratification in acute myeloid leukemia: findings from the REFOR](https://medichelpline.com/clinical-feed/medrxiv-0-molecular-landscape-and-risk-stratification-in-acute-myeloid-leukemia-insights.md)
- [Perturb‑seq reveals co‑regulated gene programs governing hematopoietic stem and progenitor cells](https://medichelpline.com/clinical-feed/biorxiv-16-perturb-seq-identifies-co-regulated-gene-programs-shaping-hematopoietic-stem.md)
- [Sequential conditioning for hematopoietic stem cell transplant in elderly high-risk myeloid malign](https://medichelpline.com/clinical-feed/frontiers-in-immunology-16-sequential-conditioning-in-hematopoietic-stem-cell-transplantation-in-elderly.md)
- [Whole-genome sequencing identifies rare coding variants and key genes associated with acute myeloi](https://medichelpline.com/clinical-feed/medrxiv-21-prioritizing-genes-and-rare-protein-coding-variants-in-acute-myeloid-leukemia.md)

## Navigation
- [← Back to Hematology Feed](https://medichelpline.com/clinical-feed/hematology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.