Hematologic diagnostics, and specifically cytomorphologic assessment of bone marrow smears, are time-intensive and require specialized expertise. The authors evaluated whether contemporary Vision Language Models (VLMs) can perform zero-shot detection of Acute Myeloid Leukemia (AML) on digitized bone marrow smears (BMS), and whether such models might be safe and accurate enough to support clinical decision-making in hematology.
Whole-slide images were obtained from bone marrow smears of 50 patients with AML and 50 bone marrow donors. From each sample, ten representative fields of view were manually extracted for analysis. These images and their corresponding human expert reports formed the reference standard for evaluating model outputs.
Three VLMs were tested: two generalist models (Qwen3.5-397B-A17B-FP8 and GLM-4.6V-FP8) and one medically adapted model (Medgemma-27b-it). All models were used in a zero-shot setting with two distinct prompting approaches:
This design assessed both how domain-aware prompting influenced predictions and whether a medically adapted model would outperform generalist models on a hematology task.
With the context-rich prompt all models showed an overwhelming tendency to label samples as leukemic, producing poor specificity. Reported key outcomes include:
These results show that detailed, domain-focused prompting did not produce clinically acceptable discrimination between leukemic and healthy bone marrow across the tested models.
When prompted without hematologic context, model accuracies improved compared with context-rich prompting. Aggregate performance with the minimal prompt was reported as follows:
Despite these relative gains with context-free prompts, none of the models achieved reliably accurate distinction between AML and healthy bone marrow across both sensitivity and specificity metrics.
Model outputs were compared against human expert reports at the level of morphologic features. Agreement was poor across all models, indicating that the VLMs failed to consistently recognize or correctly report cell-level morphologies required for accurate cytomorphologic diagnosis.
The authors suggest that the failure of these VLMs is likely driven by limited representation of hematologic images in the models’ training data. While pathology and histopathology image archives are more widely scraped and represented, digitized hematology slides and bone marrow smears are less available in public training corpora. Consequently, hematology represents an out-of-distribution use case for the evaluated models, limiting their ability to generalize to BMS AML detection.
Given the systematic tendency to overcall leukemia and the poor morphology-level agreement, the tested VLMs are currently unsuitable for clinical decision support in hematology. Misclassification of healthy donors as leukemic and failure to detect a meaningful fraction of AML cases under certain prompting conditions raise safety concerns if these models were used without extensive further development and validation.
The study was conducted within the SAL bioregistry and approved by the Institutional Review Board of the Technical University Dresden (EK 98032010). Written informed consent was obtained from all patients and donors. The authors report that all necessary ethical approvals and participant consent procedures were followed. Data produced in the study are available upon reasonable request to the authors. The report does not provide additional methodological granularities such as per-field confusion matrices or further breakdowns beyond the summary performance metrics described.
In this evaluation of three contemporary VLMs on digitized bone marrow smears, both generalist and the tested medically adapted model failed to provide reliable, morphology-aware detection of Acute Myeloid Leukemia. Models tended to overcall AML, showed poor agreement with expert morphological descriptors, and lacked robust specificity in many configurations. The authors conclude that, as configured and trained, these VLMs are unsuitable for clinical decision support in hematology and that improved representation of hematologic imaging in training data and targeted model adaptation would be required before clinical deployment.