Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, but their specialist-level accuracy and safety in hematology were not well characterized. This study benchmarked ten frontier proprietary and open-weight LLMs across two generations on a large set of board-style hematology multiple-choice questions (MCQs) to quantify performance across domains and modalities and to examine error patterns.
The evaluation set comprised 1,477 board-style MCQs derived from five publicly available educational datasets: British Society for Haematology Multiple-Choice Questions (BSH-MCQ), BSH Extended Matching Questions (BSH-EMQ), BSH Case Reports, American Society of Hematology (ASH) Hematopoiesis Case Studies, and European Society for Blood and Marrow Transplantation (EBMT) Case of the Month. Questions covered nine disease areas and six clinical skill domains and included both text-only items and multimodal case vignettes. The study used only openly available datasets; no new patient data or identifiable information were collected.
Ten frontier LLMs were benchmarked, representing both proprietary and open-weight models and spanning two generations. Named top performers reported in the source include Claude Opus 5, Gemini-3.1 Pro, Gemini-3.6 Flash, and GPT-5.6 Sol. The test set intentionally included a mix of simple recall-style items and complex clinical case vignettes to probe knowledge across subspecialist domains and clinical skills.
On text-only MCQs, the highest mean accuracy was achieved by Claude Opus 5 at 92.7%. Close behind were Gemini-3.1 Pro (91.4%), Gemini-3.6 Flash (91.0%) and GPT-5.6 Sol (89.9%). These results indicate that several frontier LLMs can reach high accuracy on board-style, text-only hematology questions drawn from established educational sources.
Performance on multimodal case vignettes was lower across the board. Reported mean accuracies for the top models were: Claude Opus 5 (76.9%), Gemini-3.1 Pro (78.7%), Gemini-3.6 Flash (74.8%) and GPT-5.6 Sol (76.7%). The drop in accuracy for multimodal items highlights that handling of multimodal clinical inputs remains more challenging than text-only question answering for these models.
Accuracy showed a significant correlation with model size for both text-only and multimodal MCQs. When comparing across model generations, open-weight models demonstrated the largest improvements in accuracy between generations, whereas proprietary models showed only marginal generational gains. The source reports these trends without presenting additional numerical breakdowns beyond the mean accuracies for named top performers.
In-depth error analysis found that top-performing models exhibited highly concordant failure patterns on the most challenging cases. This concordance suggests shared limitations across different model architectures or training data rather than isolated, model-specific weaknesses. The authors interpret this as evidence that high aggregate accuracy does not eliminate systematic blind spots that can affect multiple LLMs.
All benchmarking datasets are publicly available from the cited sources (BSH-MCQ, BSH-EMQ, BSH Case Reports, ASH Hematopoiesis Case Studies, and EBMT Case of the Month). The study used only openly accessible educational materials; no participants were recruited and no new patient-level or identifiable data were accessed. The authors state that relevant ethical guidelines were followed and that no IRB approval was required given the use of public, deidentified educational datasets.
Frontier LLMs demonstrate substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets, achieving high accuracy on many board-style MCQs. However, multimodal performance is consistently lower than text-only performance, and top models share concordant failure modes on challenging items. The authors emphasize that, despite encouraging accuracy, continuous expert-on-the-loop output monitoring is paramount when LLMs are applied in clinical or educational hematology contexts.
Several authors declared relationships with pharmaceutical companies and industry, detailed in the source. Funding for the work was declared from German Cancer Aid (grant number reported in the source). The preprint was posted on September 02, 2026, and is available with DOI: https://doi.org/10.64898/2026.09.01.26361881. No additional proprietary data or undisclosed conflicts were reported beyond the listed declarations.