Large language models (LLMs) have shown strong performance on medical licensing and board examinations, but their use in anesthesiology remained less thoroughly evaluated. Prior work did not compare contemporary models from different developers or consistently assess items requiring interpretation of figures. This study measured performance of two current-generation LLMs on a comprehensive anesthesiology in-training examination (ITE) review question bank to evaluate accuracy, multimodal capability, and potential applicability as educational tools.
A total of 1,001 single-best-answer questions were selected from a commercially published anesthesiology ITE review text covering 11 content chapters. No items were excluded. Two LLMs—Claude Opus 5 and GPT-5.6—were each administered every item once in a fresh, stateless context.
Prompts did not include tool access, retrieval, answer keys, or explanations. All 27 figure-dependent items were administered with their published figures supplied when evaluated in a multimodal context; the paper also reports the effect of withholding figures for a subset of figure-based items. Performance was analyzed by question format, and accuracy between models was compared using McNemar's test. Per-item data, prompts and analysis scripts are available from the corresponding author on reasonable request; the source question bank itself is commercially published and not redistributed.
Claude Opus 5 correctly answered 947 of 1,001 items for an accuracy of 94.6% (95% CI, 93.0–95.8). GPT-5.6 correctly answered 946 of 1,001 items for an accuracy of 94.5% (95% CI, 92.9–95.8). The difference in accuracy between the two models was not statistically significant (McNemar p=1.00).
Accuracy by general question format was similar for both models. On standard non-figure questions, accuracy was 94.3% for Claude Opus 5 and 94.6% for GPT-5.6. Both models reached 100% accuracy on image-option items presented in the bank.
The study included 27 figure-dependent items overall and specifically reported results for 18 figure-based items that were tested with figures supplied and then retested without figures. When figures were supplied, accuracy on the 18-item subset was 94.4% for Claude Opus 5 and 83.3% for GPT-5.6. Pooling the models, withholding figures from these same 18 items reduced pooled accuracy from 88.9% to 52.8%. The decline in performance when figures were omitted was statistically significant for each model (exact McNemar p=0.03 for Claude Opus 5 and p=0.04 for GPT-5.6).
These findings indicate that multimodal capability materially improved performance on many figure-dependent questions, and that availability of visual information is critical for maintaining high accuracy on such items.
The two models provided the same answer on 957 of 1,001 items (95.6% concordance). There were 33 items both models answered incorrectly; on 32 of those 33 items (97%), both models selected the same incorrect option. This high level of concordant error highlights persistent shared failure modes across different contemporary LLMs.
The authors note that current-generation LLM performance (~95% accuracy) represents a substantial improvement over earlier models evaluated in prior work. Earlier-generation models tested by the same group (reported in this paper) demonstrated lower percentile placements on ITE normative bands: for example, previous versions of ChatGPT ranged across markedly lower percentiles depending on model generation and training level. By contrast, both Claude Opus 5 and GPT-5.6 fall within the highest band of the ITE normative framework as reported by the authors, corresponding approximately to the 99th percentile across training levels.
The results support a growing role for LLMs as accessible, on-demand educational adjuncts for anesthesiology trainees. Key practical points from the study:
Multimodal LLM capability permits interpretation of many figure-dependent exam items, improving accuracy when figures are available.
Performance declines substantially when required visual information is absent, so educational use should ensure complete source material is supplied.
High concordance in incorrect answers across models suggests the need for critical review and confirmation of LLM outputs rather than blind reliance.
The authors frame these models as promising supplemental tools for medical education, contingent on careful oversight and verification of outputs.
This report is a preprint and has not been peer reviewed; the authors caution that the findings should not be used to guide clinical practice. The study design presented items in a stateless context without tool access or retrieval; results may differ with other prompting strategies or when models have access to external tools. The source question bank is commercially published and was not redistributed; the authors state that all model responses, per-item scoring, prompts, and analysis scripts are available from the corresponding author on reasonable request. The authors declared no competing interests and stated that ethical approvals and participant consent requirements were followed as applicable.
Two current-generation LLMs, Claude Opus 5 and GPT-5.6, achieved approximately 95% accuracy on a 1,001-item anesthesiology ITE review question bank. Multimodal input of figures materially improved performance on figure-dependent items, and withholding visual information produced significant declines in accuracy. High agreement between models — including concordant errors — underscores the need for critical review of LLM outputs when these systems are used as educational aids. The findings suggest that, with complete source material and appropriate oversight, contemporary LLMs may serve as useful adjuncts in anesthesiology medical education, while noting that the manuscript is an unreviewed preprint.