---
title: "Performance of Current-Generation LLMs on Anesthesiology In-Training Exams and Educational Applica"
id: "medrxiv-19-assessing-the-performance-of-artificial-intelligence-on-anesthesiology-in"
canonical_url: "https://medichelpline.com/clinical-feed/medrxiv-19-assessing-the-performance-of-artificial-intelligence-on-anesthesiology-in"
content_type: "clinical_feed_article"
specialty: "Critical Care"
source_name: "medRxiv (Clinical Preprints)"
source_url: "https://www.medrxiv.org/content/10.64898/2026.09.21.26363591v1?rss=1"
published_at: "2026-09-22T12:00:00.000Z"
evidence_level: "Verified Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Performance of Current-Generation LLMs on Anesthesiology In-Training Exams and Educational Applica
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/medrxiv-19-assessing-the-performance-of-artificial-intelligence-on-anesthesiology-in
- **Specialty:** [Critical Care](https://medichelpline.com/clinical-feed/critical-care.md)
- **Primary Source:** medRxiv (Clinical Preprints)
- **Source URL:** [Original Journal Publication](https://www.medrxiv.org/content/10.64898/2026.09.21.26363591v1?rss=1)
- **Published At:** 2026-09-22T12:00:00.000Z
- **Evidence Rating:** Verified Feed
## Executive GIST (TL;DR)
- This preprint assessed two current-generation large language models, **Claude Opus 5** and **GPT-5.6**, on 1,001 single-best-answer items drawn from an anesthesiology in-training examination (ITE) review question bank. - All questions spanning 11 content chapters were administered once in a stateless context without tool access, retrieval, answer keys, or explanatory prompts; no items were excluded. - The study included 27 figure-dependent items; figures were supplied for those items when evaluated in a multimodal context. - Claude Opus 5 answered 947/1,001 items correctly (94.6%; 95% CI, 93.0–95.8); GPT-5.6 answered 946/1,001 correctly (94.5%; 95% CI, 92.9–95.8); the difference was not statistically significant (McNemar p=1.00). - Accuracy was similar across most question formats. On standard items accuracy was ~94.3% (Claude) and 94.6% (GPT). On 18 figure-based questions with figures supplied, accuracy was 94.4% (Claude) and 83.3% (GPT). Both models achieved 100% on image-option items. - Withholding figures reduced pooled accuracy on the 18 figure-based items from 88.9% to 52.8%; model-specific declines were significant (exact McNemar p=0.03 and p=0.04 for Claude and GPT, respectively). - The two models agreed on 957/1,001 items (95.6%); of 33 items both missed, they chose the identical incorrect option on 32 (97%). - Compared to prior evaluations of earlier-generation models in anesthesiology, current-generation LLMs achieved substantially higher accuracy—approximately 95%—placing them within the highest ITE normative band (approximately 99th percentile across training levels according to the authors). - Authors conclude that multimodal capability enables interpretation of many figure-dependent questions, but performance falls when visual information is unavailable; outputs should be critically reviewed if LLMs are used as on-demand educational adjuncts. - Data (model responses, per-item scoring, prompts, analysis scripts) are available from the corresponding author on request; the source question bank is commercially published and not redistributed. - The report is a preprint and has not been peer reviewed; the authors declared no competing interests.
## Clinical Analysis & Structured Key Points
Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education | medRxiv Skip to main content Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education View ORCID Profile Andrew F Ibrahim , John F Zaki doi: https://doi.org/10.64898/2026.09.21.26363591 Andrew F Ibrahim 1 School of Medicine, Texas Tech University Health Sciences Center; Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Andrew F Ibrahim For correspondence: andrew.ibrahim{at}ttuhsc.edu John F Zaki 2 Department of Anesthesiology, Critical Care and Pain Medicine, McGovern Medical School at UTHealth Houston Find this author on Google Scholar Find this author on PubMed Search for this author on this site Abstract Info/History Metrics Supplementary material Data/Code Preview PDF Abstract INTRODUCTION Large language models (LLMs) have demonstrated substantial performance on medical licensing and board examinations, but their application to anesthesiology remains less well studied. Prior evaluations have not compared current-generation models from different developers or assessed performance on figure-dependent questions that earlier models lacked the capability to interpret. This study evaluates the performance of two current-generation LLMs on a comprehensive anesthesiology in-training examination (ITE) review question bank, including figure-dependent items. METHODS A total of 1,001 single-best-answer questions from an anesthesiology ITE review text, spanning 11 content chapters, were administered to Claude Opus 5 and GPT-5.6. No questions were excluded. Each item was presented once in a fresh, stateless context with no tool access, retrieval, answer key, or explanation provided in the prompt. Performance was analyzed by question format, and all 27 figure-dependent items were administered with their published figures supplied. Accuracy between models was compared using McNemar's test. RESULTS Claude Opus 5 answered 947/1,001 items correctly (94.6%; 95% CI, 93.0-95.8), and GPT-5.6 answered 946/1,001 correctly (94.5%; 95% CI, 92.9-95.8); the difference was not significant (McNemar p=1.00). Accuracy was similar across most question formats. On standard questions, accuracy was 94.3% and 94.6% for Claude Opus 5 and GPT-5.6, respectively; on the 18 figure-based questions with figures supplied, accuracy was 94.4% and 83.3%, respectively; and both models achieved 100% accuracy on image-option items. Withholding figures from the same 18 figure-based questions reduced pooled accuracy from 88.9% to 52.8% (exact McNemar p=0.03 and p=0.04 for Claude Opus 5 and GPT-5.6, respectively). The models agreed on 957/1,001 items (95.6%). Of the 33 items both models answered incorrectly, they selected the same incorrect option on 32 (97%). DISCUSSION Current-generation LLMs achieved approximately 95% accuracy, substantially improving on our prior results with earlier-generation models and placing both models within the highest band of the ITE normative framework, corresponding approximately to the 99th percentile across training levels. By comparison, our prior work with previous models placed ChatGPT-3.5 at the 52nd, 3rd, and 1st percentiles and ChatGPT-4.0 at the 99th, 95th, and 84th percentiles across increasing levels of training. Multimodal capability now permits successful interpretation of many figure-dependent questions, although performance declines markedly when required visual information is unavailable, and highly concordant errors remain. These findings support an increasingly promising role for LLMs as accessible, on-demand educational adjuncts for anesthesiology trainees, provided complete source material is supplied and outputs are critically reviewed. Competing Interest Statement The authors have declared no competing interest. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data Availability Data Availability Statement All model responses, per-item scoring, prompts and analysis scripts are available from the corresponding author on reasonable request. The source question bank is commercially published and is not redistributed. Copyright The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-NC-ND 4.0 International license . Back to top Previous Next Posted September 22, 2026. Download PDF Supplementary Material Data/Code Email Thank you for your interest in spreading the word about medRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education Message Subject (Your Name) has forwarded a page to you from medRxiv Message Body (Your Name) thought you would like to see this page from the medRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education Andrew F Ibrahim , John F Zaki medRxiv 2026.09.21.26363591; doi: https://doi.org/10.64898/2026.09.21.26363591 Share This Article: Copy Citation Tools Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education Andrew F Ibrahim , John F Zaki medRxiv 2026.09.21.26363591; doi: https://doi.org/10.64898/2026.09.21.26363591 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Areas All Articles Addiction Medicine (619) Allergy and Immunology (905) Anesthesia (336) Cardiovascular Medicine (4838) Dentistry and Oral Medicine (479) Dermatology (419) Emergency Medicine (651) Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1634) Epidemiology (15986) Forensic Medicine (33) Gastroenterology (1206) Genetic and Genomic Medicine (7073) Geriatric Medicine (734) Health Economics (1067) Health Informatics (5074) Health Policy (1438) Health Systems and Quality Improvement (1780) Hematology (587) HIV/AIDS (1347) Infectious Diseases (except HIV/AIDS) (16315) Intensive Care and Critical Care Medicine (1177) Medical Education (674) Medical Ethics (155) Nephrology (727) Neurology (7318) Nursing (370) Nutrition (1087) Obstetrics and Gynecology (1245) Occupational and Environmental Health (1007) Oncology (3625) Ophthalmology (1058) Orthopedics (398) Otolaryngology (458) Pain Medicine (476) Palliative Medicine (140) Pathology (709) Pediatrics (1806) Pharmacology and Therapeutics (739) Primary Care Research (769) Psychiatry and Clinical Psychology (5927) Public and Global Health (9782) Radiology and Imaging (2438) Rehabilitation Medicine and Physical Therapy (1456) Respiratory Medicine (1250) Rheumatology (644) Sexual and Reproductive Health (774) Sports Medicine (582) Surgery (783) Toxicology (107) Transplantation (306) Urology (295)
## Related Clinical Research

- [Hemoglobin-to-RDW Ratio (HRR) and ICU Outcomes in Traumatic Brain Injury Patients](https://medichelpline.com/clinical-feed/pubmed-42770772.md) (DOI: 10.1080/02699052.2026.2735358)
- [Confounder-adjusted plasma proteomics reveal robust sepsis protein signatures across perioperative](https://medichelpline.com/clinical-feed/medrxiv-16-confounder-adjusted-plasma-proteomics-identify-robust-protein-signatures-in.md)
- [Accuracy of Dexcom G7 Continuous Glucose Monitoring in Critically Ill Patients](https://medichelpline.com/clinical-feed/pubmed-42765166.md) (DOI: 10.1177/19322968261486451)
- [Impact of Medicaid Meal Deliveries Facing Budget Cuts and Policy Challenges](https://medichelpline.com/clinical-feed/kff-health-news-0-cost-saving-medicaid-meal-deliveries-threatened-by-cuts-policy-uncertainty.md)
- [First‑trimester glycolipid and inflammatory markers linked to gestational diabetes risk in rural S](https://medichelpline.com/clinical-feed/bmj-open-11-evaluation-of-the-first-trimester-maternal-glycolipid-profile-on-gestational.md)

## Navigation
- [← Back to Critical Care Feed](https://medichelpline.com/clinical-feed/critical-care.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.