---
title: "Limited benchmarks weaken conclusions comparing general-purpose and clinical AI"
id: "nature-2-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose-versus"
canonical_url: "https://medichelpline.com/clinical-feed/nature-2-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose-versus"
content_type: "clinical_feed_article"
specialty: "General"
source_name: "Nature Medicine"
source_url: "https://www.nature.com/articles/s41591-026-04638-6"
published_at: "2026-09-03T12:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Limited benchmarks weaken conclusions comparing general-purpose and clinical AI
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/nature-2-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose-versus
- **Specialty:** [General](https://medichelpline.com/clinical-feed/general.md)
- **Primary Source:** Nature Medicine
- **Source URL:** [Original Journal Publication](https://www.nature.com/articles/s41591-026-04638-6)
- **Published At:** 2026-09-03T12:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- Vishwanath et al. reported that **general-purpose large language models** outperform two specialized **clinical AI** tools across three evaluations and extended implications to procurement, reimbursement and regulation. The Matters Arising authors argue these conclusions are constrained by the benchmarks used. - The two public benchmarks that produced the study’s largest effects are acknowledged by the original authors to be confounded; the Matters Arising note that these confounders limit how much can be concluded about relative model performance. - The authors present an example (Fig. 1) showing recoverable MedQA benchmark content from a frontier LLM, illustrating risk of benchmark leakage or recoverable test content influencing results. - The blinded **real-clinical-queries (RCQ)** benchmark is recognized as a potentially valuable contribution, but the Matters Arising contend the RCQ evaluation as reported is narrow in scope and presents precision with overstated certainty. - The Matters Arising raise concerns that the real-world evaluation relies on inference-generation settings that are not sufficiently described or controlled, including nondeterminism and decoding settings that affect reproducibility. - The article calls for independent evaluation of clinical AI and more rigorous, transparent benchmark design and reporting so that claims about superiority across model classes can be assessed reliably. - The authors reference existing literature on benchmark contamination, evaluator bias, nondeterminism in hosted LLM inference, and decoding strategies to support their critique and recommend better evaluation practices. - Competing interests for the Matters Arising authors are declared but stated as unrelated to this study. The article is a Matters Arising critique of the Original Article by Vishwanath et al.
## Clinical Analysis & Structured Key Points
## Your privacy, your choice We use essential cookies to make sure the site can function. We also use optional cookies for advertising, personalisation of content, usage analysis, and social media, as well as to allow video information to be shared for both marketing, analytics and editorial purposes. By accepting optional cookies, you consent to the processing of your personal data - including transfers to third parties. Some third parties are outside of the European Economic Area, with varying standards of data protection. See our [privacy policy](https://www.nature.com/info/privacy) for more information on the use of your personal data. Manage preferences for further information and to change your choices. Accept all cookies Reject optional cookies [Skip to main content](https://www.nature.com/articles/s41591-026-04638-6#content) Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript. Advertisement [ ![Nature Medicine](https://media.springernature.com/full/nature-cms/uploads/product/nm/header-95e59e63930e5d6009bad2c23a42ab2d.svg) ](https://www.nature.com/nm) * [ View all journals ](https://www.nature.com/siteindex) * [ Saved research ](https://www.nature.com/saved-research) * [ Search ](javascript:;) ## Search Search articles by subject, keyword or author Show results from All journals This journal Search [ Advanced search ](https://www.nature.com/search/advanced) ### Quick links * [Explore articles by subject](https://www.nature.com/subjects) * [Find a job](https://www.nature.com/naturecareers) * [Guide to authors](https://www.nature.com/authors/index.html) * [Editorial policies](https://www.nature.com/authors/editorial_policies/) * [Log in](https://idp.nature.com/auth/personal/springernature?redirect_uri=https://www.nature.com/articles/s41591-026-04638-6) * [ Content Explore content ](javascript:;) ## Explore content * [ Research articles ](https://www.nature.com/nm/research-articles) * [ Reviews & Analysis ](https://www.nature.com/nm/reviews-and-analysis) * [ News & Comment ](https://www.nature.com/nm/news-and-comment) * [ Podcasts ](https://www.nature.com/nm/podcast) * [ Current issue ](https://www.nature.com/nm/current-issue) * [ Collections ](https://www.nature.com/nm/collections) * [Follow us on Facebook ](https://www.facebook.com/Nature-Medicine-193691346949/) * [Follow us on X ](https://twitter.com/naturemedicine) * [ Subscribe ](https://www.nature.com/nm/subscribe) * [Sign up for alerts ](https://journal-alerts.springernature.com/subscribe?journal_id=41591) * [ RSS feed ](https://www.nature.com/nm.rss) * [ About the journal ](javascript:;) ## About the journal * [ Aims & Scope ](https://www.nature.com/nm/aims) * [ Journal Information ](https://www.nature.com/nm/journal-information) * [ Journal Metrics ](https://www.nature.com/nm/journal-impact) * [ About the Editors ](https://www.nature.com/nm/editors) * [ Research Cross-Journal Editorial Team ](https://www.nature.com/nm/research-cross-journal-editorial-team) * [ Reviews Cross-Journal Editorial Team ](https://www.nature.com/nm/reviews-cross-journal-editorial-team) * [ Statistical Advisory Panel ](https://www.nature.com/nm/statistics-advisory-panel) * [ Our publishing models ](https://www.nature.com/nm/our-publishing-models) * [ Editorial Values Statement ](https://www.nature.com/nm/editorial-values-statement) * [ Editorial Policies ](https://www.nature.com/nm/editorial-policies) * [ Content Types ](https://www.nature.com/nm/content) * [ Web Feeds ](https://www.nature.com/nm/web-feeds) * [ Contact ](https://www.nature.com/nm/contact) * [ Publish with us ](javascript:;) ## Publish with us * [ Submission Guidelines ](https://www.nature.com/nm/submission-guidelines) * [ For Reviewers ](https://www.nature.com/nm/for-reviewers) * [ Language editing services ](https://authorservices.springernature.com/go/sn/?utm_source=For+Authors&utm_medium=Website_Nature&utm_campaign=Platform+Experimentation+2022&utm_id=PE2022) * [Open access funding](https://www.nature.com/nm/open-access-funding) * [Submit manuscript ](https://mts-nmed.nature.com/cgi-bin/main.plex) * [ Subscribe ](https://www.nature.com/nm/subscribe) * [ Sign up for alerts ](https://journal-alerts.springernature.com/subscribe?journal_id=41591) * [ RSS feed ](https://www.nature.com/nm.rss) 1. [nature](https://www.nature.com/) 2. [nature medicine](https://www.nature.com/nm) 3. [matters arising](https://www.nature.com/nm/articles?type=matters-arising) 4. article * Matters Arising * Published: 03 September 2026 # Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison * [Brett Beaulieu-Jones](https://www.nature.com/articles/s41591-026-04638-6#auth-Brett-Beaulieu_Jones-Aff1) [ORCID: orcid.org/0000-0002-6700-1468](https://orcid.org/0000-0002-6700-1468)[1](https://www.nature.com/articles/s41591-026-04638-6#Aff1) & * [Shamim Nemati](https://www.nature.com/articles/s41591-026-04638-6#auth-Shamim-Nemati-Aff2) [ORCID: orcid.org/0000-0002-0520-4948](https://orcid.org/0000-0002-0520-4948)[2](https://www.nature.com/articles/s41591-026-04638-6#Aff2) [_Nature Medicine_](https://www.nature.com/nm) (2026) [Cite this article](https://www.nature.com/articles/s41591-026-04638-6#citeas) [ Save article ](https://www.nature.com/articles/s41591-026-04638-6/save-research?_csrf=Zp_S1X9CDDqTpTMp-kLcHSJ8b2duzgh7) [ View saved research ](https://www.nature.com/saved-research) [Matters Arising](https://doi.org/10.1038/s41591-026-04637-7) to this article was published on 03 September 2026 The [Original Article](https://doi.org/10.1038/s41591-026-04431-5) was published on 12 June 2026 arising from K. Vishwanath et al. _Nature Medicine_ (2026) Vishwanath et al.[1](https://www.nature.com/articles/s41591-026-04638-6#ref-CR1 "Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat. Med. 32, 2405–2409 https://doi.org/10.1038/s41591-026-04431-5 \(2026\).") report that general-purpose frontier large language models outperform two clinical AI tools across three evaluations, and extend the finding to procurement, reimbursement and regulatory oversight. Independent evaluation of clinical AI is needed, and evaluations like the blinded real-clinical-queries (RCQ) benchmark have the potential to be a valuable contribution. We argue that the benchmarks and specific design presented constrain the conclusions that can be drawn. The two public benchmarks that produced the study’s largest effects are confounded in ways the authors acknowledge. The final real-world evaluation is limited in scope, presents results in a manner which overstates precision and relies on inference generation settings that are not adequately described or controlled. This is a preview of subscription content, [access via your institution](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04638-6) ## Access options [ Access through your institution ](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04638-6) Access Nature and 54 other Nature Portfolio journals Get Nature+, our best-value online-access subscription 27,99 € / 30 days cancel any time [Learn more](https://shop.nature.com/products/plus/?region=ROW) Subscribe to this journal Receive 12 print issues and online access 251,40 € per year only 20,95 € per issue [Learn more](https://www.nature.com/nm/subscribe) Buy this article * Purchase on SpringerLink * Instant access to the full article PDF. 39,95 € Prices may be subject to local taxes which are calculated during checkout ### Additional access options: * [Log in](https://idp.nature.com/authorize/natureuser?client_id=grover&redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04638-6) * [Learn about institutional subscriptions](https://www.springernature.com/gp/librarians/licensing/license-options) * [Read our FAQs](https://support.nature.com/en/support/home) * [Contact customer support](https://www.springernature.com/gp/contact) **Fig. 1: Example of recoverable MedQA benchmark content from a frontier large language model.** ![](https://media.springernature.com/m312/springer-static/image/art%3A10.1038%2Fs41591-026-04638-6/MediaObjects/41591_2026_4638_Fig1_HTML.png) ### Explore related subjects Discover the latest articles and news in related subjects. * [Health care](https://www.nature.com/subjects/health-care) * [Translational research](https://www.nature.com/subjects/translational-research) ## References 1. Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. _Nat. Med._ **32** , 2405–2409 (2026). 2. Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. _Appl. Sci._ **11** , 6421 (2021). [Article](https://doi.org/10.3390%2Fapp11146421) [CAS](https://www.nature.com/articles/cas-redirect/1:CAS:528:DC%2BB3MXitV2ru7vE) [ Google Scholar](http://scholar.google.com/scholar_lookup?&title=What%20disease%20does%20this%20patient%20have%3F%20A%20large-scale%20open%20domain%20question%20answering%20dataset%20from%20medical%20exams&journal=Appl.%20Sci.&doi=10.3390%2Fapp11146421&volume=11&publication_year=2021&author=Jin%2CD) 3. Arora, R. K. et al. HealthBench: evaluating large language models towards improved human health. Preprint at (2025). 4. Panickssery, A., Bowman, S. R. & Feng, S. LLM evaluators recognize and favor their own generations. In _Proc. 38th International Conference on Neural Information Processing Systems (NIPS ’24)_ 68772–68802 (Curran Associates, 2024). 5. Li, D. et al. Preference leakage: a contamination problem in LLM-as-a-judge. In _International Conference on Learning Representations_ (2026). 6. Migration guide. _Claude Platform Docs_ 7. Lab, T. M. Defeating Nondeterminism in LLM Inference. _Thinking Machines Lab_ 8. Atil, B. et al. Non-determinism of ‘deterministic’ LLM settings in hosted environments. In _Proc. 5th Workshop on Evaluation and Comparison of NLP Systems_ 135–148 (Association for Computational Linguistics, 2024). 9. Yuan, J. et al. Understanding and mitigating numerical sources of nondeterminism in LLM inference. In _39th Conference on Neural Information Processing Systems (NeurIPS 2025)_ (2025). 10. Model guidance | OpenAI API. _OpenAI Developers_ 11. Using the Messages API. _Claude Platform Docs_ 12. Holtzman, A., Buys, J., Du, L., Forbes, M. & Choi, Y. The curious case of neural text degeneration. In _International Conference on Learning Representations_ (2020). 13. Presacan, O., Nik, A., Thambawita, V., Ionescu, B. & Riegler, M. A comparative study of decoding strategies in medical text generation. In _MultiMedia Modeling. MMM 2026_ Lecture Notes in Computer Science Vol. 16413 (eds Lokoč, J. et al.) (Springer, 2026). 14. Shi, C. et al. A thorough examination of decoding methods in the era of llms. In _Proc. 2024 Conference on Empirical Methods in Natural Language Processing_ 8601–8629 (2024). 15. Manrai, A. K. Medical AI has a measurement problem. _Nature_ **655** , 1138–1139 (2026). [Article](https://doi.org/10.1038%2Fd41586-026-02125-z) [PubMed](http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=42521737) [ Google Scholar](http://scholar.google.com/scholar_lookup?&title=Medical%20AI%20has%20a%20measurement%20problem&journal=Nature&doi=10.1038%2Fd41586-026-02125-z&volume=655&pages=1138-1139&publication_year=2026&author=Manrai%2CAK) [Download references](https://citation-needed.springer.com/v2/references/10.1038/s41591-026-04638-6?format=refman&flavour=references) ## Author information ### Authors and Affiliations 1. Department of Medicine, University of Chicago, Chicago, IL, USA Brett Beaulieu-Jones 2. Department of Biomedical Informatics, University of California, San Diego, La Jolla, CA, USA Shamim Nemati Authors 1. Brett Beaulieu-Jones [View author publications](https://www.nature.com/search?author=Brett%20Beaulieu-Jones) Search author on:[PubMed](https://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=search&term=Brett%20Beaulieu-Jones)[Google Scholar](https://scholar.google.co.uk/scholar?as_q=&num=10&btnG=Search+Scholar&as_epq=&as_oq=&as_eq=&as_occt=any&as_sauthors=%22Brett%20Beaulieu-Jones%22&as_publication=&as_ylo=&as_yhi=&as_allsubj=all&hl=en) 2. Shamim Nemati [View author publications](https://www.nature.com/search?author=Shamim%20Nemati) Search author on:[PubMed](https://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=search&term=Shamim%20Nemati)[Google Scholar](https://scholar.google.co.uk/scholar?as_q=&num=10&btnG=Search+Scholar&as_epq=&as_oq=&as_eq=&as_occt=any&as_sauthors=%22Shamim%20Nemati%22&as_publication=&as_ylo=&as_yhi=&as_allsubj=all&hl=en) ### Contributions B.B.-J. and S.N. conceived the study. B.B.-J. drafted the initial manuscript. B.B.-J. and S.N. critically revised the manuscript and approved the final version. ### Corresponding author Correspondence to Brett Beaulieu-Jones. ## Ethics declarations ### Competing interests B.B.-J. serves as an advisor and consultant to Vytalize Health and is a board member of the nonprofit Association of Health Learning and Inference (AHLI). S.N. serves as an advisor and consultant to Powell & Mansfield (doing business as Neural Point) and is a founder of Clairyon. These relationships are unrelated to this study. ## Additional information **Publisher’s note** Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. ## Rights and permissions [Reprints and permissions](https://s100.copyright.com/AppDispatchServlet?title=Limited%20benchmarks%20constrain%20the%20conclusions%20of%20a%20general-purpose%20versus%20clinical%20AI%20comparison&author=Brett%20Beaulieu-Jones%20et%20al&contentID=10.1038%2Fs41591-026-04638-6&copyright=The%20Author%28s%29%2C%20under%20exclusive%20licence%20to%20Springer%20Nature%20America%2C%20Inc.&publication=1078-8956&publicationDate=2026-09-03&publisherName=SpringerNature&orderBeanReset=true) ## About this article [![Check for updates. Verify currency and authenticity via CrossMark](https://www.nature.com/articles/s41591-026-04638-6)](https://crossmark.crossref.org/dialog/?doi=10.1038/s41591-026-04638-6) ### Cite this article Beaulieu-Jones, B., Nemati, S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. _Nat Med_ (2026). https://doi.org/10.1038/s41591-026-04638-6 [Download citation](https://citation-needed.springer.com/v2/references/10.1038/s41591-026-04638-6?format=refman&flavour=citation) * Received: 16 June 2026 * Accepted: 10 August 2026 * Published: 03 September 2026 * Version of record: 03 September 2026 * DOI: https://doi.org/10.1038/s41591-026-04638-6 ### Share this article Anyone you share the following link with will be able to read this content: Get shareable link Sorry, a shareable link is not currently available for this article. Copy shareable link to clipboard Provided by the Springer Nature SharedIt content-sharing initiative [ Access through your institution ](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04638-6) [ Buy or subscribe ](https://www.nature.com/articles/s41591-026-04638-6#access-options) * Sections * Figures * References * [References](https://www.nature.com/articles/s41591-026-04638-6#Bib1) * [Author information](https://www.nature.com/articles/s41591-026-04638-6#author-information) * [Ethics declarations](https://www.nature.com/articles/s41591-026-04638-6#ethics) * [Additional information](https://www.nature.com/articles/s41591-026-04638-6#additional-information) * [Rights and permissions](https://www.nature.com/articles/s41591-026-04638-6#rightslink) * [About this article](https://www.nature.com/articles/s41591-026-04638-6#article-info) Advertisement * **Fig. 1: Example of recoverable MedQA benchmark content from a frontier large language model.** 1. Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. _Nat. Med._ **32** , 2405–2409 (2026). 2. Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. _Appl. Sci._ **11** , 6421 (2021). [Article](https://doi.org/10.3390%2Fapp11146421) [CAS](https://www.nature.com/articles/cas-redirect/1:CAS:528:DC%2BB3MXitV2ru7vE) [ Google Scholar](http://scholar.google.com/scholar_lookup?&title=What%20disease%20does%20this%20patient%20have%3F%20A%20large-scale%20open%20domain%20question%20answering%20dataset%20from%20medical%20exams&journal=Appl.%20Sci.&doi=10.3390%2Fapp11146421&volume=11&publication_year=2021&author=Jin%2CD) 3. Arora, R. K. et al. HealthBench: evaluating large language models towards improved human health. Preprint at (2025). 4. Panickssery, A., Bowman, S. R. & Feng, S. LLM evaluators recognize and favor their own generations. In _Proc. 38th International Conference on Neural Information Processing Systems (NIPS ’24)_ 68772–68802 (Curran Associates, 2024). 5. Li, D. et al. Preference leakage: a contamination problem in LLM-as-a-judge. In _International Conference on Learning Representations_ (2026). 6. Migration guide. _Claude Platform Docs_ 7. Lab, T. M. Defeating Nondeterminism in LLM Inference. _Thinking Machines Lab_ 8. Atil, B. et al. Non-determinism of ‘deterministic’ LLM settings in hosted environments. In _Proc. 5th Workshop on Evaluation and Comparison of NLP Systems_ 135–148 (Association for Computational Linguistics, 2024). 9. Yuan, J. et al. Understanding and mitigating numerical sources of nondeterminism in LLM inference. In _39th Conference on Neural Information Processing Systems (NeurIPS 2025)_ (2025). 10. Model guidance | OpenAI API. _OpenAI Developers_ 11. Using the Messages API. _Claude Platform Docs_ 12. Holtzman, A., Buys, J., Du, L., Forbes, M. & Choi, Y. The curious case of neural text degeneration. In _International Conference on Learning Representations_ (2020). 13. Presacan, O., Nik, A., Thambawita, V., Ionescu, B. & Riegler, M. A
## Related Clinical Research

- [Reinforced-count simulation: calibrating decisions under over-dispersed multi-type service demand](https://medichelpline.com/clinical-feed/plos-one-3-reinforced-count-simulation-for-decision-calibration-under-over-dispersed-multi.md)
- [Psychometric validation of the Patient-Centered Communication Scale (PCCS) in Iranian clinical nur](https://medichelpline.com/clinical-feed/plos-one-7-psychometric-features-of-the-patient-centered-communication-scale-among-iranian.md)
- [Drivers of patient satisfaction in Scottish general practice: deprivation, rurality and practice s](https://medichelpline.com/clinical-feed/bmj-open-13-patient-satisfaction-with-general-practice-in-scotland-secular-trends-and.md)
- [Using Topologically Associated Domains to Prioritize Pathogenic Non-Coding Variants in Unresolved](https://medichelpline.com/clinical-feed/biorxiv-0-identifying-putative-pathogenic-non-coding-variants-in-unresolved-rare-disease.md)
- [Jehovah’s Witnesses permit blood-derived products but keep ban on whole-blood transfusions](https://medichelpline.com/clinical-feed/stat-news-0-jehovah-s-witnesses-allow-blood-derived-products-but-keep-ban-on-whole-blood.md)

## Navigation
- [← Back to General Feed](https://medichelpline.com/clinical-feed/general.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.