---
title: "Authors' reply: Benchmarks, general-purpose AI and clinical AI comparison in Nature Medicine"
id: "nature-0-reply-to-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose"
canonical_url: "https://medichelpline.com/clinical-feed/nature-0-reply-to-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose"
content_type: "clinical_feed_article"
specialty: "General"
source_name: "Nature Medicine"
source_url: "https://www.nature.com/articles/s41591-026-04637-7"
published_at: "2026-09-03T12:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Authors' reply: Benchmarks, general-purpose AI and clinical AI comparison in Nature Medicine
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/nature-0-reply-to-limited-benchmarks-constrain-the-conclusions-of-a-general-purpose
- **Specialty:** [General](https://medichelpline.com/clinical-feed/general.md)
- **Primary Source:** Nature Medicine
- **Source URL:** [Original Journal Publication](https://www.nature.com/articles/s41591-026-04637-7)
- **Published At:** 2026-09-03T12:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- The authors thank Beaulieu-Jones and Nemati for their Matters Arising and endorse the call for additional research into AI technologies in medicine. - They acknowledge that **MedQA contamination** and **HealthBench evaluator affinity** could limit the generalizability of benchmarking results and note these concerns in their original work. - In their published paper the authors treated those benchmarks as **supplementary** and attempted to mitigate bias by using a **multi-model judging panel**. - The authors reaffirm their principal finding: under the study and deployment conditions reported, **general-purpose large language models** outperformed specialized **clinical AI tools** on the Real Clinical Query (RCQ) evaluation. - The reply indicates that further specific responses were provided in the full article, but the preview content available here does not include those detailed responses. - Funding sources and competing-interest disclosures are reported: grant support from IITP/MSIT (Republic of Korea) and support for one author from the National Cancer Institute and the W.M. Keck Foundation; one author reports equity and consulting relationships. - Author contributions, affiliations and publication metadata (received, accepted, published dates and DOI) are listed; the reply references the original Brief Communication and the Matters Arising it addresses.
## Clinical Analysis & Structured Key Points
## Your privacy, your choice We use essential cookies to make sure the site can function. We also use optional cookies for advertising, personalisation of content, usage analysis, and social media, as well as to allow video information to be shared for both marketing, analytics and editorial purposes. By accepting optional cookies, you consent to the processing of your personal data - including transfers to third parties. Some third parties are outside of the European Economic Area, with varying standards of data protection. See our [privacy policy](https://www.nature.com/info/privacy) for more information on the use of your personal data. Manage preferences for further information and to change your choices. Accept all cookies Reject optional cookies [Skip to main content](https://www.nature.com/articles/s41591-026-04637-7#content) Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript. Advertisement [ ![Nature Medicine](https://media.springernature.com/full/nature-cms/uploads/product/nm/header-95e59e63930e5d6009bad2c23a42ab2d.svg) ](https://www.nature.com/nm) * [ View all journals ](https://www.nature.com/siteindex) * [ Saved research ](https://www.nature.com/saved-research) * [ Search ](javascript:;) ## Search Search articles by subject, keyword or author Show results from All journals This journal Search [ Advanced search ](https://www.nature.com/search/advanced) ### Quick links * [Explore articles by subject](https://www.nature.com/subjects) * [Find a job](https://www.nature.com/naturecareers) * [Guide to authors](https://www.nature.com/authors/index.html) * [Editorial policies](https://www.nature.com/authors/editorial_policies/) * [Log in](https://idp.nature.com/auth/personal/springernature?redirect_uri=https://www.nature.com/articles/s41591-026-04637-7) * [ Content Explore content ](javascript:;) ## Explore content * [ Research articles ](https://www.nature.com/nm/research-articles) * [ Reviews & Analysis ](https://www.nature.com/nm/reviews-and-analysis) * [ News & Comment ](https://www.nature.com/nm/news-and-comment) * [ Podcasts ](https://www.nature.com/nm/podcast) * [ Current issue ](https://www.nature.com/nm/current-issue) * [ Collections ](https://www.nature.com/nm/collections) * [Follow us on Facebook ](https://www.facebook.com/Nature-Medicine-193691346949/) * [Follow us on X ](https://twitter.com/naturemedicine) * [ Subscribe ](https://www.nature.com/nm/subscribe) * [Sign up for alerts ](https://journal-alerts.springernature.com/subscribe?journal_id=41591) * [ RSS feed ](https://www.nature.com/nm.rss) * [ About the journal ](javascript:;) ## About the journal * [ Aims & Scope ](https://www.nature.com/nm/aims) * [ Journal Information ](https://www.nature.com/nm/journal-information) * [ Journal Metrics ](https://www.nature.com/nm/journal-impact) * [ About the Editors ](https://www.nature.com/nm/editors) * [ Research Cross-Journal Editorial Team ](https://www.nature.com/nm/research-cross-journal-editorial-team) * [ Reviews Cross-Journal Editorial Team ](https://www.nature.com/nm/reviews-cross-journal-editorial-team) * [ Statistical Advisory Panel ](https://www.nature.com/nm/statistics-advisory-panel) * [ Our publishing models ](https://www.nature.com/nm/our-publishing-models) * [ Editorial Values Statement ](https://www.nature.com/nm/editorial-values-statement) * [ Editorial Policies ](https://www.nature.com/nm/editorial-policies) * [ Content Types ](https://www.nature.com/nm/content) * [ Web Feeds ](https://www.nature.com/nm/web-feeds) * [ Contact ](https://www.nature.com/nm/contact) * [ Publish with us ](javascript:;) ## Publish with us * [ Submission Guidelines ](https://www.nature.com/nm/submission-guidelines) * [ For Reviewers ](https://www.nature.com/nm/for-reviewers) * [ Language editing services ](https://authorservices.springernature.com/go/sn/?utm_source=For+Authors&utm_medium=Website_Nature&utm_campaign=Platform+Experimentation+2022&utm_id=PE2022) * [Open access funding](https://www.nature.com/nm/open-access-funding) * [Submit manuscript ](https://mts-nmed.nature.com/cgi-bin/main.plex) * [ Subscribe ](https://www.nature.com/nm/subscribe) * [ Sign up for alerts ](https://journal-alerts.springernature.com/subscribe?journal_id=41591) * [ RSS feed ](https://www.nature.com/nm.rss) 1. [nature](https://www.nature.com/) 2. [nature medicine](https://www.nature.com/nm) 3. [matters arising](https://www.nature.com/nm/articles?type=matters-arising) 4. article * Matters Arising * Published: 03 September 2026 # Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison * [Krithik Vishwanath](https://www.nature.com/articles/s41591-026-04637-7#auth-Krithik-Vishwanath-Aff1)[1](https://www.nature.com/articles/s41591-026-04637-7#Aff1), * [Yindalon Aphinyanaphongs](https://www.nature.com/articles/s41591-026-04637-7#auth-Yindalon-Aphinyanaphongs-Aff2) [ORCID: orcid.org/0000-0001-8605-5392](https://orcid.org/0000-0001-8605-5392)[2](https://www.nature.com/articles/s41591-026-04637-7#Aff2) & * [Eric Karl Oermann](https://www.nature.com/articles/s41591-026-04637-7#auth-Eric_Karl-Oermann-Aff1-Aff3-Aff4-Aff5-Aff6) [ORCID: orcid.org/0000-0002-1876-5963](https://orcid.org/0000-0002-1876-5963)[1](https://www.nature.com/articles/s41591-026-04637-7#Aff1),[3](https://www.nature.com/articles/s41591-026-04637-7#Aff3),[4](https://www.nature.com/articles/s41591-026-04637-7#Aff4),[5](https://www.nature.com/articles/s41591-026-04637-7#Aff5),[6](https://www.nature.com/articles/s41591-026-04637-7#Aff6) [_Nature Medicine_](https://www.nature.com/nm) (2026) [Cite this article](https://www.nature.com/articles/s41591-026-04637-7#citeas) [ Save article ](https://www.nature.com/articles/s41591-026-04637-7/save-research?_csrf=MJSvBbjEDQtU_AQQSTQFBfKxTnIU6SH-) [ View saved research ](https://www.nature.com/saved-research) The [Original Article](https://doi.org/10.1038/s41591-026-04638-6) was published on 03 September 2026 replying to B. Beaulieu-Jones & S. Nemati. _Nature Medicine_ (2026) We thank Beaulieu-Jones and Nemati[1](https://www.nature.com/articles/s41591-026-04637-7#ref-CR1 "Beaulieu-Jones, B. & Nemati, S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nat. Med. https://doi.org/10.1038/s41591-026-04638-6 \(2026\).") for their interest in our work and for their revised Matters Arising. We strongly endorse their call for further research into these technologies. We also agree that possible MedQA contamination and HealthBench evaluator affinity may limit the generalizability of results with these benchmarks, which we accordingly treat as supplementary in our paper[2](https://www.nature.com/articles/s41591-026-04637-7#ref-CR2 "Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat. Med. 32, 2405–2409 https://doi.org/10.1038/s41591-026-04431-5 \(2026\)."), and to the best of our ability mitigate using a multi-model judging panel. We maintain our primary conclusion that, under our study and deployment conditions, general-purpose AI models outperformed clinical AI tools in the Real Clinical Query (RCQ) evaluation. A response to further specific concerns follows. This is a preview of subscription content, [access via your institution](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04637-7) ## Access options [ Access through your institution ](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04637-7) Access Nature and 54 other Nature Portfolio journals Get Nature+, our best-value online-access subscription 27,99 € / 30 days cancel any time [Learn more](https://shop.nature.com/products/plus/?region=ROW) Subscribe to this journal Receive 12 print issues and online access 251,40 € per year only 20,95 € per issue [Learn more](https://www.nature.com/nm/subscribe) Buy this article * Purchase on SpringerLink * Instant access to the full article PDF. 39,95 € Prices may be subject to local taxes which are calculated during checkout ### Additional access options: * [Log in](https://idp.nature.com/authorize/natureuser?client_id=grover&redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04637-7) * [Learn about institutional subscriptions](https://www.springernature.com/gp/librarians/licensing/license-options) * [Read our FAQs](https://support.nature.com/en/support/home) * [Contact customer support](https://www.springernature.com/gp/contact) ### Explore related subjects Discover the latest articles and news in related subjects. * [Health policy](https://www.nature.com/subjects/health-policy) * [Translational research](https://www.nature.com/subjects/translational-research) ## References 1. Beaulieu-Jones, B. & Nemati, S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. _Nat. Med_. (2026). 2. Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. _Nat. Med._ **32** , 2405–2409 (2026). [Article](https://doi.org/10.1038%2Fs41591-026-04431-5) [CAS](https://www.nature.com/articles/cas-redirect/1:CAS:528:DC%2BB28XhsVyqu73F) [PubMed](http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=42286322) [PubMed Central](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC13375586) [ Google Scholar](http://scholar.google.com/scholar_lookup?&title=General-purpose%20large%20language%20models%20outperform%20specialized%20clinical%20AI%20tools%20on%20medical%20benchmarks.&journal=Nat.%20Med.&doi=10.1038%2Fs41591-026-04431-5&volume=32&pages=2405-2409&publication_year=2026&author=Vishwanath%2CK) [Download references](https://citation-needed.springer.com/v2/references/10.1038/s41591-026-04637-7?format=refman&flavour=references) ## Funding E.K.O. is supported by the National Cancer Institute’s Early-Stage Surgeon Scientist Program (3P30CA016087-41S1) and the W.M. Keck Foundation. This work was supported by a grant from the Institute for Information & Communications Technology Planning and Evaluation (IITP) funded by the Ministry of Science and ICT (MSIT) of the Republic of Korea government (no. RS-2019-II190075 Artificial Intelligence Graduate School Program (KAIST); no. RS-2024-00509279, Global AI Frontier Lab). The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript. ## Author information ### Authors and Affiliations 1. Department of Neurological Surgery, NYU Langone Health, New York, NY, USA Krithik Vishwanath & Eric Karl Oermann 2. Department of Population Health, NYU Langone Health, New York, NY, USA Yindalon Aphinyanaphongs 3. Global AI Frontier Lab, New York University, New York, NY, USA Eric Karl Oermann 4. Department of Radiology, NYU Langone Health, New York, NY, USA Eric Karl Oermann 5. Center for Data Science, New York University, New York, NY, USA Eric Karl Oermann 6. Neuroscience Institute, NYU Langone Health, New York, NY, USA Eric Karl Oermann Authors 1. Krithik Vishwanath [View author publications](https://www.nature.com/search?author=Krithik%20Vishwanath) Search author on:[PubMed](https://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=search&term=Krithik%20Vishwanath)[Google Scholar](https://scholar.google.co.uk/scholar?as_q=&num=10&btnG=Search+Scholar&as_epq=&as_oq=&as_eq=&as_occt=any&as_sauthors=%22Krithik%20Vishwanath%22&as_publication=&as_ylo=&as_yhi=&as_allsubj=all&hl=en) 2. Yindalon Aphinyanaphongs [View author publications](https://www.nature.com/search?author=Yindalon%20Aphinyanaphongs) Search author on:[PubMed](https://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=search&term=Yindalon%20Aphinyanaphongs)[Google Scholar](https://scholar.google.co.uk/scholar?as_q=&num=10&btnG=Search+Scholar&as_epq=&as_oq=&as_eq=&as_occt=any&as_sauthors=%22Yindalon%20Aphinyanaphongs%22&as_publication=&as_ylo=&as_yhi=&as_allsubj=all&hl=en) 3. Eric Karl Oermann [View author publications](https://www.nature.com/search?author=Eric%20Karl%20Oermann) Search author on:[PubMed](https://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=search&term=Eric%20Karl%20Oermann)[Google Scholar](https://scholar.google.co.uk/scholar?as_q=&num=10&btnG=Search+Scholar&as_epq=&as_oq=&as_eq=&as_occt=any&as_sauthors=%22Eric%20Karl%20Oermann%22&as_publication=&as_ylo=&as_yhi=&as_allsubj=all&hl=en) ### Contributions K.V., Y.A. and E.K.O., jointly supervised the study, conceptualized the design and wrote the initial draft. K.V. performed the statistical analyses and experiments. All authors reviewed and approved the final paper. ### Corresponding authors Correspondence to Krithik Vishwanath or Eric Karl Oermann. ## Ethics declarations ### Competing interests E.K.O. reports equity in MarchAI and Artisight, spousal employment by Eikon Therapeutics and consulting for Sofinnova Partners, Google and Alphatec Holdings. The remaining authors declare no competing interests. ## Additional information **Publisher’s note** Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. ## Rights and permissions [Reprints and permissions](https://s100.copyright.com/AppDispatchServlet?title=Reply%20to%3A%20Limited%20benchmarks%20constrain%20the%20conclusions%20of%20a%20general-purpose%20versus%20clinical%20AI%20comparison&author=Krithik%20Vishwanath%20et%20al&contentID=10.1038%2Fs41591-026-04637-7&copyright=The%20Author%28s%29%2C%20under%20exclusive%20licence%20to%20Springer%20Nature%20America%2C%20Inc.&publication=1078-8956&publicationDate=2026-09-03&publisherName=SpringerNature&orderBeanReset=true) ## About this article [![Check for updates. Verify currency and authenticity via CrossMark](https://www.nature.com/articles/s41591-026-04637-7)](https://crossmark.crossref.org/dialog/?doi=10.1038/s41591-026-04637-7) ### Cite this article Vishwanath, K., Aphinyanaphongs, Y. & Oermann, E.K. Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. _Nat Med_ (2026). https://doi.org/10.1038/s41591-026-04637-7 [Download citation](https://citation-needed.springer.com/v2/references/10.1038/s41591-026-04637-7?format=refman&flavour=citation) * Received: 24 June 2026 * Accepted: 10 August 2026 * Published: 03 September 2026 * Version of record: 03 September 2026 * DOI: https://doi.org/10.1038/s41591-026-04637-7 ### Share this article Anyone you share the following link with will be able to read this content: Get shareable link Sorry, a shareable link is not currently available for this article. Copy shareable link to clipboard Provided by the Springer Nature SharedIt content-sharing initiative [ Access through your institution ](https://wayf.springernature.com?redirect_uri=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-026-04637-7) [ Buy or subscribe ](https://www.nature.com/articles/s41591-026-04637-7#access-options) ## Associated content ### [General-purpose large language models outperform specialized clinical AI tools on medical benchmarks](https://www.nature.com/articles/s41591-026-04431-5) * Krithik Vishwanath * Anton Alyakin * Eric Karl Oermann Nature Medicine Brief Communication Open Access 12 Jun 2026 * Sections * References * [References](https://www.nature.com/articles/s41591-026-04637-7#Bib1) * [Funding](https://www.nature.com/articles/s41591-026-04637-7#Fun) * [Author information](https://www.nature.com/articles/s41591-026-04637-7#author-information) * [Ethics declarations](https://www.nature.com/articles/s41591-026-04637-7#ethics) * [Additional information](https://www.nature.com/articles/s41591-026-04637-7#additional-information) * [Rights and permissions](https://www.nature.com/articles/s41591-026-04637-7#rightslink) * [About this article](https://www.nature.com/articles/s41591-026-04637-7#article-info) Advertisement 1. Beaulieu-Jones, B. & Nemati, S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. _Nat. Med_. (2026). 2. Vishwanath, K. et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. _Nat. Med._ **32** , 2405–2409 (2026). [Article](https://doi.org/10.1038%2Fs41591-026-04431-5) [CAS](https://www.nature.com/articles/cas-redirect/1:CAS:528:DC%2BB28XhsVyqu73F) [PubMed](http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?cmd=Retrieve&db=PubMed&dopt=Abstract&list_uids=42286322) [PubMed Central](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC13375586) [ Google Scholar](http://scholar.google.com/scholar_lookup?&title=General-purpose%20large%20language%20models%20outperform%20specialized%20clinical%20AI%20tools%20on%20medical%20benchmarks.&journal=Nat.%20Med.&doi=10.1038%2Fs41591-026-04431-5&volume=32&pages=2405-2409&publication_year=2026&author=Vishwanath%2CK) Nature Medicine (_Nat Med_) ISSN 1546-170X (online) ISSN 1078-8956 (print) ## nature.com footer links ### About Nature Portfolio * [About us](https://www.nature.com/npg_/company_info/index.html) * [Press releases](https://www.nature.com/npg_/press_room/press_releases.html) * [Press office](https://press.nature.com/) * [Contact us](https://support.nature.com/support/home) ### Discover content * [Journals A-Z](https://www.nature.com/siteindex) * [Articles by subject](https://www.nature.com/subjects) * [protocols.io](https://www.protocols.io/) * [Nature Index](https://www.natureindex.com/) ### Publishing policies * [Nature portfolio policies](https://www.nature.com/authors/editorial_policies) * [Open access](https://www.nature.com/nature-research/open-access) ### Author & Researcher services * [Reprints & permissions](https://www.nature.com/reprints) * [Research data](https://www.springernature.com/gp/authors/research-data) * [Language editing](https://authorservices.springernature.com/language-editing/) * [Scientific editing](https://authorservices.springernature.com/scientific-editing/) * [Nature Masterclasses](https://masterclasses.nature.com/) * [Research Solutions](https://solutions.springernature.com/) ### Libraries & institutions * [Librarian service & tools](https://www.springernature.com/gp/librarians/tools-services) * [Librarian portal](https://www.springernature.com/gp/librarians/manage-your-account/librarianportal) * [Open research](https://www.nature.com/openresearch/about-open-access/information-for-institutions) * [Recommend to library](https://www.springernature.com/gp/librarians/recommend-to-your-library) ### Advertising & partnerships * [Advertising](https://partnerships.nature.com/product/digital-advertising/) * [Partnerships & Services](https://partnerships.nature.com/) * [Media kits](https://partnerships.nature.com/media-kits/) * [Branded content](https://partnerships.nature.com/product/branded-content-native-advertising/) ### Professional development * [Nature Awards](https://www.nature.com/immersive/natureawards/index.html) * [Nature Careers](https://www.nature.com/naturecareers/) * [Nature Conferences](https://conferences.
## Related Clinical Research

- [Reinforced-count simulation: calibrating decisions under over-dispersed multi-type service demand](https://medichelpline.com/clinical-feed/plos-one-3-reinforced-count-simulation-for-decision-calibration-under-over-dispersed-multi.md)
- [Psychometric validation of the Patient-Centered Communication Scale (PCCS) in Iranian clinical nur](https://medichelpline.com/clinical-feed/plos-one-7-psychometric-features-of-the-patient-centered-communication-scale-among-iranian.md)
- [Drivers of patient satisfaction in Scottish general practice: deprivation, rurality and practice s](https://medichelpline.com/clinical-feed/bmj-open-13-patient-satisfaction-with-general-practice-in-scotland-secular-trends-and.md)
- [Bacterial secreted products selectively inhibit non-symbiotic fungi in stingless bee larval diet](https://medichelpline.com/clinical-feed/biorxiv-9-bacterial-secreted-products-selectively-inhibit-non-symbiotic-fungi-in-bees.md)
- [High-altitude exercise alters immune landscape: cytotoxic suppression, humoral compensation, neutr](https://medichelpline.com/clinical-feed/medrxiv-2-high-altitude-exercise-orchestrates-a-divergent-immune-landscape-cytotoxic.md)

## Navigation
- [← Back to General Feed](https://medichelpline.com/clinical-feed/general.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.