Vishwanath et al. reported that general-purpose large language models outperform two specialized clinical AI tools across three evaluations and discussed implications for procurement, reimbursement and regulatory oversight. The original study used two public benchmarks and a blinded real-clinical-queries (RCQ) evaluation to support its conclusions.
The Matters Arising authors acknowledge the potential value of independent evaluations of clinical AI but contend that the specific benchmarks and study design limit the conclusions that can be drawn from Vishwanath et al.’s results.
The Matters Arising note that the two public benchmarks responsible for the study’s largest observed effects are confounded in ways the original authors themselves acknowledge. These confounders reduce confidence that the measured differences reflect intrinsic model superiority rather than artifacts of benchmark design or contamination.
The critique emphasizes that when public benchmarks contain recoverable material or are otherwise contaminated, models with broader pretraining or access to public content may appear to perform better without demonstrating genuine clinical reasoning or generalizable capability on unseen clinical tasks.
The article includes an illustrative example (Fig. 1) showing recoverable MedQA benchmark content from a frontier LLM. This example serves to demonstrate how benchmark material can be reproduced or leaked by largescale models, which confounds comparisons that assume benchmarks are held-out or novel to evaluated systems.
The Matters Arising point to this recoverability as a specific mechanism that can inflate apparent performance on public benchmarks and argue that such risks should be controlled for or explicitly tested in comparative evaluations.
The blinded real-clinical-queries (RCQ) benchmark is highlighted as a potentially meaningful approach to assess performance on real clinical questions. However, the authors argue the RCQ evaluation in the original study is limited in scope and that results are presented in a way that overstates precision.
Specifically, the Matters Arising assert that the RCQ-derived findings cannot be generalized broadly because the sample, the clinical contexts assessed, and the evaluation metrics reported were not shown to represent the diversity or complexity of real-world clinical deployment scenarios.
The Matters Arising raise concerns about the inference-generation settings used in the original real-world evaluation. They note that nondeterminism in hosted LLM environments, decoding strategies and undocumented model or API settings can materially affect outputs and measured performance.
Citing literature on nondeterminism and decoding effects, the authors argue that inference parameters and environment details must be adequately described and controlled to allow reproducible comparisons. Where settings are not reported or controlled, apparent differences between systems may reflect implementation or deployment differences rather than model capability.
The Matters Arising conclude that independent evaluation of clinical AI is needed and that benchmarks like a blinded RCQ could be a valuable contribution if designed and reported with greater rigor. The authors recommend that future comparative studies address benchmark contamination, control inference settings, describe decoding and API parameters, and avoid overstating precision from limited real-world samples.
They also reference prior work documenting evaluator bias, preference leakage, and methodological challenges in LLM evaluation to support their recommendation for more robust, transparent evaluation practices.
The Matters Arising lists author contributions: B.B.-J. and S.N. conceived the study; B.B.-J. drafted the manuscript; both authors revised and approved the final version. Competing interests are disclosed: B.B.-J. advises and consults for Vytalize Health and is a board member of the nonprofit Association of Health Learning and Inference (AHLI); S.N. advises Powell & Mansfield (Neural Point) and is a founder of Clairyon. The authors state these relationships are unrelated to this Matters Arising critique.