The authors open by thanking Beaulieu-Jones and Nemati for their Matters Arising and state their endorsement of the call for further research into clinical and general-purpose AI technologies. They emphasize agreement with key concerns raised and reiterate that open inquiry and additional studies are important for validating and extending comparisons between AI systems.
The authors acknowledge two specific limitations noted by Beaulieu-Jones and Nemati: potential MedQA contamination of benchmark data and possible HealthBench evaluator affinity that could affect how broadly results generalize. They accept that these factors may constrain the conclusions that can be drawn from the benchmarks cited and explicitly treated these benchmarks as having limited generalizability in their original paper.
In the reported Brief Communication, the authors used these benchmarks but characterized them as supplementary to the principal evaluation. They state that, because of the recognized limitations, benchmark results were not the sole basis for the paper's main conclusions. The authors reference their own earlier article in which benchmark performance formed part of a broader comparative assessment of general-purpose large language models and specialized clinical AI tools.
To address evaluator and dataset concerns, the authors report that they employed a multi-model judging panel as a mitigation strategy in the study. This approach was intended to reduce single-evaluator bias and to limit the influence of any one benchmark’s idiosyncrasies on the overall comparison. Beyond stating that a multi-model panel was used, the preview content available here does not provide additional procedural details about the panel composition, adjudication process, or other experimental safeguards.
The authors maintain their principal conclusion: under the study and deployment conditions they described, general-purpose AI models outperformed specialized clinical AI tools in the Real Clinical Query (RCQ) evaluation. They emphasize that this conclusion applies to the specific experimental and deployment settings reported in their paper, consistent with their caution that benchmark limitations could affect generalizability.
The preview of the reply available in this source provides the authors’ high-level positions, acknowledgements of limitations, and restatement of their main conclusion. The text indicates that “a response to further specific concerns follows,” but the detailed responses are not included in the preview content provided here. Readers wishing to review the full, detailed reply and specific counterpoints should consult the full article via the journal or institutional access channels, as noted in the publication metadata and access options.
The authors declare funding and support sources: one author (E.K.O.) is supported by the National Cancer Institute’s Early-Stage Surgeon Scientist Program and the W.M. Keck Foundation, and the work also received grant support from the Institute for Information & Communications Technology Planning and Evaluation (IITP) funded by the Ministry of Science and ICT of the Republic of Korea. The authors state that funders did not influence study design, data collection and analysis, publication decisions, or manuscript preparation.
Competing-interest disclosures are reported: E.K.O. declares equity in MarchAI and Artisight, spousal employment by Eikon Therapeutics, and consulting relationships with Sofinnova Partners, Google and Alphatec Holdings. The remaining authors declare no competing interests.
Affiliations listed for the authors include departments at NYU Langone Health and the Global AI Frontier Lab at New York University. The authors report joint supervision, study conceptualization, drafting and statistical analysis contributions among K.V., Y.A. and E.K.O. Publication metadata available in the source shows receipt, acceptance and publication dates and the DOI for the reply. The reply cites the original Brief Communication comparing general-purpose large language models and specialized clinical AI tools and the Matters Arising it addresses.
Note: This rewritten summary and article-style presentation are based solely on the preview content available in the provided source. Specific methodological clarifications and the authors’ detailed responses to individual points raised by Beaulieu-Jones and Nemati were referenced but are not present in the source excerpt reviewed here.