Ezekiel J. Emanuel and Abe Baker-Butler argue that autonomous AI will likely outperform both unaided physicians and AI-aided physicians on several core cognitive medical tasks by as early as 2030. Their claim rests on a review of studies published since January 2024 comparing autonomous AI against physicians with or without AI assistance. The five cognitive tasks highlighted are: patient information gathering, differential diagnosis, selection of cost-efficient tests, prescribing guideline-concordant treatments, and chronic disease management.
The authors conclude that autonomous AI is already superior in many published comparisons and that, based on trends from medicine and other fields, once autonomous AI surpasses unaided humans it tends to also surpass human–AI hybrids. They contend the appropriate response to the evidence is to develop licensure, liability, and real-world testing pathways rather than to reject autonomous AI on principle.
To illustrate how the medical establishment can resist effective innovations, the authors recount Joseph Lister’s introduction of antiseptic technique in the 19th century. Despite strong data and demonstrations, many American surgeons were skeptical; that skepticism contributed to poor outcomes in notable cases such as President Garfield’s post‑shooting care. The authors use this analogy to warn that rejecting autonomous AI despite evidence of superiority risks denying patients better care.
The authors summarize multiple published studies where autonomous AI outperformed physicians on specific tasks:
Additionally, across 13 studies since January 2024 that directly compared autonomous AI with AI‑aided physicians, nine showed autonomous AI to be superior. The authors note one counterintuitive finding: when AI is highly capable, physician involvement can introduce more errors and harmful corrections than benefits.
The authors summarize evidence suggesting that adding physicians to a highly capable AI system can reduce overall performance. Data cited indicate that physicians sometimes introduce errors or false corrections when interacting with competent AI outputs. That phenomenon is expected to intensify as AI capability improves, leading to situations in which autonomous AI provides more accurate diagnostic or treatment recommendations than those produced by clinicians using AI as an adjunct.
The authors summarize three principal criticisms, notably advanced by AMA CEO John Whyte, and provide responses:
Licensure and liability: Critics argue current licensing and liability frameworks are inadequate for autonomous AI. The authors agree the existing structures are insufficient but contend the remedy is to create appropriate licensure and liability mechanisms rather than to forgo autonomous AI. They reference prior JAMA proposals for licensing and a liability framework currently under review.
The “art of medicine”: Critics maintain that trust, emotional awareness, compassion, patient relationships, and complex communication require humans in the loop. The authors counter with data from a 2025 review by Howcroft and colleagues showing that 13 of 15 studies reported higher empathy ratings for AI than for human health professionals. In one study patient actors reported feeling more at ease and better listened to by Google’s AMIE (ratings reported as 97% vs. 65% and 95% vs. 72%). The authors acknowledge clinicians will remain needed for tasks that require physical presence or uniquely human actions, such as holding a patient’s hand.
Simulated vs. real‑world testing: Critics point out that much medical AI literature is based on simulations and may not generalize to clinical practice. The authors argue that the paucity of real‑world autonomous AI trials is largely due to regulatory, legal, and clinician resistance to such testing. They cite an Annals of Internal Medicine study of 461 real patient visits in which AI generated treatment recommendations via structured online chat; physicians who then saw the patients and had access to the AI recommendations produced worse treatment recommendations than the AI.
The authors concede that autonomous AI cannot perform physical procedures—surgeries, deliveries, and other interventions requiring manual skills remain outside AI’s scope. They emphasize that while autonomous AI may surpass clinicians on cognitive tasks, many roles for human clinicians remain essential for procedural care and for physical presence. The authors further note that humans may still be uniquely required for certain noncognitive aspects of care even if AI demonstrates superior empathy metrics in studies.
Emanuel and Baker-Butler call for developing robust licensing and liability frameworks to permit safe, regulated deployment and testing of autonomous AI. They assert that the answer to concerns about regulation is not rejection but structured oversight. The authors view expanded, real‑world trials as necessary to validate autonomous AI outside simulations and to ensure patients receive care supported by the best available science.
The authors argue that available data indicate autonomous AI will be better than physicians working with AI at core cognitive medical tasks and that failure to test and implement proven systems could harm patients by withholding superior care. They recommend creating licensing and liability structures and conducting real-world deployments and trials rather than relying on skepticism rooted in historical resistance to innovation.
The article frames its conclusions as a call to action: build regulatory mechanisms, run real‑world tests, and evaluate autonomous AI’s role so that patients can benefit from tools that, according to the cited studies, already outperform clinicians on specific cognitive tasks.