Across medicine, clinicians are increasingly using artificial intelligence models to predict patient risk for outcomes such as sepsis, falls, and death. These prediction tools are frequently embedded directly into electronic health records, which makes incorporating their outputs into workflows straightforward. That ease of use can also mean predictions are accepted and acted upon without sufficient interrogation of their strengths and limits.
James Deardorff, a geriatrician and assistant professor in the division of geriatrics at the University of California San Francisco, argues that models must be evaluated not only for their overall accuracy but also for performance in relevant subgroups — including older adults. Older patients may differ from the populations on which models were trained in ways that affect calibration, discrimination, and clinical relevance. According to Deardorff, clinicians need to understand both technical performance metrics and the real-world implications of acting on model outputs in geriatric care.
The source notes that Deardorff has developed several prediction models aimed specifically at older adults, addressing outcomes from mortality to the need for nursing home care. His background in building geriatric-focused models informs his perspective: models intended for general adult populations do not automatically translate to appropriate or safe use among older patients. The publicly available portion of the article emphasizes his dual role as a model developer and a clinician concerned about responsible deployment.
Deardorff recently penned a commentary responding to a large analysis published in JAMA Network Open that evaluated Epic’s proprietary end-of-life prediction model. The STAT piece summarizes his central point: even when a model demonstrates statistical accuracy in predicting outcomes such as one-year mortality, that accuracy alone does not guarantee good clinical outcomes once predictions are used to guide care. The JAMA Network Open analysis is cited as the empirical prompt for Deardorff’s commentary; the STAT article links to both the analysis and his written response.
Deardorff and his co-author draw a clear distinction between lower-risk and higher-risk applications of predicted mortality. If a predicted probability of one-year mortality is used to trigger an open-ended conversation about goals of care, the potential downside is limited and clinicians can use the information as one input among many. In contrast, using the same prediction to inform high-stakes decisions — for example, to adjust transplant priority or restrict access to interventions — could produce significant, potentially harmful consequences for patients. The STAT summary underscores that the ethical and practical impact of model-driven decisions depends heavily on the use case.
The publicly available excerpt from STAT emphasizes practical cautions for clinicians: assess algorithmic performance in relevant patient subgroups, understand how predictions were validated, and consider the consequences of different uses of model outputs. Because predictions can be easily integrated into electronic records and clinical workflows, clinicians and health systems should avoid treating algorithms as infallible and should design safeguards around high-stakes applications.
This STAT article was written by Katie Palmer and published Sept. 18, 2026. It summarizes Deardorff’s views and links to the JAMA Network Open analysis and his commentary, but the piece is presented as a STAT+ exclusive. The publicly available portion provides the core argument and context; however, additional details, examples, and further reporting referenced by STAT are behind the STAT+ paywall and were not included in the source material available for this rewrite. Therefore, any specifics beyond what is cited above were not reported in the source and are not represented here.
References and source notes