Accurate identification of cardiovascular events from the electronic medical record (EMR) is essential for retrospective outcomes research and quality measurement. This study compared the diagnostic accuracy of four automated EMR retrieval approaches — including a zero-shot large language model (LLM)-assisted workflow — against clinician manual chart adjudication for several cardiovascular events across multiple sites in a single US tertiary health system.
This was a retrospective diagnostic accuracy study conducted at three sites within one US tertiary health system. The analysis used two previously adjudicated adult cohorts to evaluate how well automated retrieval methods reproduce clinician-adjudicated cardiovascular outcomes.
Two distinct patient cohorts were included:
Both cohorts had prior clinician-adjudicated cardiovascular outcomes that served as the reference standard for this validation exercise.
The study used clinician manual chart adjudication as the reference standard. Outcomes assessed were:
These outcomes were compared between the automated retrieval methods and the manual adjudication results.
Four automated approaches to identify events from the EMR were evaluated:
The study assessed how each method performed relative to the clinician-adjudicated reference.
Performance measures reported included area under the receiver operating characteristic curve (AUC), sensitivity, specificity, and net reclassification improvement. AUCs and 95% confidence intervals were provided for key outcomes and used to compare methods within and between cohorts.
In Cohort 1 (n=2,258), the LLM-assisted workflow achieved the highest reported AUC for several outcomes:
For heart failure, ICD-based retrieval achieved a slightly higher AUC than the LLM:
When comparing the LLM and ICD approaches in this cohort, differences in AUC were not statistically significant across evaluated outcomes according to the reported results.
In Cohort 2 (n=1,426), the LLM-assisted workflow achieved the highest AUC for all evaluated endpoints:
In this cohort the LLM showed statistically significant higher AUCs than ICD-based retrieval for stroke and composite MACE as reported.
Across both cohorts, the LLM-assisted workflow demonstrated strong performance overall, but with variation by outcome and by cohort. In Cohort 1, although the LLM had the highest AUCs for stroke, MI, and MACE, the differences versus ICD were not statistically significant. In Cohort 2, the LLM had the highest AUCs for all outcomes and significantly outperformed ICD-based retrieval for stroke and composite MACE. ICD-based retrieval remained competitive, particularly for HF in Cohort 1.
The multisite retrospective validation indicates that a zero-shot LLM-assisted EMR extraction workflow can achieve strong diagnostic accuracy for identifying cardiovascular events, but performance is context-dependent. Results varied by specific event type and by patient cohort, and traditional ICD-based methods remained competitive for certain outcomes. The findings support a complementary role for LLM-assisted extraction in retrospective cardiovascular outcomes research rather than sole replacement of existing codified approaches.
Note: The source provided the performance metrics listed above; additional operational details of the LLM workflow, model specifications, or implementation parameters were not reported in the source article text provided.