This study used linked electronic health record and pharmacy dispensing data from 7,625 adults treated for hypertension. The primary objective was to determine whether the temporal structure in a patient’s antihypertensive medication history—fills, gaps, and regimen changes—contains prognostic information that is lost when treatment is reduced to binary indicators or aggregate adherence measures such as PDC (proportion of days covered).
Investigators defined a 1-year medication-history window and sought to predict a composite clinical outcome over a fixed 3-year horizon. The team kept cohort selection, clinical covariates, prediction horizon, and validation strategy constant while comparing alternative representations of dispensing data and model classes.
Three representations of medication exposure were compared:
The study also used factorial analyses to separate the impact of richer representations (data encoding) from changes in model class (statistical learning method), enabling assessment of how much of any observed performance gain came from preserved temporal information versus model complexity.
The outcome was a composite of myocardial infarction, stroke, or all-cause death occurring within three years after the 1-year medication-history window. Prediction used the same clinical covariates and validation approach across representation types to isolate the incremental value of medication-history encoding.
Model discrimination, measured by the concordance index (C-index), improved progressively with richer medication representations:
The sequence model produced an overall ΔC of 0.0283 compared with clinical factors alone (95% CI, 0.0189–0.0382) and exceeded engineered summaries by 0.0123 (95% CI, 0.0040–0.0202). Factorial decomposition attributed roughly three times more discrimination gain to richer representations (≈0.0181–0.0191) than to changing the model class (≈0.0059–0.0078).
A notable subgroup finding was that among patients with high aggregate adherence (PDC ≥0.80), PDC ranked risk poorly (C-index = 0.4628). In contrast, strata derived from the sequence model had broad 3-year event rates ranging from 3.6% to 22.8%, indicating that temporal sequencing recovered substantial risk heterogeneity hidden by aggregate metrics.
The authors interpret these results to mean that a medication history’s prognostic value mainly resides in its temporal structure—the order, timing, and changes in fills—information that is lost when dispensing records are aggregated. The performance advantage of the sequence model therefore reflects preserved information in the representation rather than merely increased model complexity.
Practically, these findings suggest that dispensing timelines tokenized into ordered sequences can reveal treatment signatures that improve cardiovascular risk stratification from existing records. This could affect risk prediction in research settings and potentially inform clinical decision support if validated and implemented appropriately.
The authors emphasize the need for external validation before broader application. Data used in the study are not publicly available because they contain protected health information; de-identified analytic datasets may be provided to qualified researchers through the corresponding author subject to institutional approval and a data-use agreement.
Ethical oversight was provided by the University of Michigan Medical School Institutional Review Board (HUM00244597) with a waiver of informed consent. One author disclosed a one-time advisory board role for Boehringer Ingelheim; the remaining authors declared no competing interests.
Overall, the study reports that tokenized antihypertensive dispensing sequences recover prognostic information lost by aggregate adherence measures and improve 3-year cardiovascular risk discrimination in this single-center dataset, with external validation required before clinical adoption.