Cardiovascular disease (CVD) continues to be a leading cause of preventable morbidity and mortality. Most widely used CVD prediction tools rely on a baseline survival framework, where predictors are represented by single measurements taken at cohort entry. These static approaches can be re-run with updated values but do not explicitly exploit the longitudinal, irregularly recorded data captured in primary-care electronic medical records (EMRs).
This protocol aims to develop and validate encounter-level prediction models that provide updated 5-year CVD risk estimates at each general practitioner (GP) encounter. The study compares sex-specific transformer-based sequence models with sex-specific dynamic landmark survival models that use time-updated covariates, and benchmarks both approaches against existing static CVD prediction tools. The objective is to evaluate whether dynamic modelling can support risk assessment and monitoring of changes over time in routine primary-care settings.
The analysis will use the NSW Lumos linked health data asset, a large privacy-preserving linkage of GP EMRs with hospital admissions, mortality and other administrative datasets across New South Wales, Australia. Eligible GP encounters included in the study begin from 2018 onwards. Analyses will be performed within a secure data environment using de-identified linked data under NSW Health governance arrangements.
The primary outcome is the first fatal or non-fatal CVD event occurring within 5 years of each eligible GP encounter. Events will be identified using ICD-10-AM codes recorded in the linked hospital and mortality datasets. The protocol specifies using these administrative codes to ascertain incident CVD events for outcome definition at the encounter level.
Predictor variables are drawn from routinely collected information available at or before each GP encounter in the linked data. These include demographics, documented chronic conditions, clinical measurements, prescribed medications and patterns of healthcare utilisation. The data are longitudinal and irregularly recorded; the modelling approaches explicitly accommodate time-updated covariates or sequential encounter histories rather than single baseline values.
One modelling strategy comprises sex-specific dynamic landmark survival models, such as Cox proportional hazards models, fit at prespecified landmark times. These models will include time-updated covariates derived from information available up to each landmark. Both full and LASSO-regularised versions of the landmark survival models will be developed to explore predictor selection and shrinkage.
Repeated resampling will be used for internal validation of the landmark survival models. The dynamic landmark approach provides a transparent, survival-analysis-based framework for encounter-level risk updating by treating each landmark as a new prediction time with covariates summarised to that point.
The protocol also describes developing sex-specific transformer-based sequence models that take encounter histories as sequential inputs and are fine-tuned to predict 5-year CVD risk at each encounter. These models are intended to leverage the longitudinal, irregular nature of EMR data to generate updated risk estimates without reducing prior information to a single baseline measurement.
For sequence models, internal validation will be implemented using patient-level and temporal separation to reduce information leakage and better simulate prospective performance. The protocol focuses on fine-tuning transformer architectures to the specific task of encounter-level 5-year risk prediction.
Internal validation strategies differ by modelling approach: repeated resampling for landmark survival models and patient-level plus temporal separation for sequence models. Geographic transportability will be examined using internal–external validation across primary health networks (PHNs). This approach assesses model performance when transported between different PHN subpopulations within the NSW Lumos asset.
Model performance will be evaluated using standard prognostic metrics: discrimination and calibration. Decision-curve analysis will assess potential clinical utility. Subgroup analyses will explore heterogeneity of performance across relevant patient groups. Both dynamic approaches will be compared to existing static predictive models to benchmark incremental gains from using longitudinal encounter-level data.
When required, missing data will be imputed. The protocol specifies comparing a range of imputation approaches to balance computational efficiency with predictive generalisability. Details of the specific imputation methods compared are not reported in this summary of the protocol and will be chosen to optimise both model performance and feasibility within the secure analysis environment.
The study is conducted under NSW Health governance arrangements with approval from the NSW Population and Health Services Research Ethics Committee (2019/ETH00660). Analyses will take place within a secure environment using de-identified linked data. Reporting will follow TRIPOD-AI guidance. Results will be disseminated through peer-reviewed publications, scientific conferences and policy and consumer forums.