Long-horizon prediction of kidney function trajectories in people with type 2 diabetes is commonly reported as point estimates or probabilities of events, but clinicians also need to know how reliable an individual prediction is. This study performed a secondary prognostic analysis of the ACCORD trial to develop an uncertainty-aware model that predicts 48-month change in eGFR, and to test whether conformal prediction interval width provides a clinically useful, patient-level signal of prediction reliability.
The analysis used ACCORD participants with both baseline and 48-month eGFR measurements (n = 6,853). The target outcome was the annualized change in eGFR, defined as (eGFR at 48 months minus baseline eGFR) divided by four years.
For the primary baseline feature set the investigators intentionally excluded serum creatinine (because eGFR is derived from creatinine) and excluded urine biomarkers. Models were trained, calibrated, and tested using fixed data partitions. Algorithms compared included random forest, gradient boosting, penalized linear models, and XGBoost. Prediction performance metrics reported include R2 and mean absolute error (MAE).
Conformal prediction methods were applied to produce prediction intervals for individual predictions. Both split conformal intervals and locally adaptive conformal intervals were evaluated for empirical coverage and interval width. Analyses of interval width were repeated after conditioning on baseline eGFR to test whether interval width conveyed information beyond baseline kidney function.
Multiple machine learning approaches were compared on the same training, calibration, and test splits. The authors selected a best-performing primary model based on predictive metrics.
Conformal prediction was used to generate calibrated prediction intervals at the patient level. Two conformal approaches were evaluated: a split conformal method and a locally adaptive conformal method. Performance of these intervals was assessed by empirical coverage (the proportion of true outcomes lying within the interval) and mean interval width (reported in mL/min/1.73 m2).
The best-performing primary model was a random forest, which achieved R2 = 0.382 and MAE = 3.271 mL/min/1.73 m2 for predicting annualized eGFR change over 48 months. Split 90% conformal prediction intervals achieved empirical coverage of 0.917. Locally adaptive 90% conformal intervals achieved empirical coverage of 0.909 with a mean interval width of 13.759 mL/min/1.73 m2. These results indicate that conformal intervals were well calibrated to empirical coverage targets in the test data from ACCORD.
In unadjusted analyses, wider conformal intervals were associated with larger realized prediction errors and with more rapid eGFR decline. To separate the effect of baseline kidney function from interval width, the investigators assigned interval-width quintiles within baseline-eGFR strata. After this conditioning, wider intervals remained associated with greater realized prediction error: the authors report an annual adjusted increase in error of 0.151 mL/min/1.73 m2 per interval-width quintile. This suggests that interval width carries information about prediction reliability beyond baseline eGFR alone.
Beyond baseline eGFR, wider prediction intervals were associated with a set of baseline clinical characteristics: younger age, female sex, higher HbA1c, higher triglycerides, and higher systolic blood pressure. These associations indicate that standard baseline clinical variables can predict not only the expected kidney trajectory but also the model’s uncertainty for an individual, defining an interpretable “reliability phenotype.”
Using baseline clinical variables available in ACCORD and excluding serum creatinine and urine biomarkers, the authors achieved reasonable long-horizon prediction of 48-month eGFR change. Application of conformal prediction produced calibrated, patient-specific prediction intervals. Importantly, interval width behaved as an informative reliability measure rather than a random modeling artifact: wider intervals were linked to greater realized error and to clinical features associated with higher risk.
The authors propose an uncertainty-aware framing for kidney trajectory prediction in which both the predicted rate of decline and the prediction interval width are used together to identify patients with potentially rapid and uncertain decline from baseline clinical data.
The deidentified ACCORD data used in this study are available through the NHLBI BioLINCC repository, subject to repository application and approval. Analysis code and derived aggregate outputs are available from the corresponding author on reasonable request. The study received ethical approval from the Johns Hopkins Medicine Institutional Review Board (IRB00256340). The manuscript is a medRxiv preprint and has not been peer reviewed; the authors note it should not be used to guide clinical practice at this stage.