The study reports the development of the Comorbidity Risk Score (CRS), a risk-prediction framework trained on linked electronic health records (EHRs) for 13 million individuals aged 40–69 from the entire population of England. CRS estimates the conditional effect of prior diagnoses on subsequent disease outcomes, providing an interpretable mapping from comorbid diagnoses to risk for a broad set of diseases, including COVID-19 hospitalisation.
CRS was trained at what the authors describe as near-saturated sample size, enabling precise estimation of the influence of many prior diagnoses on multiple outcomes simultaneously.
The CRS was built using linked EHR datasets covering the full English population subset aged 40–69, comprising 13 million individuals. The study received approvals to access data in NHS England's SDE service from the Advisory Group for Data (formerly IGARD) via Data Access Request Service (ref. DARS-NIC-381078-Y9C5K) and project approval (CCU022) from the CVD-COVID-UK/COVID-IMPACT Approvals & Oversight Board.
The data used were anonymised and made available to accredited researchers only. The authors state that ethical guidelines and reporting standards were followed, and that necessary patient/participant consent and institutional approvals were obtained. No competing interests were declared.
CRS estimated the conditional effects of 212 prior diagnosis categories on 88 disease outcomes (one outcome being COVID-19 hospitalisation, plus 87 other diseases). The model framework quantified the effect of each prior diagnosis on an outcome while conditioning on the presence of all other prior diagnoses, aiming to separate independently predictive comorbidities from associations that arise indirectly through correlated diagnoses.
This conditional-estimation approach was used to improve interpretability compared with many previous models that do not report per-diagnosis effects conditional on other diagnoses.
Across the evaluated outcomes, CRS identified approximately 5% of the population who had an elevated risk profile. This top 5% subgroup had, on average, a 3.4-fold higher risk for the outcomes assessed. Selected outcome-specific relative risks reported by the authors include:
These estimates derive from using prior diagnoses alone as predictors.
Using only prior diagnoses, CRS outperformed state-of-the-art clinical models for predicting COVID-19 outcomes referenced by the authors. The study additionally compared CRS performance with previously reported machine-learning and linear models trained on considerably smaller cohorts: the state-of-the-art AI and linear models cited were trained on about 0.5 million individuals each, whereas CRS used 13 million. The authors report that CRS substantially outperformed those models, and they infer that training sample size can outweigh model complexity for improving predictive performance.
The authors evaluated transferability of CRS across self-reported ethnic groups and report near-perfect transferability. An example provided is an AUROC ratio for Black versus White participants of 97.3%, indicating consistent discriminative performance across these groups in the study data.
A key feature of CRS is distinguishing diagnoses that remain independently predictive after conditioning on other comorbidities from diagnoses whose associations are indirect. The authors give the example of lipid metabolism disorder: after conditioning on other prior diagnoses, lipid metabolism disorder remained a strong predictor of myocardial infarction risk but was not predictive of ischaemic stroke. This demonstrates how CRS can refine understanding of which comorbidities are likely to have direct versus indirect relationships with specific outcomes.
Additionally, the study computed correlations of CRS-derived effect sizes across outcomes. For example, the correlation of effect sizes between myocardial infarction and hyperlipidaemia was 0.76, closely matching the corresponding genetic correlation of 0.79. The authors interpret such concordance as evidence that comorbidity architecture derived from diagnoses captures elements of disease aetiology.
CRS provides an interpretable, diagnosis-level resource for estimating disease risk from prior diagnoses and may support clinical decision-making during health emergencies or routine care by identifying patients at elevated risk for outcomes including COVID-19 hospitalisation and major chronic diseases.
Because CRS explicitly estimates conditional effects of many diagnoses, it can help prioritize which comorbidities to target for prevention or monitoring and can clarify when apparent associations are mediated by other diagnoses.
The authors suggest that the strong performance of CRS relative to smaller-sample complexity-focused models argues for prioritising larger training cohorts when developing risk-prediction tools.
Data used in this study are derived from NHS England's SDE service for England and were made available to accredited researchers under approvals described above. The CVD-COVID-UK/COVID-IMPACT programme obtained the necessary approvals to access these data and the study authors indicate that those wishing to access similar data should follow the application procedures of the relevant national data custodian. The manuscript declares compliance with ethical guidelines and reporting standards and reports no competing interests.
The manuscript reports the development and performance characteristics of CRS as described above. Details beyond those reported in the article—such as exact model specification, thresholds used to define the top 5% high-risk group, internal validation metrics across all 88 outcomes, or full lists of the 212 diagnoses and outcome-specific effect estimates—are not reproduced here; readers should consult the original preprint and supplementary materials for complete methodological and numerical detail where available.