Hip fractures are linked to substantial morbidity, disability, increased mortality and high healthcare costs. Effective identification of persons at high risk is essential to direct preventive measures such as osteoporosis treatment and fall prevention. Existing risk tools (for example, FRAX, QFracture) typically require patient-entered data (BMI, smoking, alcohol) and therefore are less suited for automated, large-scale screening. The authors aimed to develop and internally validate a clinical decision support tool, FRACTURE-ML, capable of accurately predicting short- and long-term hip fracture risk using routinely collected national health data without the need for in-person assessment.
The study used nationwide Swedish registers. From a base population of all persons born in 1981 or earlier and alive in 2005, each individual was assigned a random baseline date between 2011 and 2013. Inclusion for analysis required age ≥50 years at baseline and no prescription for osteoporosis medication in the preceding 2 years. The final study population comprised 3,542,647 individuals. Participants were followed through the end of 2021. During follow-up, 142,327 individuals sustained a hip fracture.
A broad, unconditional variable-generation strategy produced 139,980 candidate predictors. These variables encompassed diagnostic codes, medication exposures, procedures, demographic factors and socioeconomic measures, captured across multiple historical windows and at varying levels of granularity. The approach sought to leverage the full depth of routinely collected data available in the national registers to maximize predictive signal while allowing downstream model-based feature selection or dimensionality reduction.
The dataset was partitioned into discovery (25%), development (65%) and holdout (10%) cohorts. Multiple modeling approaches were trained and compared: traditional time-to-event Cox models, gradient-boosted trees (XGBoost), and a neural-network survival model (DeepSurv). Performance was assessed using discrimination (area under the receiver-operating characteristic curve, AUC) at prespecified time horizons (for example 1, 2 and 5 years) and calibration analyses at the individual level. The authors also generated reduced models with fewer predictors to assess trade-offs between complexity and performance.
The primary FRACTURE-ML model, implemented with DeepSurv and using 2,500 predictors, achieved an AUC of 0.89 (95% CI 0.88–0.89) at 1 year and 0.88 (95% CI 0.87–0.88) at 2 years. A parsimonious DeepSurv model reduced to 35 predictors produced similar discrimination: AUC ≈0.87 at 2 years and 0.85 at 5 years. Traditional Cox models with 35 and 400 predictors reached comparable AUCs to the machine-learning approaches. Calibration plot analyses indicated excellent agreement between predicted and observed risk at the individual level for both DeepSurv and Cox models, supporting reliable absolute risk estimation in the studied population.
The authors compared FRACTURE-ML performance against a commonly advocated secondary-prevention screening approach—targeting individuals with a recent fracture as implemented by Fracture Liaison Services (FLS). The FLS criterion achieved an AUC of 0.55 (95% CI 0.54–0.55) at 2 years. For 2-year prediction, FRACTURE-ML identified nearly seven times more individuals at risk: sensitivity 0.84 (95% CI 0.82–0.85) for FRACTURE-ML versus 0.12 (95% CI 0.11–0.13) for the FLS approach. This increased sensitivity was associated with a lower but still reasonable specificity (0.79 vs 0.98 for FLS). The authors note FRACTURE-ML could thus serve as a complementary strategy for primary prevention and broader population screening beyond the FLS secondary-prevention focus.
Key strengths include the nationwide scope, large sample size (3.5 million adults ≥50), long follow-up to 2021, and extensive candidate predictor construction from linked registries. Both machine-learning and traditional statistical approaches were evaluated and shown to perform well, and reduced models retained high discrimination, enhancing potential implementability.
Important limitations reported by the authors include the lack of external validation and absence of implementation studies; these are required to confirm generalizability and clinical utility in other settings and health systems. Data underlying the analyses cannot be made publicly available because of Swedish confidentiality legislation, but the authors provide links to the code repositories used to generate results. Funding sources and competing interests are reported in the manuscript.
FRACTURE-ML demonstrated high discriminative performance (AUC up to 0.89) and strong calibration for predicting hip fracture using only routinely collected registry data. The tool identified substantially more high-risk persons than an FLS-based secondary-prevention strategy, suggesting value for population-level screening without in-person assessment. The authors conclude that FRACTURE-ML could serve as a resource-efficient approach to improve primary prevention of hip fracture, but stress that external validation and implementation research are necessary next steps to establish clinical usefulness and feasibility.