This analysis harmonized individual-level baseline questionnaire data and incident invasive breast cancer (BC) diagnoses from 21 cohorts in North America, Europe, and Australia participating in the Breast Cancer Risk Prediction Project (BCRPP). The pooled dataset included 1,595,977 women aged 20–75 years who were enrolled between 1976 and 2015. Within five years of exposure assessment, 19,062 invasive BC cases (1.2% of the sample) were ascertained.
All participating cohorts obtained ethical approvals from their respective institutional review boards, and cohort-specific approvals are detailed in the source. Data access is facilitated through the BCRPP Data Platform under the Data Coordinating Center policies.
The study evaluated five established general-population breast cancer risk prediction models. For each cohort and model, the investigators estimated the five-year absolute risk of invasive BC using classical risk factors only (baseline questionnaire variables common to the cohorts).
Performance assessment focused on two complementary measures:
Model-specific and cohort-specific performance metrics were combined using meta-analysis methods. Metaregression was used to examine whether cohort characteristics (for example, age distribution, birth year, racial composition, or variable missingness) were associated with model performance.
Across cohorts and models, discrimination was modest and consistent. The pooled age-adjusted AUCs for the five models ranged narrowly from 0.57 to 0.58. This indicates similar ability across models to separate women who developed invasive breast cancer within five years from those who did not, but overall discriminatory performance remained limited when models used only classical risk factors.
No substantial differences in discrimination were observed across cohorts or by cohort-level characteristics reported in the analysis.
Calibration varied substantially between models and cohorts. Pooled expected-to-observed (E/O) ratios by model ranged from 0.83 to 1.25, reflecting both underestimation and overestimation of absolute risk depending on the model and cohort context.
A notable finding was systematic overestimation among individuals assigned to higher predicted risk categories. Specifically, overestimation was common among those with predicted five-year risk exceeding 3%. The magnitude and direction of calibration errors differed across cohorts, contributing to the observed heterogeneity in E/O ratios.
Performance metrics (AUCs and E/O ratios) were meta-analyzed across the 21 cohorts and models. Metaregression examined whether cohort characteristics explained between-cohort variation in model performance. The analysis did not identify appreciable associations between model performance and cohort age, birth year, race, or variable missingness. In other words, these cohort-level factors did not systematically account for the observed variation in discrimination or calibration.
The analysis reported improved calibration after assigning race-specific incidence rates to risk estimates. This finding indicates that tailoring the background incidence rates used to derive absolute risks can materially affect calibration, particularly for populations with different underlying incidence patterns.
Key interpretations from this pooled evaluation include:
Taken together, these observations support the rationale for developing a unified risk model designed for diverse populations that explicitly leverages appropriate incidence rates to produce better-calibrated absolute risk predictions across settings. The findings also underscore limitations of models based solely on classical risk factors for discriminating individual-level short-term risk.
All participating cohorts obtained IRB or ethics committee approvals, as detailed in the source text. The Institutional Review Board of the National Institutes of Health and other named institutional review boards provided approvals for individual cohorts where applicable. The authors declared no competing interests.
Data produced in this work are reported within the manuscript and supplementary materials. Data access for the BCRPP is available through the BCRPP Data Platform in accordance with data transfer agreements and the policies of the BCRPP Data Coordinating Center at the Division of Cancer Epidemiology and Genetics, National Cancer Institute. Funding was declared from the National Cancer Institute and intramural research funds as noted in the source.