Cancer prevalence varies geographically and reflects the interaction of multiple, measurable risk factors. Prior county-level studies using global modeling approaches can mask local variation. This study aimed to develop and compare geographically weighted machine learning and traditional models to predict cancer prevalence at the census tract level across the United States and to identify local determinants of cancer burden.
The investigators first conducted a scoping review to assemble a list of measurable drivers of cancer relevant to the US context. Using that list, they extracted variable data for 84,415 census tracts from the Centers for Disease Control and Prevention PLACES dataset and other publicly accessible sources. The assembled predictors included demographic characteristics, preventive behaviors, and indicators of metabolic health among others identified in the scoping exercise.
Several modeling strategies were implemented and compared. Global, non-spatial methods included Ordinary Least Squares (OLS), Random Forest, XGBoost, and a Deep Neural Network. Spatial approaches included Geographically Weighted Regression (GWR) and geographically weighted counterparts of the machine learning algorithms. The geographically weighted variants allow model relationships and feature effects to vary across space, capturing location-specific associations between predictors and cancer prevalence.
Models were compared using the Coefficient of Determination (reported as pseudo-R2 for some spatial models), Root Mean Square Error (RMSE), and Absolute Error. Performance comparisons focused on both overall predictive accuracy and consistency across geographic units. The analysis sought to determine whether geographically weighted machine learning delivered measurable improvements in predictive performance over global techniques and classical spatial regression.
Across the set of evaluated models, geographically weighted approaches outperformed the global counterparts. The geographically weighted XGBoost model demonstrated the strongest and most consistent overall performance, with reported pseudo-R2 values ranging between 0.89 and 0.98. These results indicate a high level of explained variation for cancer prevalence when allowing model terms to vary spatially and when using a tree-based boosting algorithm adapted for geographic heterogeneity.
Feature importance analysis from the geographically weighted XGBoost model showed that the relative influence of predictors changed by location. Variables identified as influential across various areas included older age distributions, racial composition, preventive health behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol. The spatially varying importance underscores that drivers of cancer prevalence are not uniform nationwide and that different factors can dominate in different places.
Findings from this mixed-methods, spatially explicit modeling approach highlight the utility of localized prediction at small geographic scales. Geographically weighted machine learning can detect regional risk patterns and shifting driver importance, which may inform targeted public health surveillance and more efficient allocation of preventive and diagnostic resources to hotspot areas. The study suggests that consideration of geographic heterogeneity improves both predictive accuracy and interpretability of dominant local risk factors.
The authors declared no competing interests and confirmed compliance with ethical guidelines, IRB or ethics committee approvals, and participant consent procedures as applicable. The preprint was posted on medRxiv on August 21, 2026. The authors report that all data produced in the study are available upon reasonable request.
This analysis of 84,415 US census tracts found that geographically weighted machine learning models—particularly geographically weighted XGBoost—outperformed global models and classical spatial regression in predicting cancer prevalence. The models revealed spatially varying importance of predictors such as age, racial composition, preventive behaviors, and metabolic conditions. The authors conclude that localized modeling at the census tract scale can improve understanding of regional cancer burden and help guide allocation of resources to areas of greatest need.