The authors present LLMPopSim, a generative population simulation framework that uses a large language model (LLM) to simulate individual health behaviors and aggregate those behaviors to produce community-level estimates. The framework is designed to integrate publicly available demographic and health data to build geographically grounded synthetic individuals whose behavior can be simulated over time and evaluated against observed community prevalence estimates. The stated aim is to assess whether agentic, individually represented LLM-based agents can scale to reproduce measurable features of real-world health behavior when aggregated at local geographic units.
LLMPopSim constructs synthetic populations using aggregate, publicly available data. Demographic and socioeconomic inputs were drawn from the U.S. Census Bureau American Community Survey (ACS) 5-year estimates (2018, 2020, and 2022). Community-level health characteristics and preventive behavior estimates were taken from the Centers for Disease Control and Prevention (CDC) PLACES project (2020, 2022, and 2024 releases). The authors emphasize that only aggregate, non-identifiable human data were used and that individual-level populations analyzed in simulation were synthetically generated from these aggregate inputs.
To demonstrate the framework, the authors selected two preventive behaviors as proof-of-concept outcomes: colorectal cancer screening and mammography. These outcomes were simulated at the individual-agent level within synthetic populations and then aggregated to compare against observed prevalence estimates at the ZIP Code Tabulation Area (ZCTA) level.
The framework was developed using historical data from Hawaiʻi. Temporal and geographic generalizability were evaluated using held-out 2022 cohorts from Hawaiʻi and New York State. The evaluation focused on how well simulated aggregate outcomes matched observed ZCTA-level prevalence for the two screening behaviors across these held-out datasets and geographies. The authors compared multiple performance dimensions including absolute error, correlation with observed geographic prevalence, and preservation of between-community variation and geographic ranking.
Across the four state–outcome evaluations reported, mean absolute error (MAE) between simulated and observed community prevalence ranged from 3.5 to 15.0 percentage points. Correlations between simulated and observed ZCTA-level prevalence ranged from 0.26 to 0.69. These metrics varied by outcome and by geography: colorectal cancer screening and mammography produced different trade-offs between absolute accuracy and preservation of geographic patterns.
Performance differed across multiple dimensions of population fidelity. For colorectal cancer screening, simulated predictions better preserved geographic ranking but tended to systematically overestimate prevalence and produced compressed geographic variation compared with observed data. For mammography, simulations achieved lower absolute error on average but showed weaker geographic correlation and inconsistent preservation of between-community variability. Prediction errors were largest in communities with lower observed screening prevalence. Error patterns also varied across community characteristics, and the authors did not observe a uniform socioeconomic gradient explaining variation in simulation performance.
The authors identify several challenges for generative population simulation. Calibration of predicted prevalences, fidelity to the true distribution of behaviors across communities, and variable performance across subgroups and community types are highlighted as key limitations. The framework produced measurable signals of real-world behavior but also produced systematic biases (for example, overestimation and compression of variation for some outcomes) that would need addressing before deployment for policy analysis or intervention planning.
The study demonstrates that individually represented LLM-based synthetic agents can aggregate into population-level patterns that retain some measurable features of real-world preventive health behavior and can generalize across time and geography to a degree. LLMPopSim is presented as an empirical foundation for further development of synthetic-population approaches, with potential future application to simulate heterogeneous responses to public health interventions. The authors emphasize that calibration, improved distributional fidelity, and better subgroup performance must be priorities for future work. They also note that all source data used are publicly available and that the work is a preprint that has not been peer reviewed and therefore should not guide clinical practice.
The authors declare no competing interests. They report following relevant ethical guidelines and confirm that the study used only openly available aggregate data from ACS and CDC PLACES. No restricted-access or individually identifiable human data were used. Data availability details and source links are provided in the manuscript, and the synthetic individual-level populations were generated from those aggregate data sources as described.