Accurate and reproducible phenotyping in large biobanks is essential for investigating biological and environmental contributors to disease. The authors designed a flexible electronic health record (EHR)-based framework, the Phenotyping Algorithm for Cases and matched Controls using EHR-based Rules (PACER), to identify disease cases and generate matched control cohorts suitable for downstream clinical, socioeconomic, and genomic analyses. This report describes PACER's application to female breast cancer (BC) phenotyping within the All of Us Research Program Curated Data Repository (CDR v8.0).
PACER uses rule-based EHR criteria implemented against the All of Us CDR. For female BC phenotyping, the analytic cohort included participants recorded as female at birth. Cases were defined by the presence of at least two breast cancer–associated diagnostic Observational Medical Outcomes Partnership (OMOP) concept IDs documented at least 30 days apart. This requirement was implemented to increase specificity of case ascertainment by reducing false positives from single diagnostic codes or provisional entries.
Clinical, socioeconomic, and genomic data available in the All of Us CDR were integrated for subsequent validation and comparative analyses. The PACER pipeline is described as customizable and reproducible, and the authors indicate it will be made available on the All of Us Workbench as a community workspace resource.
Applying PACER to All of Us CDR v8.0, the investigators identified 10,225 female BC cases. A one-to-one matched control cohort of equal size was constructed by jointly matching on the following variables: sex, age, genetic ancestry, and state-level residency. This joint matching strategy was used to align demographic and geographic factors between cases and controls and to minimize confounding in downstream analyses.
To evaluate concordance with existing approaches, the PACER-defined BC cohort was compared with a phecodeX-based breast cancer cohort; agreement between the two methods was 91.03%. In survey-based validation, among participants who responded to relevant All of Us survey items, 80.86% of PACER-classified cases self-reported a personal history of breast cancer, compared with 1.89% of controls. These orthogonal validations—algorithmic comparison and participant self-report—support the face validity of PACER case definitions in this dataset.
The authors integrated genomic data from the All of Us CDR to assess genetic enrichment consistent with breast cancer status. PACER cases showed enrichment for breast cancer–associated variants catalogued in genome-wide association study (GWAS) resources, a higher prevalence of pathogenic mutations in known risk genes, and elevated polygenic risk scores relative to matched controls. These genomic signals provide biological validation for the phenotypes produced by PACER.
The data analyzed are accessible through the All of Us Researcher Workbench under the program's data access policies and requirements. The PACER pipeline code is publicly available on GitHub at https://github.com/ChambweLab/PACER_All_of_Us and has been archived on Zenodo at https://doi.org/10.5281/zenodo.21829117. The authors state that the pipeline will be released on the All of Us Workbench as a community workspace resource for other researchers.
The study used openly available controlled All of Us data (CDR v8.0). The authors report adherence to relevant ethical guidelines and state that necessary approvals or exemptions were obtained for use of these data. Participants' consent procedures and data access restrictions associated with All of Us are noted. The authors declared no competing interests. Funding support reported included a 2024–2026 award from the Feinstein Institutes for Medical Research's Advancing Women’s Science and Medicine Awards.
Concordance across a phecodeX-based cohort, participant self-reported history, and multiple genomic analyses supports the validity of PACER-defined female breast cancer cohorts in the All of Us CDR v8.0. PACER is presented as a reproducible, adaptable EHR phenotyping pipeline that can be customized for other diseases and research questions, facilitating risk modeling and precision medicine studies in large, integrated datasets.
Limitations specific to this report (for example, performance metrics, sensitivity, specificity, or details on excluded participants) were not fully enumerated in the abstracted content provided here. Interested users should consult the full preprint and the publicly shared code and workspace resources for implementation details, parameter settings, and supplementary analyses.