Large-scale biobank whole-genome sequencing (WGS) cohorts offer an opportunity to quantify how both rare and common germline variants contribute to cancer risk. This study leveraged WGS and harmonized phenotypes from the All of Us (AoU) Research Program to prioritize genes and rare protein-coding variants associated with acute myeloid leukemia (AML). The authors aimed to perform gene- and gene-set-level rare-variant analyses alongside single-variant tests for common variants to identify loci and aggregated burdens relevant to AML susceptibility.
Analyses used a European-like ancestry subset of the AoU Curated Data Repository (version 8). Association analyses compared Ncases = 265 AML cases against Ncontrols = 169,706 controls for rare-variant set tests, and Ncases = 265 versus Ncontrols = 169,705 for single-variant common-variant tests. Variant annotation was performed using VEP version 115 and associated plugins. The authors augmented germline data with public somatic and expression resources: clinical and somatic mutation frequency data and bulk RNA-seq from the Genomic Data Commons (GDC), single-cell RNA-seq (GSE116256), and ClinVar annotations. All code used within the AoU workspace and summary statistics are available on GitHub under the repository named by the authors, with AoU-imposed reporting limits for small counts.
Rare-variant set-based association tests were conducted using SAIGE-GENE+, a tool designed to account for case-control imbalance and relatedness in large cohorts. Single-variant association tests were performed for common variants in the same European-like subset. In addition to single-gene tests, the investigators constructed rare-variant burden risk scores across predefined gene-sets to evaluate aggregated rare coding variation for enrichment in AML cases. Multiple testing correction used a Bonferroni approach applied to the set-based Cauchy p-values.
Four genes achieved statistical significance in the rare-variant set-based tests after Bonferroni correction: DNMT3A, TET2, SRSF2, and IDH2 (Bonferroni-corrected Cauchy p-value < 0.05). These genes are reported as having a statistically significant excess burden of rare protein-coding variants in AML cases relative to controls in this AoU sample. The abstract does not report per-gene effect sizes, variant counts, or individual-variant p-values; those details are expected to be present in the full text, supplementary materials, or the accompanying summary statistics on GitHub.
To capture aggregated signals beyond individual genes, the study generated multiple rare-variant burden risk scores using different biologically informed gene-sets. Two GDC-derived gene-sets yielded a statistically significant rare-variant burden after Bonferroni correction: (1) genes observed to harbor somatic mutations in AML and (2) genes observed to harbor somatic mutations across all cancer types. These results indicate enrichment of rare germline protein-coding variation within gene-sets defined by somatic mutation occurrence in cancer, supporting a link between germline rare variants and genes implicated somatically in malignancy.
The study used publicly available resources. Individual-level AoU data were accessed under the program's data use agreement. Annotation used VEP release 115. GDC clinical and RNA-sequencing data, single-cell RNA-seq (GSE116256), and ClinVar records were additionally queried on stated dates. The authors have provided the code and corresponding summary statistics in a public GitHub repository. Per AoU publication policy, participant counts below 20 and allele counts below 40 are not reported in the public summary statistics; the AoU does not permit sharing of AoU workspaces, so reproducibility outside the AoU environment depends on the provided code and summary outputs.
These findings demonstrate that WGS in a large biobank context can identify genes and gene-sets with an elevated burden of rare protein-coding variants in AML. The statistically significant genes—DNMT3A, TET2, SRSF2, and IDH2—are notable because they are also known to be recurrently mutated in myeloid malignancies at the somatic level, which may indicate convergence between somatic and germline variation in pathways relevant to leukemogenesis.
Limitations that are not detailed in the abstract include effect size estimates, variant-level statistics, adjustment covariates, and replication in independent cohorts; these items are not reported in the abstract and should be checked in the full manuscript or supplementary files. The work is a preprint and has not undergone peer review, and the authors explicitly note that the research should not guide clinical practice until validated. The authors declare no competing interests and report that appropriate approvals and data use agreements were in place for use of AoU data.
By prioritizing genes and gene-sets through rare-variant aggregation in WGS data, the study contributes to understanding genetic architecture underlying AML risk. The overlap between genes with germline rare-variant burden and genes observed somatically in AML or across cancers supports further investigation into shared mechanisms, potential biomarkers, and functional follow-up. Full variant-level results and methods details are available in the main paper, supplementary material, and the public GitHub repository referenced by the authors.