Standard polygenic risk scores (PRS) rely on additive single-nucleotide polymorphism (SNP) effects, yet genetic liability for complex diseases also includes interactions among loci and between genes and the environment. The authors propose an extended PRS (ePRS) framework that explicitly incorporates non-additive, locus-by-locus effects not captured by additive single-locus PRS or by linkage disequilibrium (LD) tagging.
The ePRS framework decomposes additional genetic risk into distinct, biologically interpretable components so that multiple interaction-driven contributions can be captured at the individual level.
Two classes of multilocus genetic interactions are modelled within ePRS. The first, cumulative multilocus burden (denoted G+G), represents summed allele counts across loci to capture an aggregated burden effect beyond single-variant additive weights. The second, statistical epistasis (denoted GxG), uses products of allele counts between locus pairs to capture pairwise non-additive interactions.
These representations target genetic variation that contributes to broad-sense heritability but is not recovered by conventional additive PRS methods.
In addition to multilocus genetic terms, the ePRS approach models gene–environment effects derived from routinely collected cardiometabolic variables in electronic health records. These gene–environment components aim to capture interaction effects between genotype and clinical or metabolic exposures that may influence disease risk.
By building separate models for different interaction-driven components, the framework can identify distinct high-risk individuals who may not be captured by additive scores alone.
The primary analyses used data from 235,000 participants in the UK Biobank. The research was conducted under UK Biobank project 86965, and ethical approval was provided by the North West Multi-centre Research Ethics Committee of the Health Research Authority (REC reference 21/NW/0157). The authors note that the manuscript is a preprint and has not been peer reviewed.
All code used in the analyses is publicly available in the authors' GitHub repository: https://github.com/keri/prsInteractive.
Across the UK Biobank cohort, the investigators derived five complementary ePRS models that each capture different non-additive or interaction-driven components of genetic risk for type 2 diabetes. These models identified largely non-overlapping sets of high-risk individuals, indicating that different interaction terms highlight different at-risk subgroups.
This complementarity suggests that a single additive PRS may miss clinically relevant individuals who are identified when multilocus burden, epistatic interactions, or gene–environment components are considered.
The authors combined the complementary ePRS components into a composite score. This composite improved case detection beyond standard clinical predictors, according to the report, and identified individuals at elevated genetic risk even when their routine clinical measures fell within clinically normal ranges.
Details on specific metrics, effect sizes, or statistical performance measures were not reported in the abstract and are contained in the full manuscript and supplementary materials.
To assess generalizability beyond type 2 diabetes, the authors applied the same modeling strategy to celiac disease. They observed similar complementarity across ePRS models for celiac disease, supporting the idea that multilocus burden and epistatic components can be informative across distinct complex diseases.
Prospective clinical validation and replication in independent cohorts are noted as necessary steps before any clinical application.
All analysis code is available at the linked GitHub repository. The authors emphasize that this work is reported as a preprint and has not undergone peer review. They state that the findings have potential for clinical use but require prospective validation.
The abstract and article note that additive single-locus PRS and LD-based approaches do not encompass all heritable risk, and that including multiple interaction-driven biological components may be key to improving individual-level genetic risk prediction. Specific limitations, detailed results, and methodological parameters are presented in the full manuscript and supplementary materials; any unreported numerical performance details in the abstract should be consulted directly in those sources.