The authors present an omnigenic neural network, a genome-scale neural architecture designed to reflect biological process hierarchies and the omnigenic model of complex traits. The architecture is intended to bridge gaps between conventional additive polygenic methods and domain-specific neural networks by providing a biologically informed structure that remains interpretable at systems level.
Genetic prediction of complex phenotypes frequently relies on additive linear models. Those models scale well to large variant sets but cannot represent non-additive effects, complex interactions, or integrate heterogeneous molecular and clinical data in a unified framework. Generic neural network architectures have advanced other domains but are challenging to apply at genome scale because genotypes are sparse and high-dimensional, effective sample sizes are constrained, and off-the-shelf architectures lack interpretability.
The omnigenic network is biologically structured to learn hierarchical representations of biological processes. It is explicitly designed to accommodate multimodal inputs, to enable transfer learning, and to support multitask prediction across multiple related phenotypes. These design choices aim to (1) capture nonlinear effects and interactions among genetic variants, (2) integrate other molecular or clinical modalities when available, and (3) leverage shared information across outcomes to improve predictive performance.
Models were trained using data from the UK Biobank and evaluated for external validity in the All of Us cohort. The authors report training and assessment procedures that test generalization across cohorts; further dataset-level details, including sample sizes and specific preprocessing steps, are provided in the original manuscript and supplementary materials (aggregate data and code are available on GitHub, and individual-level records are available via the respective biobank platforms).
The omnigenic models were developed and evaluated for multiple complex phenotypes. For ischemic heart disease, type 2 diabetes, and schizophrenia, models trained in UK Biobank and evaluated in All of Us outperformed published polygenic risk scores from the PGS Catalog and the PRS-CSx method. The authors attribute this improvement to the architecture’s ability to represent nonlinear effects, integrate hierarchical biological signals, and exploit multimodal inputs where available.
Beyond single-phenotype models, the team trained a multitask omnigenic model across 36 cardiovascular endpoints. The multitask model exceeded the performance of corresponding single-phenotype models and baseline approaches, demonstrating the potential benefit of shared representation learning when outcomes are biologically related.
A core contribution of the architecture is systems-level interpretability. The model provides attributions that quantify the contribution of biological processes to predictions. According to the authors, these process-level attributions were consistent with established disease mechanisms, supporting biological plausibility of the learned representations and offering a route to mechanistic insight beyond black-box predictions.
The omnigenic network captures non-linear interactions between variants. To interrogate these interactions, the authors applied an attribution method termed Integrated Hessians. This analysis revealed interaction patterns concordant with previously reported epistatic associations, indicating the model can recover pairwise and higher-order interaction signals that linear methods miss.
The authors report that individual-level records are available through UK Biobank and All of Us. Aggregate data and code are provided in supplementary files and on GitHub. Specific external datasets cited include genotypes for an SLE cohort available at a referenced dbGaP accession and scRNA-seq data hosted at a public collection URL. Ethical approval for the work was provided by the Ethics Committee of Charité Universitätsmedizin Berlin, and the authors confirm participant consent procedures and adherence to reporting guidelines.
Several authors disclosed consulting roles, ownership interests, or company affiliations; the remaining authors declared no conflicts. The manuscript is posted as a medRxiv preprint and has not undergone peer review; the authors explicitly caution that it should not be used to guide clinical practice. For full competing interest details and methodological specifics, the original preprint and supplementary materials should be consulted.