Genetic fine-mapping aims to identify causal variants inside loci flagged by genome-wide association studies, but strong linkage disequilibrium (LD) and large-scale summary datasets make the sparse variable-selection problem difficult. The SuSiE algorithm is widely used for its computational speed, variational inference, and output of posterior inclusion probabilities (PIPs) and credible sets. However, a single variational fit can fail to resolve LD ambiguity, become trapped in a poor local optimum, or underrepresent uncertainty among competing causal configurations.
The authors introduce SuSiNE (Sum of Single Non-central Effects) to address these limitations by incorporating signed functional annotations into the generative prior while retaining desirable computational and inferential properties of SuSiE.
SuSiE's variational approach is efficient and supplies PIPs and credible sets, but a single fit can misrepresent uncertainty when the posterior surface has multiple basins. Standard post-processing steps such as purity filtering — intended to remove low-confidence credible sets — can inadvertently discard informative signal and may degrade overall performance according to the authors' analyses.
SuSiNE extends the SuSiE single-effect framework by adding a prior-mean channel μ0 = c a, where a denotes signed functional annotations and c is a scalar hyperparameter. This construction preserves conjugacy of the single-effect updates, maintains credible-set construction, and keeps sufficiency of summary-statistics for inference. The prior-mean channel permits the model to leverage external signed annotation information while working in the same computational framework as SuSiE.
A key model-level behavior is that the resulting single-effect Bayes factor effectively self-gates on agreement between annotation sign and observed association direction. When annotation sign and association direction agree, the annotation contributes weight; when they disagree, the mechanism limits annotation influence. This self-gating reduces the impact of noisy or mis-specified annotations on effect estimates.
The authors propose new effect-level diagnostics to characterize model fits and variational behaviour. These diagnostics measure (1) concentration of posterior mass, (2) accuracy of effect estimation, and (3) movement of the fitted basis between fitted variational basins. These metrics are intended to give deeper insight into convergence issues and whether alternative basins reflect meaningful alternative explanations.
To explore and summarize multiple variational basins, SuSiNE is combined with a grid-based ensembling strategy and cluster-weight aggregation. The ensemble approach runs the model across a grid of prior settings and clusters resulting fits, aggregating cluster weights to produce pooled PIPs and summaries that reflect multiple plausible basins instead of relying on a single variational solution.
In oligogenic simulations calibrated to AlphaGenome eQTL benchmarks, the ensemble improved pooled area under the precision-recall curve (AUPRC) for recovery of the largest-effect causal variants from 0.2474 (SuSiE-equivalent baseline) to 0.3130, a delta of 0.0656 with a reported 95% paired-bootstrap confidence interval [0.0591, 0.0722]. At 75% precision, recall increased from 11.9% to 19.3%, a relative gain of 61.7%.
The authors report AUPRC gains were robust across variable annotation quality and across different sparse and diffuse genetic architectures. They note, however, that if null annotations are strongly aligned with association in the wrong direction, ensemble gains can reverse.
Using GTEx Lung summary statistics and reference LD, SuSiNE placed nontrivial weight on annotation-informed fits at 7 of 20 loci and altered which variants received high PIP compared with SuSiE-equivalent fits. The ARSA locus showed the clearest durable shift in PIP allocation attributable to annotation-informed fits, while a large shift at the YDJC locus coincided with a reference-LD discrepancy. The authors used internal diagnostics and found little evidence of strong annotation confounding in this panel. They emphasize these analyses use reference rather than in-cohort LD and are intended to illustrate method behavior rather than provide definitive variant-level discoveries.
This work is presented as a bioRxiv preprint (not peer reviewed). Supplementary materials, data, and code are referenced in the preprint via multiple DOIs and GitHub repositories provided by the authors. The AlphaGenome team at Google DeepMind supplied API access used to generate the signed variant-effect annotations; they did not influence study design, analysis, interpretation, or publication decisions.
The preprint discloses that one author is employed by Calico Life Sciences and contributed in a personal capacity; the AlphaGenome team provided annotation access but no role in the study. Funding sources and institutional support are listed in the preprint. The authors reiterate that the manuscript is a preprint and has not been certified by peer review.