Self-supervised representation learning is increasingly used to convert medical images into quantitative phenotypes for downstream biological and genetic discovery. However, statistical reproducibility of learned features does not guarantee the features correspond to the intended anatomy. This study investigated how field-of-view confounding and acquisition-related variation affect genetic discovery from self-supervised cardiac-imaging phenotypes.
The authors trained a video masked-autoencoder on 69,932 UK Biobank cardiac cine-MRI studies. The model produced a latent representation for each study, which the investigators used as image-derived phenotypes for genome-wide association analysis. Details of model architecture and training parameters beyond the use of a video masked-autoencoder were not reported in the source.
Genome-wide association analysis of the latent axes showed that 18 of 20 leading latent axes were heritable and returned well-calibrated association statistics. These initial diagnostics suggested genetic signal was present in the learned representation; standard genomic-control and LD-score regression diagnostics appeared appropriate by conventional metrics.
Despite well-calibrated association statistics, the latent representation encoded substantial non-cardiac information. Specifically, the representation contained signals related to body size, stature and the imaging centre where scans were acquired. A linear-probe model reported an R² of 0.55 for imaging site, indicating strong encoding of acquisition site information within the representation.
Notably, the standard genomic-control and LD-score diagnostics used in GWAS did not identify this source of phenotype-level confounding. This indicates that phenotype-level nuisance variation arising from image field of view and acquisition can escape commonly used post hoc statistical checks.
The investigators evaluated strategies to attenuate nuisance information upstream of association testing. Two key preprocessing steps proved effective: restricting the image field of view to the heart, and residualising out body and acquisition covariates prior to dimensionality reduction. Applying these corrections substantially reduced both linear and non-linear nuisance signals in the learned representation while retaining cardiac signal.
Residualising covariates and limiting the imaged region focuses the model on the target anatomy and reduces the extent to which principal components or latent axes capture variation driven by body habitus or scanner/site differences. The study emphasizes that such upstream corrections can change the basis on which the model learns and thus the downstream genetic associations derived from those axes.
The authors compared the upstream correction approach with an alternative strategy: adjusting body and acquisition covariates only during association testing, after representations were learned. While this latter approach attenuated some nuisance associations at the GWAS stage, it recovered substantially less cardiac-associated genetic signal than the upstream correction.
This finding is consistent with the interpretation that nuisance variation influenced the latent basis during dimensionality reduction; once the learned representation had incorporated field-of-view and acquisition effects, later statistical adjustment could not fully recover cardiac-specific variation that had been mixed into the principal-component axes.
Using the corrected latent representation — restricted to the heart and residualised for body and acquisition covariates prior to dimensionality reduction — the study identified new associated loci beyond those detected using supervised phenotypes at matched sample size. These loci shared genetic architecture selectively with cardiac-conduction traits and were localized to cardiac structures within the imaged field of view, supporting the notion that correcting learned representations can improve the specificity of genetic discovery for target anatomy.
The source reports that aggregate per-axis summary statistics and figure-source tables are available from the corresponding author on request, but does not provide the full list of newly identified loci or locus-level statistics within the text provided.
The main implication is that confounding in learned medical-image phenotypes can arise upstream of association testing — during model training and dimensionality reduction — and thus may not be detectable by standard GWAS diagnostics. The work argues for auditing learned representations for nuisance signals related to field of view, body habitus and acquisition centre, and for applying corrective measures where appropriate prior to genetic association testing.
Practically, this suggests that researchers using self-supervised or unsupervised image representations for genetic discovery should consider image cropping or masking to focus on target anatomy and should residualise known acquisition and anthropometric covariates before learning lower-dimensional representations.
Data used were from UK Biobank and are accessible to approved researchers under application 65439. External GWAS and eQTL resources used in the study are publicly available (CARDIoGRAMplusC4D, Aragam et al. coronary artery disease GWAS, GTEx v8 cis-eQTLs, and reference GWAS used for genetic-correlation analyses). Aggregate per-axis summary statistics and figure-source tables are available from the corresponding author on request.
Ethical oversight was provided by UK Biobank (North West Multi-centre Research Ethics Committee, REC reference 11/NW/0382) and participants provided written informed consent at recruitment. The authors declared no competing interests.
Limitations reported in the source include that certain methodological details and the complete list of newly identified loci were not provided in the text available here. The source emphasizes that standard genomic-control and LD-score checks did not detect the phenotype-level confounding, underscoring the need for representation-level audits that were the focus of this work.