This study examined the transferability of 141 genetic scores for metabolomic traits that were originally developed in the INTERVAL Study sample of European (EUR) ancestries. The authors evaluated these scores in the Mexico City Prospective Study (MCPS), an admixed American (AMR) cohort comprising 132,336 participants. When INTERVAL-trained scores were applied to MCPS, the median predictive R2 across evaluated metabolic traits was 0.027, indicating limited predictive performance of EUR-derived scores in this AMR population.
The analysis highlights the challenge that metabolomic profiling and genetics-driven prediction face when moved across ancestry groups. The authors note that metabolomics remains infrequently performed in non-European ancestries despite these groups bearing a disproportionate burden of metabolic disease.
To address the observed performance gap, the investigators trained Bayesian ridge prediction models using MCPS data. Models were evaluated on a withheld 20% subset of MCPS participants. Training within the AMR cohort substantially improved prediction: the median R2 rose to 0.083 on the withheld subset. This within-cohort training approach therefore produced a marked increase in explained variance compared with applying INTERVAL-derived models directly.
The results demonstrate that retraining or recalibration of genetic prediction models in the target admixed population can meaningfully increase predictive accuracy for metabolomic traits compared with using models developed in ancestrally distinct populations.
The study further tested transferability by comparing MCPS-trained and INTERVAL-trained models in an independent set of UK Biobank participants of AMR ancestry (n=600). In this comparison, the MCPS-trained models achieved a median R2 of 0.070, while the INTERVAL-trained models reached a median R2 of 0.046. Thus, models trained in MCPS transferred better to an external AMR sample in UK Biobank than models trained in the EUR INTERVAL cohort.
This cross-cohort comparison supports the value of training prediction models in admixed target populations to improve external validity and portability to other AMR datasets.
The authors assessed downstream implications for metabolome-wide association studies (MWAS) by applying the prediction models to AMR participants in the All of Us cohort. Using MCPS-trained models to predict metabolomic traits—rather than INTERVAL-trained models—yielded five times as many significant associations after false discovery rate (FDR) correction (FDR-corrected P<0.05) across three cardiometabolic diseases: ischaemic heart disease, type 2 diabetes, and chronic kidney disease.
This increase in significant associations suggests improved statistical power for genetics-driven MWAS when prediction models are trained in ancestrally matched or admixed cohorts. The finding underscores that model training population matters for discovery of metabolomic biomarkers linked to disease in AMR groups.
The study provides empirical evidence that genetics-driven metabolomic prediction benefits from training within admixed American cohorts. Key implications include:
Retraining genetic scores in the target admixed cohort can substantially increase explained variance for metabolomic traits compared with applying European-trained models directly.
Improved predictive performance translates into greater power for metabolome-wide association testing and more detected disease associations in AMR participants.
Making cohort-trained genetic scores available publicly may help reduce disparities in omics research by enabling better-powered analyses across diverse AMR datasets.
The authors make the genetic scores openly available via the OmicsPred portal (www.OmicsPred.org) to facilitate broader use and to help address global inequities in metabolomic and genetics research.
The analyses used data from the Mexico City Prospective Study, the UK Biobank, and the All of Us Research Program. The Mexico City Prospective Study data are available to bona fide academic researchers and include an ancestry-specific allele frequency browser (https://rgc-mcps.regeneron.com/). UK Biobank and All of Us data were accessed under appropriate application and authorised user procedures.
Ethical approvals and participant informed consent were obtained for all included datasets. The MCPS received approvals from the Mexican Ministry of Health, the Mexican National Council for Science and Technology, the Central Oxford Research Ethics Committee, and the Medical Ethics Committee of the National Autonomous University of Mexico; all participants provided written informed consent. UK Biobank and All of Us research governance and consent processes were followed in accordance with each resource's oversight arrangements.
Additional methodological details, model specifications, and supplementary materials are available in the study resources and the authors' supplementary material. The authors reported no competing interests. The genetic scores and related resources are accessible on OmicsPred to support application of these models across AMR cohorts.