Pharmacogenetic (PGx) testing can inform safer and more effective prescribing, but the utility of any PGx result depends on the genomic assay used. Genotyping arrays such as the Illumina Global Screening Array (GSA) v3 assay interrogate predefined probes and therefore miss variants not represented on the array. In contrast, low-pass whole-genome sequencing (LP-WGS) at approximately 1× coverage does not rely on fixed probe design and can recover additional variants after imputation. This study directly compares the two approaches for PGx variant detection and phenotype inference in a real-world hospital biobank sample.
The analysis used samples from 500 participants in the CHUV Genomic Biobank (BGC) selected for electronic health record evidence of exposure to pharmacogenetically actionable drugs and reported adverse drug reactions. The study protocol and amendments were approved by the Ethics Committee of canton Vaud (BGC ref 144/12; project-ID: 2024-00901) and complied with the Declaration of Helsinki and the Swiss Human Research Act. Authors report that participants provided general consent for use of coded clinical data and biological samples for research.
The two genomic assays compared were the Illumina GSA v3 genotyping array and LP-WGS at ~1× coverage. Concordance between imputed array genotypes and LP-WGS-derived genotypes was evaluated at multiple levels: genome-wide single-site concordance, across 20 actionable pharmacogenes using PharmCAT to derive star-alleles and metabolizer phenotypes, and at HLA loci using available imputation strategies. The analysis emphasized both shared-site concordance and the additional variant recovery afforded by LP-WGS, with particular attention to rare alleles and gene structural complexity.
Genome-wide concordance between the imputed array and LP-WGS data was high. The reported median concordance was 99.63% with an interquartile range of 99.59% to 99.64%. This indicates that, at loci covered by both approaches and after imputation, agreement is strong.
LP-WGS captured a larger fraction of pharmacogenetically relevant variants overall, particularly rare alleles that were not present on the GSA array. At shared sites, concordance remained high, but LP-WGS extended the set of observed variants beyond those available from the array, potentially expanding downstream PGx interpretability for loci where relevant variants were absent from the genotyping array.
Predicted metabolizer phenotype concordance across the tested pharmacogenes exceeded 98% for most genes. However, the authors observed gene-specific differences in phenotype classification, indicating that high overall concordance does not eliminate discrepancies at individual loci or for particular allele definitions.
LP-WGS reduced missing phenotype assignments for certain loci—most notably CYP2C19 and NAT2—by improving the resolution of star-allele structure. This reduced the number of unassignable phenotypes for those genes in the study cohort. By contrast, for structurally complex or incompletely characterized genes such as CYP2C9 and CYP2D6, the broader variant recovery afforded by LP-WGS sometimes increased indeterminate classifications rather than yielding clearer clinical interpretations. Thus, expanded variant detection does not uniformly translate to improved, actionable phenotype calls for all pharmacogenes.
Concordance at HLA loci varied depending on imputation strategy. In this dataset, SNP2HLA applied to the GSA array performed marginally better than the LP-WGS-based HLA approach used by the authors. This result highlights that HLA inference remains sensitive to both input data type and the chosen imputation tool.
Overall, LP-WGS at ~1× provides broader pharmacogenetic variant coverage and improves phenotype resolution for selected genes, especially where array probe content is lacking. Nevertheless, LP-WGS did not resolve all clinically important loci and in some complex genes produced more indeterminate classifications. The authors conclude that LP-WGS merits further evaluation as a scalable PGx screening approach, particularly in settings that value long-term reuse of comprehensive genomic data, but it is not a universal substitute for locus-specific, high-resolution genotyping or sequencing when clinical interpretation depends on complex structural alleles.
Individual-level BGC data used in this study contain sensitive personal information and cannot be openly shared under Swiss privacy legislation. Coded, non-identifiable individual-level data may be available on request to qualified investigators who meet the Biobank's access criteria; requests are reviewed by the Biobank's Scientific Committee. The report is presented as a medRxiv preprint and has not been peer-reviewed; the authors declared no competing interests.