Inter-individual variability in drug response is influenced by genetic and non-genetic factors. Pharmacogenomics (PGx) links inherited variation in pharmacogenes to predictable differences in drug metabolism, transport, and exposure, enabling more individualized drug selection and dosing. The Clinical Pharmacogenetics Implementation Consortium (CPIC) and PharmVar provide harmonized genotype-to-phenotype translation and prescribing guidance. However, allele and phenotype frequencies vary by ancestry and local prescribing, so implementation priorities should be informed by population-specific data and by medication exposure within local healthcare systems.
This study addresses gaps in Thailand by combining scalable, array-based genotyping with electronic medical record (EMR) medication data to estimate both population-level PGx variation and the intersection of actionable phenotypes with observed drug exposure (referred to as “realized actionability”). The analysis used an Asian-optimized SNP array and CPIC level A/B gene–drug relationships to identify high-yield targets for pre-emptive PGx implementation.
The analysis comprised secondary data from two prospective studies enrolling adults affiliated with Siriraj Hospital and surrounding communities. Genotyping used the Infinium Asian Screening Array (ASA) v1.0 (Illumina). The study obtained institutional approvals from the Faculty of Medicine Siriraj Hospital, Mahidol University (COA no. Si 235/2026 and cohort-specific COAs). Written informed consent had been obtained in the original studies. Genotype data were accessed for research on 4 June 2024 and EMR data on 1 April 2026 after IRB approval.
A pre-specified PGx panel captured on the ASA included 11 clinically relevant genes represented by 26 markers. Gene selection aligned to CPIC/PharmVar definitions and focused on CPIC level A/B gene–drug relationships. The authors applied a hybrid required/optional calling policy for diplotype and phenotype assignment to reflect the constraints of array coverage while maintaining CPIC-aligned phenotype definitions.
Standard genotype quality control steps were applied: participants with sex discrepancies were excluded; duplicated variants removed; variants with call rate <90% and participants with call rate <97% were excluded. Kinship estimation using KING was used to remove one individual from each related pair (second-degree or closer) to avoid bias in frequency estimates. Variant coordinates were normalized to GRCh38. No genotype imputation was performed, and Hardy–Weinberg filtering was not applied to PGx panel variants to avoid removing phenotype-defining markers.
Across the panel, overall callability for gene results was 98.62%, exceeding 99% for most genes. Two exceptions with lower callability were CYP2C19 at 95.99% and NUDT15 at 90.28%, reflecting array marker coverage and calling policy constraints. Among the nine phenotype-coded genes, 95.99% of callable individuals carried at least one CPIC-actionable result, with a median of 2 actionable phenotypes per person (IQR 2–3). Gene-level actionable prevalence was highest for CYP3A5 (58.54%) and CYP2C19 (56.67%), followed by ABCG2 (45.10%) and UGT1A1 (27.37%). These prevalence estimates reflect the cohort profiled with the ASA and the CPIC-aligned phenotype definitions used in the study.
CPIC level A/B gene–drug relationships were linked to hospital EMR prescription and dispensation records to quantify realized actionability — the overlap between actionable phenotypes and actual medication exposure. EMR linkage identified 1,529 participants (32.58%) who had been exposed to at least one study medication captured in the analysis. The most commonly prescribed or dispensed medications among these participants were proton pump inhibitors and statins, specifically omeprazole (n = 658), atorvastatin (n = 606), and simvastatin (n = 603).
When medication exposure was considered, actionable phenotypes among users were common for specific gene–drug pairs. For omeprazole users, actionable CYP2C19 phenotypes were present in 55.02% of exposed participants. For statin users, actionable SLCO1B1 phenotypes occurred in 21.95–23.05% of exposed participants depending on the statin. These overlaps identify high-yield targets where pre-emptive genotyping could affect prescribing decisions for commonly used medications in this Thai hospital population.
The study demonstrates that an Asian-optimized SNP array can deliver high callability for a targeted PGx panel and produce clinically interpretable diplotype and phenotype calls at scale. By linking genotype-derived phenotypes to EMR medication records, the analysis quantifies realized actionability and highlights which gene–drug pairs are most likely to be encountered in routine care locally. The authors emphasize CYP2C19 (proton pump inhibitors) and SLCO1B1 (statins) as priority targets for pre-emptive implementation in this setting, given both genotype prevalence and medication utilization patterns.
All summary data supporting the study findings are included in the manuscript and supporting files. Individual-level genotype and EMR data are not publicly available due to institutional ethics and privacy restrictions; de-identified minimal datasets may be requested from the Siriraj Health Study subject to participant consent, institutional requirements, and IRB approval. The study funding source had no role in design, analysis, or reporting, and the authors declared no competing interests.
In summary, array-based PGx profiling using the Infinium ASA and CPIC/PharmVar-aligned phenotype definitions provided high callability and revealed that most individuals in this Thai cohort carried at least one actionable PGx phenotype. EMR linkage quantified the clinical reach of actionable phenotypes and identified CYP2C19 and SLCO1B1 gene–drug pairs as high-yield targets for pre-emptive testing, driven by both genotype frequency and local prescribing patterns. These findings support using scalable array platforms and EMR integration to inform national or institutional PGx implementation priorities. Limitations specific to array coverage and the calling policy (including lower callability for some genes) were reported, and individual-level data access is controlled to protect participant privacy.