Cytochrome P450 (CYP) enzymes are responsible for metabolizing roughly three-quarters of clinically used drugs, so genetic variation in these genes contributes substantially to interindividual differences in drug response. Most modern variant-effect predictors operate on a homology-based paradigm, scoring missense variants by evolutionary conservation. The authors examined whether that paradigm holds for human CYP pharmacogenes and report systematic failure of conservation-based approaches on these enzymes.
The study highlights that pharmacogenes, including CYPs, violate the core assumption underlying many variant-effect models: that evolutionary conservation reliably indicates functional importance. As a consequence, predictors trained or calibrated on proteome-wide conservation signals may misclassify or produce ambiguous assessments for CYP variants.
Two state-of-the-art models were evaluated: AlphaMissense (AM) and Evolutionary Scale Modeling 2 (ESM-2). AM assigns variants to an "ambiguous" class at nearly twice the proteome-wide rate across six CYP proteins, indicating a systematic tendency to withhold confident functional labels for CYP missense changes. Within AM's ambiguous class, the reported scores show negligible correlation with experimental activity measurements from CYP2C9 deep mutational scanning (DMS), with Spearman's rho = 0.069. ESM-2 exhibits a similar failure pattern when applied to CYP variants.
These observations indicate that both a leading homology-informed predictor (AM) and a large language/model embedding approach (ESM-2) do not reliably map to experimentally measured functional effects for CYP variants under their default frameworks.
To probe whether features beyond conservation could disambiguate calls, the authors tested non-homology features including sequence position, substitution chemistry, distance to the binding site, and secondary structure. These features had only weak explanatory power for variant effects in CYP2C9 when evaluated against DMS activity. The weak signal suggests that simple structural or chemical annotations do not substitute for whatever critical information is missing from homology-based assessments of pharmacogene variation.
Using CYP2C9 DMS activity as the experimental ground truth, the team built a k-nearest-neighbors (k-NN) model operating over ESM-2 embeddings. They ensembled that k-NN model with AlphaMissense and ESM-2 masked marginal probability to address AM's ambiguous-class failures.
This ensemble produced marked improvements in correlation with experimental activity for variants initially labeled ambiguous by AM: ambiguous-class Spearman's rho rose from 0.069 to 0.715. Overall correlation across evaluated variants increased from rho = 0.638 to rho = 0.825. These results demonstrate that embedding-based local similarity (k-NN over ESM-2) combined with existing predictors can substantially recover signal tied to DMS activity for CYP2C9.
Although the ensemble substantially improved alignment with CYP2C9 DMS measurements, the authors report that this increased accuracy did not translate into agreement with clinical variant annotations. In other words, better prediction of experimental activity did not necessarily equate to concordance with clinical interpretations or annotations, which may reflect additional factors considered in clinical classification that are not captured by DMS or by the modeled features.
The authors draw attention to the multi-substrate nature of many CYP enzymes and suggest that a variant's functional effect can depend on the specific drug (substrate) being metabolized. They hypothesize that substrate identity is the key missing feature in current prediction models and propose that predicting function for multi-substrate enzymes like CYPs may require redefining function as substrate-conditioned rather than substrate-agnostic.
If substrate-conditioning is necessary, models would need to incorporate or be conditioned on specific ligands or substrates to produce clinically relevant effect predictions for CYP variants.
The preprint is available as a bioRxiv entry and the authors provide supplementary material and links to data and code. The authors note no competing interests. Because this work is a preprint, it has not been peer reviewed.
Findings underscore limitations of applying proteome-wide homology assumptions to pharmacogenes. For clinical pharmacogenomics, the results signal caution when relying on homology-derived variant-effect scores for CYP genes. Future model development may need to integrate substrate-specific biochemical context and to validate predictions against both experimental DMS data and clinical annotation frameworks to improve translational utility.