Protein‑ligand affinity (PLA) prediction is a central task in AI‑driven drug discovery, but many high‑accuracy methods depend on explicit modeling of interactions and require costly conformation preparation and specialized data encoding. These preparation steps limit throughput and practical scalability for large virtual screening campaigns. The study introduces a preparation‑free approach that seeks to reconcile predictive accuracy with computational efficiency by eliminating the need for explicit interaction encoders and by leveraging existing pre‑trained molecular representations.
The authors first evaluate whether off‑the‑shelf, pre‑trained molecular representation models can substitute for complex interaction encoders in PLA tasks. They perform a unified and diverse assessment of sequence‑, graph‑, and image‑based pre‑trained encoders. The evaluation highlights generally strong overall predictive performance from pre‑trained models, while also revealing variability in performance across protein families. These results are positioned as the first practical guidance for selecting encoder types when constructing interaction‑free PLA predictors.
To maintain expressive capacity without large computational cost, the study adopts a mixture‑of‑experts (MoE) strategy, a sparse, dynamic parameterization approach borrowed from large language model architectures. Systematic ablation experiments are reported to identify the key design choices needed to apply MoE effectively in molecular prediction settings. The ablations investigate architectural and routing design elements that influence both predictive performance and parameter efficiency.
Building on the encoder benchmarking and MoE design findings, the authors present HydrAffinity, an interaction‑free model that combines pre‑trained encoders with a dynamic, sparse MoE backbone. HydrAffinity uses pre‑trained representations for molecular inputs and a sparse routing mechanism to activate a subset of expert parameters per example, enabling parameter‑efficient learning without explicit interaction computations or conformation preparation. The model is described as designed for high computational efficiency while retaining expressive predictive power.
According to the manuscript, HydrAffinity outperforms all existing interaction‑free methods on the CASF‑2016 benchmark and achieves performance on par with state‑of‑the‑art interaction‑based methods. This comparative result is presented as evidence that a preparation‑free, MoE‑based strategy can deliver competitive affinity prediction without the overhead of interaction modeling. The paper frames HydrAffinity as closing the gap between efficient, high‑throughput models and more expensive interaction‑driven approaches.
The authors analyze the routing behavior of the MoE and report that expert activation follows distinct, family‑specific patterns. This observation is presented as interpretable evidence that the MoE dynamically parameterizes predictions in a manner aligned with protein class differences. The family‑wise activation patterns are used to support claims that the model adapts its computational pathway based on protein attributes, which may aid interpretability and targeted improvement.
HydrAffinity was also evaluated in zero‑shot settings on the DUDE‑Z and LIT‑PCBA benchmarks. The study reports strong EF5% performance in these tests, indicating effective early enrichment when used as a screening pre‑filter. These results are highlighted to support the model's proposed role as a practical, scalable early‑stage tool that can rapidly triage candidates prior to heavier computational or experimental follow‑up.
The manuscript emphasizes HydrAffinity's computational efficiency and scalability, arguing that an interaction‑free, MoE‑driven model can serve as an effective early filter in drug discovery pipelines. The article page links to data and code resources via external links; specific repository names or implementation details are provided on the article's data/code section. As this work is a preprint, it has not been peer reviewed. The authors declare no competing interests and acknowledge funding support from the National Natural Science Foundation of China and the Natural Science Foundation of Gansu Province.
HydrAffinity demonstrates that combining pre‑trained encoders with a mixture‑of‑experts architecture yields a preparation‑free, parameter‑efficient model capable of competitive PLA prediction. By avoiding explicit interaction modeling and conformation preparation, the approach promises higher throughput for early‑stage virtual screening while retaining accuracy comparable to interaction‑based methods on established benchmarks. Routing behavior that aligns with protein family distinctions offers an interpretable aspect to the model's dynamic parameterization. Because the study is reported as a preprint, further validation through peer review and independent replication will be important to confirm practical utility and robustness in diverse screening contexts.