Target-based and structure-guided approaches remain central to drug discovery, but they can be limited when predefined targets or binding pockets do not capture the full biology of disease. The authors argue that system-level readouts such as gene-expression signatures provide scalable phenotypic information that can guide molecule design when conventional target definitions are insufficient. The main challenge is maintaining the connection between generated chemical structures and the intended biological response during de novo molecular generation.
The study introduces Tx2Mol, a framework that translates gene-expression signatures (transcriptomes) into candidate molecules while keeping biological guidance active throughout the generation process. Tx2Mol is described as a transcriptome-guided de novo molecular design system intended to produce chemically plausible structures that remain connected to the phenotypic input signal.
Tx2Mol was evaluated across three biological contexts: bulk gene perturbation datasets, single-cell perturbation profiles, and patient-derived disease signatures. These settings were selected to test the method's applicability from controlled perturbation experiments to noisy single-cell measurements and clinically relevant patient transcriptomes.
The authors assessed Tx2Mol along three validation dimensions: (1) chemical plausibility of generated molecules, (2) structural compatibility with known ligands and predicted binding behavior, and (3) preservation of the intended phenotypic transcriptional responses after design. These complementary dimensions were used to evaluate both the molecular realism and the biological relevance of the output.
On a set of 10 cancer-relevant bulk gene-perturbation benchmarks, Tx2Mol was compared against nine transcriptome-guided baseline methods. Tx2Mol outperformed these baselines overall. Quantitatively, the framework improved maximum Tanimoto similarity to known ligands by 24.10% on average across the benchmarks. A particularly large improvement was observed for HDAC1, where maximum Tanimoto similarity increased by 50.67%.
Beyond similarity metrics, structure-based analyses were performed and supported that Tx2Mol generated structurally novel candidates that had favorable predicted binding to targets of interest. These analyses are reported as additional evidence that Tx2Mol can produce molecules that are both novel and compatible with target binding expectations.
The authors evaluated Tx2Mol on single-cell perturbation profiles, which are inherently noisier than bulk measurements. Tx2Mol generalized to these noisy inputs, indicating that the transcriptome-guided generation approach can accommodate variability typical of single-cell datasets while maintaining the phenotype-directed design objective.
To test whether generated molecules retained the intended transcriptional effects, in silico drug-perturbation validation was performed. Tx2Mol preserved drug-induced transcriptional responses in these computational validations, supporting the framework’s capacity to maintain phenotypic guidance through the generation pipeline.
When provided with patient-derived disease signatures, Tx2Mol steered molecular generation toward chemical space occupied by approved drugs. This result is presented as evidence that clinically relevant transcriptomes can bias generation toward molecules with properties similar to known therapeutics, supporting candidate prioritization for translational follow-up.
The authors conclude that gene-expression phenotypes can serve as actionable guidance signals for phenotype-directed molecular design. Tx2Mol demonstrated improved similarity to known ligands relative to transcriptome-guided baselines, produced structurally credible and novel candidates with favorable predicted binding, generalized to single-cell perturbations, preserved transcriptional responses in silico, and guided designs toward approved-drug chemical space when using patient data.
Note on manuscript status
This work is reported as a preprint on bioRxiv and has not been certified by peer review. The authors declared no competing interests. Details beyond the abstract and summary results reported here (for example, training data specifics, model architecture, hyperparameters, exact baseline methods, and quantitative results beyond those cited) were not reported in the provided source excerpt.