Interpreting large-scale single-cell transcriptomic data to reveal disease mechanisms remains challenging. The authors present scGENet, a computational framework designed to translate gene embeddings from single-cell foundation models into context-specific, interpretable gene interaction networks. scGENet is built to operate on pretrained single-cell models and to produce transcriptome-scale gene modules that represent biologically meaningful cellular programs in a given context, here focused on the human midbrain and Parkinson's disease (PD).
The core objective is to leverage rich representations learned by foundation models across millions of cells and make those representations actionable for biological discovery, by constructing networks and modules that can be benchmarked against curated pathways, genetic risk loci, and independent transcriptional signatures.
scGENet involves fine-tuning pretrained single-cell foundation models on context-relevant transcriptomic data — in this study, data from human midbrain organoids. After fine-tuning, the framework derives gene embeddings and builds interaction networks and gene modules at transcriptome scale.
The authors benchmarked networks derived from multiple foundation models. They report that networks produced from a fine-tuned scGPT brain model exhibited the highest concordance with reference biological information: curated neuronal pathways, Parkinson's disease genetic risk loci, and independent patient-derived transcriptional signatures. This comparative benchmarking was used to prioritize the model-derived networks that best captured known neuron-specific and PD-relevant biology.
(Specifics of the foundation models compared, the number of models, training parameters, and quantitative benchmarking metrics were not reported in the abstract provided.)
Using scGENet applied to human induced pluripotent stem cell (iPSC)-derived PD midbrain organoids, the authors identified transcriptomic modules associated with several key biological programs. Prominent modules included those linked to neuronal differentiation, synaptic signaling, and cell-cycle regulation. These modules were interpreted as capturing molecular programs relevant to midbrain neuronal development and function that are altered in PD organoid models.
The use of midbrain organoids provided a controlled, human-derived cellular system to probe how disease-associated genetic backgrounds or perturbations impact neurodevelopmental and synaptic programs at single-cell resolution.
To connect transcriptional modules with cellular phenotypes, the authors used single-nucleus RNA sequencing. This analysis associated the scGENet-derived programs with changes in cellular composition within PD organoids. Key observations included:
These links between modules and cell-type composition provide cellular context for the dysregulated molecular programs uncovered by scGENet.
The authors extended their analysis by integrating scGENet findings from PD midbrain organoids with independent human substantia nigra datasets. This cross-dataset integration identified a conserved neurogenic program that is disrupted across both genetic and idiopathic forms of PD, indicating that the transcriptional signatures observed in organoids reflect disease-relevant biology present in human tissue.
The conserved program implicates disrupted neurogenic or developmental processes as a shared feature across PD etiologies in the midbrain, linking organoid-based observations to patient-derived tissue data.
The study establishes a strategy for extracting biologically interpretable gene networks from single-cell foundation models, demonstrating utility in uncovering disease-relevant molecular programs. By combining foundation-model embeddings, fine-tuning on context-specific data, module construction, and integration with single-nucleus and bulk tissue datasets, scGENet aims to enable systematic discovery of molecular programs across diverse tissues and disease contexts.
The authors position scGENet as a generalizable framework that can translate large-scale single-cell representations into interpretable networks for research into cellular programs and disease mechanisms.
The present summary is based on the article abstract and front matter. Specific methodological details — including dataset sizes, organoid sample counts, exact foundation models compared, hyperparameters for fine-tuning, network construction algorithms, statistical tests, effect sizes, and availability of code and data — were not reported in the provided source text. The manuscript is a preprint and has not undergone peer review.
Readers should consult the full preprint for complete methods, quantitative results, and supporting figures.