Models that couple tissue imaging with pretrained gene features aim to predict spatial gene expression for genes not used when fitting downstream predictors. Success in predicting held-out genes can reflect at least two distinct capabilities: estimating a gene's overall mean expression across tissue locations and recovering its spatial variation across the tissue. The study assessed how broadly pretrained gene representations transfer these two components in the context of virtual spatial transcriptomics.
The evaluation spanned four cohorts that covered three human brain regions and HER2-positive breast cancer. Held-out genes were evaluated on held-out individuals to test cross-gene generalization in realistic conditions. The analysis examined spatial predictors that used fixed pretrained gene representations produced by Decima or scGPT, and compared performance against matched random vectors and mean-only baselines.
To understand what pretrained gene representations contribute, the authors separated prediction performance for held-out genes into two components. The first component is the gene's mean expression across tissue locations—essentially a location-averaged level. The second component is the gene's spatial pattern, the deviation from that mean across locations. By decomposing performance, the analysis distinguishes gains driven by better mean estimation from gains driven by improved spatial recovery.
Across the tested settings, the dominant contribution of pretrained gene representations was to transfer mean expression. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. In other words, most of the predictive improvement over random baselines derived from better estimation of each gene's average expression level across locations.
Independently fitted mean-only models that used the same pretrained representations but no tissue images retained between 90% and 99% of the corresponding gain in full-matrix correlation. This finding shows that a large fraction of the representational benefit can be captured without any image-based spatial modeling, indicating strong transfer of mean-level information from the pretrained embeddings.
By contrast, gains in recovering spatial expression patterns were smaller on average. Spatial improvements were detectable but more limited than mean-level gains, and they depended on context. Specifically, spatial gains increased when the expression variation in the training tissue was larger, indicating that the ability to learn and transfer spatial patterns benefits from stronger spatial signal in training data.
Moreover, spatial recovery differed across cohorts and across the pretrained representations evaluated. This heterogeneity shows that some datasets and some pretrained embeddings are more effective at supporting spatial generalization than others.
The balance between transferred mean expression and transferred spatial structure varied across the four cohorts and between the two pretrained representations examined (Decima and scGPT). While mean-level transfer was broadly observed across settings, the magnitude of spatial gains was cohort- and representation-specific and correlated with the degree of expression variability present in training tissues. These results imply that dataset composition and the nature of the pretrained representation both influence what aspects of gene expression can be generalized across genes.
The findings indicate that cross-gene generalization in virtual spatial transcriptomics is not a single capability. Pretrained gene representations broadly and reliably convey mean expression information across genes, which explains most improvements over random baselines. Improvements in spatial-pattern recovery are smaller, selective, and dependent on training-data variability and the chosen representation. Practically, this implies that methods relying on pretrained gene embeddings should not assume uniform transfer of fine-grained spatial signals; instead, users and developers should evaluate mean-level and spatial-level generalization separately when benchmarking or deploying virtual spatial transcriptomics models.
The source reports links to public data repositories and supporting materials. The authors provided GEO accession identifiers and DOIs for supplementary materials and code repositories. Specific accession numbers and DOIs were listed in the source.