Speciesformer is presented as a cross-species, generative single-cell foundation model aimed at representing, describing and predicting cellular states across biological contexts. The model's stated objective is to separate evolutionarily conserved biological programs from species-, tissue- and context-specific variation, enabling transfer of knowledge across species while modeling conditional state transitions.
The authors position Speciesformer as an advance over existing single-cell foundation models by combining evolutionary representation learning with a unified generative framework. According to the abstract, this design supports both representation probing and conditional generation of cellular states.
Speciesformer was pretrained on a collection termed SpeciesCorpus. The corpus reportedly includes 131 million single-cell profiles drawn from 11 species, representing 154 tissues and more than 923 cell types. This wide-ranging dataset is used to map species-specific gene information into a common, evolution-informed gene space and to expose the model to diverse cellular contexts.
The source abstract does not provide further breakdowns of the SpeciesCorpus (for example, per-species cell counts, sequencing modalities, quality control procedures, or tissue-to-cell-type mappings). Those implementation and dataset curation details are not reported in the provided excerpt and would need to be consulted in the full article or supplementary materials.
A core component of Speciesformer is a mapping from species-specific genes into a shared gene embedding space informed by evolutionary relationships. The model's encoder learns representations for both cells and genes that are intended to be transferable across species and contexts.
These learned representations are described as enabling biological representation probing and cross-species knowledge transfer, and as providing a common state space to resolve which cellular programs are conserved versus which are context dependent. The abstract emphasizes that the shared representation enables the model to distinguish transferable biological principles from context-specific variation.
The abstract does not report the precise methods used to derive the evolution-informed gene space (for example, whether orthology mappings, sequence similarity, phylogenetic priors, or other constraints were applied), nor does it report architecture details of the encoder. Such methodological specifics are not present in the provided source text.
Built upon the shared representation, Speciesformer employs a unified generative architecture to model cellular states and state transitions under both semantic and interventional conditions. The generative framework is presented as capable of producing cellular states conditioned on textual or semantic descriptions and of modeling transitions that reflect interventions.
The model is described as unified in that it supports multiple generative tasks from the same representation space, rather than being limited to discriminative analyses or specialized generation modes. The abstract does not detail the internal generative mechanisms, loss functions, or training objectives used.
One reported capability of Speciesformer is bidirectional generation between transcriptomic states and biological text descriptions. In other words, the model is stated to generate text-like descriptions conditioned on transcriptomic inputs and, conversely, to generate transcriptomic-like outputs conditioned on biological semantics.
The source abstract does not provide examples, evaluation metrics, or benchmarks demonstrating the quality of these bidirectional generations. Those demonstrations and quantitative assessments would be expected in the full manuscript but are not present in the abstract.
Speciesformer is also described as able to predict post-perturbation transcriptomes from initial cell states together with descriptions of interventions. The authors claim this predictive ability extends to previously unobserved cellular contexts, suggesting the model can generalize intervention effects across species and cell types by leveraging the shared evolutionary representation.
The abstract does not include specific validation results, error rates, comparative baselines, or case studies showing these predictions. Details on how interventions are encoded, the types of perturbations tested, and the degree of generalization achieved are not reported in the provided source text.
By unifying evolutionary variation, biological semantics and conditional state transitions, Speciesformer is presented as a step toward a generative virtual cell framework. The model aims to provide a common platform for representing conserved cellular programs, describing cellular states in natural language, and predicting state changes across biological contexts.
Potential applications implied by the abstract include cross-species knowledge transfer in single-cell analysis, generation of human-readable descriptions of cell states, and in silico prediction of perturbation outcomes. The abstract frames these as capabilities enabled by combining a large cross-species corpus with evolution-informed gene mappings and a unified generative architecture.
This work is a preprint posted to bioRxiv and has not been certified by peer review. The authors declare no competing interests. The abstract is the primary source for the summary above; many technical and evaluation details (model architecture, training procedures, quantitative performance, dataset provenance and curation, ablation studies, and examples of generated outputs) are not reported in the provided excerpt. Readers seeking implementation specifics and validation results should consult the full preprint and supplementary materials referenced by the authors.
Speciesformer is described as a large-scale cross-species single-cell foundation model pretrained on a 131-million-cell SpeciesCorpus. It maps species-specific genes into a shared evolution-informed gene space, learns transferable cell and gene representations, and implements a unified generative architecture supporting bidirectional transcriptome–text generation and intervention-conditioned transcriptome prediction. The abstract positions the model as advancing generative virtual cell modeling across species, while noting that detailed methods and performance data are not provided in the abstract and remain to be reviewed in the full manuscript.