scDIVA is a semi-supervised deep generative model that adapts the Domain Invariant Variational Autoencoder to single-cell RNA-seq for fine-grained label transfer in tumor-immune atlases. The model decomposes each cell's transcriptomic profile into three latent subspaces using three separate encoders: one for cell type, one for batch, and one for residual variation. A single decoder reconstructs the original expression profile from the concatenated latent embeddings.
Auxiliary classifiers are attached to the cell type and batch embeddings during training to encourage each encoder to capture its intended source of variation. Because the architecture explicitly separates batch and cell type embeddings and uses classifier supervision, scDIVA's cell type latent space is designed to be batch-invariant by construction rather than relying on explicit correction or adversarial approaches.
The semi-supervised design lets scDIVA leverage labeled reference atlases for annotation transfer while retaining flexibility to represent query-specific variation in the residual embedding. The source reports that model implementation and supporting data resources were made available by the authors.
The authors benchmarked scDIVA against four established reference-mapping approaches: Harmony/Symphony, scANVI with scArches, scPoli with scArches, and Seurat label transfer. These comparisons were performed across six tumor-immune atlases that together span five cancer types.
According to the preprint, scDIVA attained the highest mean macro-F1 score in five of the six atlases tested, indicating improved overall label-transfer performance across diverse tumor-immune datasets. In addition, scDIVA achieved the highest biological conservation scores relative to the comparator methods reported in the source, suggesting better preservation of biologically meaningful structure during integration and annotation.
The benchmarking design therefore assessed both classification accuracy (macro-F1) and conservation of biological signal, and scDIVA outperformed the listed alternatives on these metrics in the majority of tested atlases.
To identify cell populations present only in the query (out-of-reference, OOR), the authors adapted the Milo differential abundance (DA) framework for atlas-scale use with scDIVA embeddings. Their modifications included the addition of a directional test, a correction to the spatialFDR weighting procedure, and parallelization of neighborhood distance computations to scale to large atlases.
This adapted DA pipeline was applied to scDIVA's cell type embeddings with reference-versus-query membership as the condition of interest. The design enabled detection of cell states that are absent from the reference rather than forcing them into the nearest reference label. The authors report that this approach correctly identified cell types intentionally held out from the reference as OOR.
An illustrative result reported in the preprint was that the procedure flagged exhausted CD8 T cells from a tumor-immune atlas as out-of-reference when compared with a healthy pan-tissue immune reference, supporting the pipeline's capacity to detect disease- or tumor-associated states that are not represented in healthy references.
Conversely, when applied to two independent colorectal cancer cohorts, the adapted DA procedure confirmed a conserved immune landscape between the cohorts, showing that the method can both discover novel, query-specific populations and verify shared biology across studies.
The source describes two principal application outcomes demonstrating scDIVA's utility. First, detection of OOR populations highlights the importance of distinguishing truly novel or disease-associated cell states from mislabeled or batch-driven artifacts during label transfer. The example of exhausted CD8 T cells being identified as OOR relative to a healthy reference underscores how scDIVA plus the adapted Milo DA can reveal clinically relevant immune states present in tumors.
Second, the confirmation of conserved immune composition across independent colorectal cancer cohorts illustrates that scDIVA can preserve and recover shared biological signals across studies, supporting comparative analyses and meta-atlas construction.
Together, these examples show that coupling a disentangled latent representation with an FDR-controlled, atlas-scaled DA test addresses two complementary challenges in single-cell reference mapping: accurate fine-grained annotation and controlled detection of reference-missing states.
The preprint lists authors affiliated with Memorial Sloan Kettering Cancer Center, Bristol Myers Squibb, and Chan Zuckerberg Biohub. The manuscript reports that code and data resources supporting the work are available through the authors' provided repositories. The study acknowledges funder support as declared in the source.
Competing interest disclosures provided in the preprint state that one author is currently an employee of Bristol Myers Squibb but contributed while a student at the primary institution and that another author completed an unrelated internship; additional disclosures include a board/advisory relationship and IP held by one author unrelated to this work. The remaining authors declared no competing interests. The preprint was posted on bioRxiv and is made available under a CC-BY-NC-ND 4.0 International license.