Accurate annotation of clusters derived from single-cell RNA sequencing (scRNA-seq) is an essential and recurring step in scRNA-seq analysis workflows. The preprint describes celltypeEnrich, a cluster-level annotation tool designed to address limitations of some existing methods, namely time-consuming workflows, limited reproducibility, and restricted tissue or species coverage. The authors position celltypeEnrich as a reproducible, broadly applicable approach that synthesizes evidence across multiple reference resources to produce a consensus cell-type label for each cluster.
celltypeEnrich takes as input gene lists for clusters and tests for enrichment of cell-type–specific genes drawn from reference datasets. Enrichment is evaluated using a hypergeometric test, which assesses whether the overlap between an input gene list and a reference cell-type gene set is greater than expected by chance. Enrichment results are computed across multiple references and then combined to assign a consensus annotation to the queried cluster.
The approach is cluster-level rather than cell-level, meaning the tool annotates the aggregated marker lists produced for clusters rather than labeling each individual cell. The consensus mechanism integrates outputs from up to 26 reference datasets, providing broader coverage across tissues and species compared with tools that rely on a single reference.
The authors report that enrichment is performed against as many as 26 reference datasets. The diversity of references is intended to improve tissue and species coverage and to reduce dependence on any single atlas or marker collection. The preprint indicates that relying on multiple sources allows the tool to deliver annotations even when individual reference datasets have incomplete coverage for the tissue or species under study.
To evaluate performance, the authors benchmarked celltypeEnrich using scRNA-seq datasets from three tissues across two species. The manuscript compares annotation accuracy of celltypeEnrich to that of other tools, noting differences in accuracy, tissue coverage, and required parameter tuning. Specific details about the identities of the tissues, species, exact datasets, preprocessing steps, and comparator tools were reported in the preprint and its supplementary material; readers should consult the source for the full experimental setup.
Across the benchmark sets, celltypeEnrich achieved annotation accuracy in the range of 62–72%. According to the preprint, this performance generally exceeded that of other evaluated methods. The comparisons highlighted three areas where competing tools underperformed relative to celltypeEnrich: lower overall accuracy, incomplete tissue coverage from the available reference resources, and the need for manual parameter optimization to reach acceptable results.
The consensus-based strategy and the use of multiple reference datasets are presented as contributors to the improved accuracy and broader applicability observed for celltypeEnrich.
The authors evaluated the stability of celltypeEnrich when input gene lists were reduced in size. Down-sampling the marker lists to 25% of their original length produced stable performance for the tool, suggesting robustness to sparser marker information. This characteristic may be advantageous when cluster markers are limited in number or when differential expression yields small marker lists for some clusters.
celltypeEnrich is provided as an R Shiny web application and is freely accessible at https://celltypeenrich.gdcb.iastate.edu. The source code and related resources are hosted on GitHub at https://github.com/Tuteja-Lab/celltypeEnrich. The application is distributed under the MIT license for non-profit academic use, according to the preprint.
Users interested in reproducing the benchmarking or applying the tool to their own data should consult the project website and the GitHub repository for usage instructions, supported reference datasets, and any supplementary material provided with the preprint.
The authors declared no competing interests in the preprint. Funding support noted in the manuscript includes grants from the Eunice Kennedy Shriver National Institute of Child Health and Human Development; specific grant numbers and details are reported in the source. The preprint was posted to bioRxiv on September 21, 2026.