This preprint reports an empirical evaluation of statistical power in single-cell differential-expression studies. The authors used sex-biased differential expression as a real-world contrast to derive power estimates grounded in empirical data rather than solely simulations or small pilot samples. The goal was to quantify how commonly used experimental setups perform when seeking to detect small-effect genes in brain cell types.
The analysis used data from 1,494 donors and focused on three brain cell types. The sex-biased differential-expression framework served as the biological contrast to measure detectable effects across donors and cell populations. The authors made their data-processing and analysis code available via a public GitHub repository linked in the source.
Using the large donor set, the authors derived power estimates for a variety of experimental configurations. They found a substantial lack of power for reliably detecting small-effect genes — effects that are commonly reported in single-cell studies — even in experiments containing as many as 600 donors. This indicates that increasing donor numbers alone, under typical single-cell sequencing conditions, may still leave many small but biologically relevant expression differences undetected.
The study demonstrates that sequencing characteristics and per-cell sampling critically influence discovery. When the astrocyte dataset was reduced to mimic the poorer sequencing characteristics and lower cell counts typical of microglia, more than half of the differentially expressed genes (DEGs) that were detected in the full astrocyte set were no longer identified. This result highlights how lower sequencing depth and fewer cells per sample can substantially erode detectable signal and the resulting gene lists.
Even when using standard multiple-testing corrections and adjusted p-value thresholds, the authors observed limited reproducibility among detected genes. They report that only the top quartile of significant genes were reproducible across empirical comparisons. This finding suggests that commonly applied significance thresholds may allow many false positives or context-specific findings, and that reported DEG lists should be interpreted with caution unless supported by stronger evidence or replication.
The authors fitted a predictive model to empirical outcomes to identify variables that most strongly determine the probability of detecting differential expression. The model identified cell count and gene expression level as strong determinants of power. In other words, genes that are more highly expressed and cell types with larger per-sample cell counts are more likely to yield reproducible differential-expression signals under the examined experimental conditions.
Based on these empirical findings, the authors advocate several design and analysis changes for future single-cell differential-expression studies:
These recommendations are presented as pragmatic responses to the empirical observation that many current single-cell study designs are underpowered to detect small-effect genes reliably.
The authors provided a link to a public repository containing code and resources used in the analysis. The preprint states that the authors have declared no competing interests. The manuscript was posted to the preprint server on September 19, 2026.
All conclusions and recommendations in this summary are drawn from the results reported in the source preprint. The source does not provide every methodological detail in this summary; readers wishing to evaluate specific modeling choices, downsampling methods, or threshold definitions should consult the original preprint and the linked repository for full technical details.