Probe-based genomics methods are increasingly used to analyze molecular content in intact tissues and fixed cells, but approaches for encoding and decoding complex experimental conditions in cellular RNA remain limited. The authors report a comprehensive, modular framework that addresses these limitations by enabling the scalable construction and deployment of combinatorial DNA barcodes for use with probe-based genomics platforms.
The core of the reported approach is a combined software-and-reagent workflow. Custom software tools guide the design of barcode sequences and their assembly architectures, while purpose-built cloning reagents and optimized plasmids enable physical construction of barcode libraries. This integrated pipeline is intended to streamline the steps from in silico design through physical validation and downstream deployment.
Combinatorial barcodes in this framework are defined as spatially adjacent collections of known sequences. By arranging a set of defined sequence elements in combinatorial patterns, the system can generate a very large number of distinct barcode identities from a relatively small set of probes. This combinatorial strategy allows efficient discrimination of millions of unique molecules using probe-based readout methods.
To ensure accurate construction, the workflow pairs optimized assembly plasmids with whole-plasmid long-read sequencing for structural validation. Long-read sequencing of assembled plasmids provides high-fidelity confirmation of barcode architecture and sequence integrity across diverse constructs. The source emphasizes this combination as critical for reliable large-scale barcode library production and for identifying structural errors that could compromise decoding.
Once barcode libraries are assembled and validated, they are transferred into user-modified expression vectors to support different experimental applications. This transfer step is presented as flexible, enabling researchers to adapt the same assembled barcode repertoire for varied delivery formats, expression systems, or downstream assays as required by their probe-based genomics platform.
The authors demonstrate the framework by assembling two structurally distinct combinatorial barcode libraries. Each of these libraries contains millions of unique sequences, showcasing the scalability and architectural flexibility of the platform. The preprint summary does not provide additional protocol-level details, sequence designs, or exact library sizes beyond the description that each library contains millions of unique entries.
To validate functional performance in a biological context, the authors delivered a combinatorial barcode library using rabies virus to the mouse brain. They report successful in vivo decoding of a combinatorial barcode architecture via probe-based in situ sequencing, with the described capacity to distinguish approximately 16.3 million expressed RNAs. The abstract presents this as a demonstration of both the decoding capability of the barcode architecture and compatibility with probe-based in situ sequencing in tissue.
The framework is framed as filling a technically demanding niche by providing cost-effective, flexible, and accurate reagents for multiplexed experiments on current and evolving probe-based genomics platforms. Key practical advantages highlighted include modularity (software and reagents), structural validation via long-read sequencing, and the ability to port assembled libraries into different expression backbones.
Limitations in the source text: this preprint abstract does not report detailed sequencing metrics, error rates, per-barcode expression variability, experimental controls, or step-by-step protocols. As a preprint, the work has not been peer reviewed; readers should consult the full manuscript or contact the authors for methodological specifics and for performance benchmarks that are not included in the abstract.
The authors disclose funding sources that include One Mind Rising Star Award, an Alfred P. Sloan Foundation fellowship, a Simons Foundation ASD Genomics Grant, and NIH/BRAIN support (MH130464), among others. A competing interest statement notes that four authors are employees of Stellaromics, Inc.; the remaining authors declared no competing financial interests. The source document is a preprint posted on bioRxiv and explicitly states it has not undergone peer review.