The study performed paired long-read RNA sequencing (lrRNA-seq) and short-read RNA-seq (srRNA-seq) on whole blood samples from 20 individuals with rare diseases and their unaffected biological parents. On average the lrRNA-seq data comprised 13.4 million full-length non-chimeric reads per sample. The cohort and sequencing strategy were designed to compare transcriptome coverage, isoform discovery, fusion transcript detection, and the ability to detect splicing outliers between lrRNA-seq and srRNA-seq in a heterogeneous rare disease setting.
The authors report that lrRNA-seq provides more uniform coverage across transcripts compared with srRNA-seq. A notable difference in transcript length capture was observed: 20.2% of long-read transcripts exceeded 10 kb in length, whereas fewer than 5% of transcripts from paired srRNA-seq reached that length. This indicates lrRNA-seq's advantage in resolving full-length and long transcripts that are underrepresented or fragmented in short-read data. Analyses included comparisons overall and focused on known disease-associated (DA) genes to assess coverage in clinically relevant loci.
From lrRNA-seq the study identified a mean of 24,439 isoforms per sample. Of these isoforms, 18.5% were unannotated in GENCODE, indicating a substantial proportion of transcript diversity not captured in current annotations. Importantly, 74.3% of unannotated isoforms were located in genes previously classified as disease-associated, highlighting the potential clinical relevance of previously unreported isoforms. The authors emphasize that lrRNA-seq can reveal isoform-level complexity that may inform variant interpretation and functional studies in rare disease cases.
The lrRNA-seq datasets contained a mean of 13 unique fusion transcripts per sample. All detected fusion transcripts were intrachromosomal. The authors compared fusion transcript calls to paired long-read DNA sequencing and did not find supporting structural variants that would indicate a genomic rearrangement cause. They interpret these findings as likely reflecting stochastic transcriptional read-through to adjacent genes rather than bona fide genomic fusion events. The absence of corresponding DNA structural variants suggests caution in interpreting fusion transcripts from RNA data alone.
The study highlights a single individual with a molecular diagnosis of ReNU syndrome due to a de novo RNU4-2 variant, a disorder of the major spliceosome. In this case, lrRNA-seq revealed a transcriptome-wide spliceopathy pattern characterized by variation at 5′ splice sites, a pattern that paired srRNA-seq did not detect. This example illustrates lrRNA-seq's potential to identify global splicing aberrations arising from spliceosomal defects, providing diagnostic and mechanistic insight not visible with short-read approaches.
All data generated in this study are available via dbGaP under accession number phs003047. Ethical approval for the work was provided by the IRB of Mass General Brigham under protocol 2016P001422, and the authors confirm informed consent and compliance with research reporting guidelines. Competing interests reported include that one author received research support from Pacific Biosciences for this project, and another author has served on the Scientific Advisory Board of an industry company. Funding sources declared include grants from the National Human Genome Research Institute, the Chan Zuckerberg Initiative, and the National Institute of Arthritis and Musculoskeletal and Skin Diseases.
The paired lrRNA-seq and srRNA-seq cohort establishes a resource to evaluate the utility of long-read transcriptomics in rare disease. The data demonstrate lrRNA-seq's strengths: improved full-length transcript recovery, greater detection of long transcripts, discovery of many isoforms including unannotated transcripts in disease-associated genes, and the ability to detect transcriptome-wide splicing defects in at least one spliceosome disorder. Limitations and interpretive challenges are also evident: fusion transcripts detected by lrRNA-seq lacked DNA structural support and may reflect transcriptional read-through rather than pathogenic rearrangements; srRNA-seq missed certain splicing patterns identified by lrRNA-seq; and practical considerations for clinical implementation—such as scalability, cost, and analytic pipelines—remain. The authors position lrRNA-seq as a complementary tool that can expand transcriptomic resolution in rare disease research and potentially in diagnostics, while acknowledging current technical and interpretive hurdles.