Field isolates of Plasmodium falciparum show substantial genomic variation, including sequence divergence and copy number variation, that is not fully represented by the laboratory-derived Pf3D7 reference genome. The degree to which the choice of reference genome influences mapping of RNA-seq reads and downstream expression inferences in natural infections has been unclear. This study generated a sample-matched genome assembly from an infected human carrier and paired it with single-cell RNA sequencing (scRNAseq) data to directly compare expression inference when aligning to the conventional Pf3D7 reference versus a matched assembly.
The authors produced a new isolate-specific assembly, designated ML52, from a P. falciparum-infected carrier in Mali. Annotation of the ML52 assembly was performed using the Companion annotation pipeline with Pf3D7 supplied as the reference guide. The creation of ML52 provided a sample-matched genomic scaffold to test whether a matched reference improves mapping and expression inference for scRNAseq from a natural infection.
scRNAseq reads derived from the natural infection were separately aligned to the canonical Pf3D7 genome and to the sample-matched ML52 assembly. The authors performed locus-level inspection of alignments to assess concordance and discrepancies in expression inference attributable to the choice of reference. This approach allowed identification of mapping artefacts, reads aligning to unplaced contigs, and loci where divergence between the sample and Pf3D7 influenced read placement and gene-level expression calls.
For the majority of conserved genes examined, expression inferences were concordant whether reads were aligned to Pf3D7 or ML52, indicating that scRNAseq expression estimation is generally robust for conserved coding loci even when using an unmatched laboratory reference. Many multigene family loci also showed concordant expression across both references, suggesting that for several gene families existing reference sequences are sufficient to capture expression patterns in single-cell data from natural infections.
Where expression differed between alignments to Pf3D7 and ML52, the authors investigated underlying causes. Some discrepancies were attributable to mapping artefacts inherent to short read alignment against divergent sequences. Other differences arose from reads mapping to unplaced contigs in the assembly; these contigs reflected divergent haplotypes consistent with co-infecting strains present in the natural infection sample. The presence of multiple haplotypes in the infection therefore contributed to complex alignment behaviors and to locus-specific inconsistencies when using a single laboratory reference.
The most pronounced reference-dependent differences were observed for var genes, the extremely polymorphic antigenic loci in P. falciparum. The study found considerable mis-mapping of var reads when aligned to Pf3D7: reads either incorrectly aligned to 3D7 var loci or did not map to the Pf3D7 reference at all, with approximately two-thirds of var reads failing to map to Pf3D7. These results indicate that var gene expression cannot be reliably inferred from scRNAseq data unless a sample-matched assembly that captures the specific var repertoires present in the infection is available.
The authors conclude that scRNAseq expression inference in Plasmodium falciparum is robust for conserved genes and for most multigene families when comparing results obtained with a matched versus an unmatched reference genome. However, for the highly variable var genes, a sample-matched assembly is required to evaluate expression accurately. This distinction underscores the importance of considering reference genome choice when interpreting transcriptomic analyses, particularly in organisms that possess highly variable antigenic loci or where co-infection and haplotype diversity are expected.
Note: This summary is based on the information reported in the source preprint. Details such as exact assembly statistics, sample collection specifics, sequencing depth, or quantitative mapping metrics were not reported in the supplied source text and therefore are not included here.