Wastewater surveillance provides a community-level signal of infectious agents by sampling a pooled, composite matrix that includes contributions from symptomatic and asymptomatic individuals. During the COVID-19 pandemic, wastewater monitoring has been applied to detect and track SARS-CoV-2 and its evolving lineages. A major analytic challenge is resolving co-occurring variants and subvariants in complex wastewater samples where mutations are shared across lineages.
This study evaluated two bioinformatic strategies for identifying SARS-CoV-2 variants and subvariants in Oregon wastewater collected from February 7, 2021 through February 26, 2022. The objectives were to (1) compare the multilocus sequence typing (MLST) adaptation used previously for wastewater variant identification with the Freyja depth-weighted approach, and (2) correlate variant and subvariant relative abundances inferred from wastewater with clinical sequence data submitted to GISAID during the same period.
Composite 24-hour influent samples were collected more than once per week from up to 43 Oregon wastewater treatment facilities participating in the statewide SARS-CoV-2 wastewater surveillance program coordinated by the Oregon Health Authority and Oregon State University. Sample collection was voluntary with utilities’ permission. The study operated under an institutional determination that wastewater surveillance was exempt from human subjects regulations. Samples were vacuum-filtered onsite using 0.45-µm mixed cellulose ester electronegative filters as part of the processing workflow described in the publication.
The MLST-based adaptation for SARS-CoV-2 uses curated sets of lineage-defining mutations derived from clinical sequences. In wastewater samples, these specific mutations at multiple genomic loci are sought and matched to assign and quantify circulating variants. MLST has previously demonstrated strong correlations between wastewater-derived variant abundances and clinical surveillance but has limited ability to resolve nested subvariants in complex samples.
In contrast, the Freyja approach estimates relative abundances of variants and subvariants by leveraging a collection of marker mutations that represent nodes of the SARS-CoV-2 phylogeny. Freyja examines single-nucleotide variant frequencies across these markers and incorporates sequencing depth to weight the contribution of observed mutations. Using a depth-weighted, least absolute deviation statistical method, Freyja produces an abundance estimate for each variant and subvariant represented in its marker set. Prior evaluations cited in the study indicate improved detection accuracy and computational performance for Freyja relative to some other tools.
The study compared wastewater-derived relative abundances from both MLST and Freyja to clinical surveillance data available in GISAID for Oregon during the study period. Clinical SARS-CoV-2 sequences used for comparison were publicly available through GISAID EpiCoV and are referenced by an EPI_SET identifier in the source. Correlation analyses focused on overall variant agreement and, for Delta, on hierarchical subvariant levels and clades identified during the study window.
Both MLST and Freyja identified SARS-CoV-2 variants in wastewater at relative abundances that closely agreed with clinical surveillance data. A distinctive finding was Freyja’s ability to resolve subvariant structure: it identified over 200 Delta subvariants across three Delta clades (21A, 21I and 21J) and organized these into two Pango-based hierarchical levels (Level 1 and Level 2) for correlation with clinical data.
At the broader (Level 1) Delta subvariant grouping, Freyja-derived relative abundances showed strong agreement with clinical surveillance (reported Spearman correlations ranged approximately from rs = 0.892 to 0.944). Correlations at the finer (Level 2) Pango subvariant resolution were more inconsistent, with reported rs values spanning approximately rs = 0.324 to 0.903, indicating variable concordance for more granular subvariant assignments.
The MLST approach matched clinical relative abundances for major variants but lacked the ability to resolve the depth of Delta subvariant diversity detected by Freyja in wastewater samples.
The authors conclude that the Freyja bioinformatics approach provides enhanced resolution of SARS-CoV-2 variants and subvariants in wastewater relative to an MLST-based method, while producing relative abundances that agree with clinical surveillance. This added resolution is highlighted as a critical advantage for public health surveillance as the virus continues to evolve and share mutations across lineages.
All raw sequencing reads generated in the study are deposited in the NCBI Sequence Read Archive under Bioproject PRJNA938474. The custom Python notebook used to group variants, calculate relative abundances, and perform analyses is available in the Freyja_Correlation_Study GitHub repository and archived on Zenodo as documented in the article. Clinical comparison sequences were obtained from GISAID and are accessible via the cited EPI_SET identifier.
Limitations and notes reported in the source
The article emphasizes methodological differences between MLST and Freyja and reports variable correlation at fine-grain subvariant resolution (Level 2), indicating that while Freyja improves subvariant identification in wastewater, some subvariant-level assignments may show inconsistent agreement with clinical submissions. No additional limitations beyond those described in the source are invented here; readers should consult the full manuscript for detailed methods, figures, and supplementary analyses.