---
title: "SARS-CoV-2 variant resolution in Oregon wastewater using Freyja (Feb 2021–Feb 2022)"
id: "plos-one-19-correlation-assessment-of-sars-cov-2-variants-and-their-subvariants-present-in"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-19-correlation-assessment-of-sars-cov-2-variants-and-their-subvariants-present-in"
content_type: "clinical_feed_article"
specialty: "Infectious Disease"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357591"
published_at: "2026-09-09T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# SARS-CoV-2 variant resolution in Oregon wastewater using Freyja (Feb 2021–Feb 2022)
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-19-correlation-assessment-of-sars-cov-2-variants-and-their-subvariants-present-in
- **Specialty:** [Infectious Disease](https://medichelpline.com/clinical-feed/infectious-disease.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357591)
- **Published At:** 2026-09-09T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- Wastewater surveillance captures SARS-CoV-2 signal from symptomatic and asymptomatic individuals and can track variant distributions at a community level. - Two bioinformatic approaches were compared for identifying SARS-CoV-2 variants in Oregon wastewater collected from February 7, 2021 to February 26, 2022: a multilocus sequence typing (**MLST**) adaptation and the **Freyja** depth-weighted approach. - Samples comprised 24-hour composite influent wastewater collected more than once per week from up to **43** Oregon wastewater treatment facilities participating in the statewide surveillance program. - The MLST approach uses curated lineage-defining mutations from clinical sequences matched to mutations detected at multiple genomic loci in wastewater; it previously showed strong correlations with clinical relative abundances but lacked subvariant resolution in complex samples. - The Freyja approach uses genome-wide marker mutation profiles and sequencing depth with a depth-weighted least absolute deviation method to estimate relative abundances of variants and **subvariants**. - Both MLST and Freyja identified SARS-CoV-2 variants in wastewater at relative abundances that closely agreed with clinical surveillance data from GISAID. - Freyja uniquely resolved more than 200 Delta subvariants during the study period, organized into three clades (21A, 21I and 21J) and two hierarchical Pango-based levels; Level 1 Delta subvariants showed strong correlations with clinical data (rs = 0.892–0.944), while Level 2 correlations were more variable (rs = 0.324–0.903). - The authors conclude that Freyja provides enhanced resolution of variants and subvariants in wastewater, an advantage for public health surveillance as SARS-CoV-2 continues to evolve and share mutations across lineages. - All raw sequencing reads are deposited under Bioproject PRJNA938474; analysis code and notebooks are available in the Freyja_Correlation_Study GitHub repository and archived on Zenodo. - Funding came from a CDC cooperative agreement; authors declared no competing interests.
## Clinical Analysis & Structured Key Points
Correlation assessment of SARS-CoV-2 variants and their subvariants present in clinical and wastewater samples in Oregon, USA (February 7, 2021 - February 26, 2022) using the Freyja bioinformatics approach | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Reader Comments Figures Figures Abstract Background Wastewater surveillance is a valuable tool for monitoring SARS-CoV-2 at the community level. As the virus diversified into many variants and subvariants that share overlapping mutations, resolving them accurately from wastewater becomes a key bioinformatic challenge. Objectives and aims This study evaluated two distinct bioinformatic approaches, multilocus sequence typing (MLST) and Freyja, for identifying SARS-CoV-2 variants and subvariants in Oregon wastewater samples collected from February 2021 to February 2022. Methods The MLST approach identified SARS-CoV-2 variants using unique mutations curated from clinical samples. In contrast, the Freyja approach resolved variant and subvariant abundances using genome wide mutation profiles weighted by sequencing depth. In this study, the variant and subvariants relative abundances produced by both approaches were compared against those observed in clinical surveillance data. Results Both approaches identified SARS-CoV-2 variants at relative abundances that agreed closely with those observed in clinical surveillance data. However, only the Freyja approach identified over 200 Delta subvariants, divided into three clades (21A, 21I and 21J) and two levels (Level 1 and 2) based on Pango subvariants. Delta subvariants showed strong agreement at Level 1 subvariants (r s = 0.892–0.944), while agreement at Level 2 subvariants was inconsistent (r s = 0.324–0.903). Conclusions The Freyja approach provided enhanced resolution of SARS-CoV-2 variants and subvariants in wastewater, at abundances that agreed with clinical surveillance. This added resolution is a critical advantage for public health surveillance as SARS-CoV-2 continues to evolve and share mutations across variants and subvariants. Citation: Bhatia A, Elser J, Carrell SJ, Kronmiller B, Sutton M, Kelly C, et al. (2026) Correlation assessment of SARS-CoV-2 variants and their subvariants present in clinical and wastewater samples in Oregon, USA (February 7, 2021 - February 26, 2022) using the Freyja bioinformatics approach. PLoS One 21(9): e0357591. https://doi.org/10.1371/journal.pone.0357591 Editor: Nagarajan Raju, Emory University, UNITED STATES OF AMERICA Received: June 15, 2026; Accepted: August 18, 2026; Published: September 9, 2026 Copyright: © 2026 Bhatia et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: All raw sequencing reads generated in this study are deposited in the NCBI Sequence Read Archive under Bioproject: PRJNA938474. The custom Python notebook used to group variants and subvariants and to calculate relative abundances is publicly available at the GitHub repository Freyja_Correlation_Study ( https://github.com/AnirudhB7/Freyja_Correlation_Study) and is permanently archived in the Zenodo repository under DOI: https://doi.org/10.5281/zenodo.21384161 . Clinical SARS-CoV-2 sequence data used for comparison were obtained from the GISAID EpiCoV database and are available under EPI_SET_260715qr ( https://doi.org/10.55876/gis8.260715qr) . The data can be retrieved directly from GISAID using the EPI_SET identifier. Funding: Funding for this work came from the Centers for Disease Control and Prevention (cooperative agreement no. CK-19-1904). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing interests: The authors have declared that no competing interests exist. Introduction Wastewater comprises a complex blend of diverse biomarkers that provide insights into a wide range of human activities. Wastewater surveillance has been used to monitor the use of illicit drugs and other substances by the community and assess the communal burden of pathogens, including polio, SARS-CoV-2, influenza, respiratory syncytial virus and measles [ 1 – 5 ]. Wastewater surveillance has the benefit of encompassing both symptomatic and asymptomatic individuals and includes those individuals who have not undergone testing [ 6 ]. Additionally, wastewater surveillance can identify the distributions of SARS-CoV-2 variants within a community [ 7 , 8 ]. To monitor the rapid acquisition of mutations in the SARS-CoV-2 virus during the COVID-19 pandemic and investigate the impact of genetic variants on disease severity, bioinformatic approaches that include variant calling tools were used to identify low-frequency mutations found in raw sequence data from wastewater samples [ 7 – 11 ]. One of these bioinformatic approaches was the multilocus sequence typing (MLST) approach. The MLST approach was originally used to identify bacterial strains by sequencing multiple housekeeping genes [ 12 ]. At each gene locus, all unique sequences identified across a bacterial species are assigned distinct allele numbers, and the combination of allele numbers across all loci defines a sequence type (ST) that serves as a standardized identifier for strain classification and epidemiological tracking [ 12 – 14 ]. When adapted for SARS-CoV-2 genomic surveillance in wastewater, rather than housekeeping genes, the MLST approach sets lineage-defining mutations identified from clinical sequences are matched to mutations detected at multiple genomic loci in wastewater samples to identify and quantify circulating variants [ 11 , 15 , 16 ]. In addition to identification, the MLST approach has demonstrated strong correlations between SARS-CoV-2 variant relative abundances in wastewater and clinical data [ 15 ]. However, the MLST approach was unable to identify subsequent subvariants within SARS-CoV-2 lineages in complex wastewater samples. In contrast, the Freyja approach can estimate the relative abundances of SARS-CoV-2 variants and their subvariants in wastewater samples by utilizing a collection of unique markers representing different lineages of the SARS-CoV-2 phylogenetic tree [ 17 – 19 ]. This approach examines the frequency of single nucleotide variants (SNVs) within each unique marker, representing a specific mutation in a particular lineage. Furthermore, Freyja considers how often distinct parts of the genomes have been sequenced ( i.e. , sequencing depth). By taking both sequencing depth and the presence of SNVs into account and by utilizing depth-weighted least absolute deviation-based statistical methods, Freyja achieves an estimation of the relative abundance of each variant and subsequent subvariants in a sample [ 20 ]. Additionally, the Freyja approach has been shown to have superior performance relative to other variant identification tools, with improvements in detection accuracy and computational efficiency [ 9 , 17 , 20 – 23 ]. However, the extent to which SARS-CoV-2 subvariants identified in wastewater by the Freyja approach correspond to those identified in GISAID clinical submissions has yet to be thoroughly examined [ 24 ]. In this study, the relative abundances of SARS-CoV-2 variants in Oregon wastewater identified with the Freyja approach were correlated with the relative abundances of SARS-CoV-2 variants in Oregon wastewater previously identified by the MLST approach [ 15 ]. Additionally, the relative abundances of SARS-CoV-2 variants and subsequent subvariants found in Oregon wastewater samples with the Freyja approach were correlated with the relative abundances of SARS-CoV-2 variants and subsequent subvariants identified in Oregon clinical samples submitted to GISAID. The subvariant correlation focused on all clades of Delta variants present during the study period from February 7, 2021 – February 26, 2022. Materials and methods Wastewater collection, concentration and RNA extraction From February 7, 2021, through February 26, 2022, 24-hour composite samples were collected > 1 time per week from the influents of up to 43 Oregon-based wastewater treatment facilities ( Fig 1 ). Wastewater samples were collected from municipal wastewater treatment facilities participating in the Oregon statewide SARS-CoV-2 wastewater surveillance program, a public health surveillance effort coordinated by the Oregon Health Authority and Oregon State University. Sample collection was conducted with the knowledge and permission of each participating utility as part of their voluntary participation in this program. No specific permits were required for this study. The Oregon State University Institutional Biosafety Committee determined wastewater surveillance is exempt from the human subjects regulations set forth by the U.S. Department of Health and Human Services (45 CFR 46). All the clinical sequencing data were publicly available on GISAID. Download: PNG larger image TIFF original image Fig 1. Workflow for comparative relative abundance of SARS-CoV-2 variants/subvariants in wastewater and clinical samples. https://doi.org/10.1371/journal.pone.0357591.g001 The samples were vacuum-filtered onsite using a 0.45-µm, 47-mm diameter mixed cellulose ester electronegative filter which was subsequently placed into 1 mL of DNA/RNA Shield (Zymo Scientific, USA) for RNA stabilization. The stabilized filters underwent bead beating with 0.7-mm garnet beads for 2 minutes, and RNA was extracted using the MagMAX Viral/Pathogen Nucleic Acid Isolation Kit (Thermo Fisher Scientific). SARS-CoV-2 RNA concentrations were quantified using reverse transcription droplet digital PCR (RT-ddPCR) on a QX200 ddPCR system (Bio-Rad Laboratories, Hercules, CA) using the 2019-nCoV CDC ddPCR Triplex Probe Assay and One-Step RT-ddPCR Advanced Kit protocol [ 11 , 15 ]. Amplicon libraries were generated using the Swift Amplicon SARS-CoV-2 Panel and Swift Amplicon Combinatorial Dual indexed adapters (Integrated DNA Technologies [IDT] Swift Biosciences). Sequence reads from 2161 samples were preprocessed and demultiplexed with zero index mismatches using bcl2fastq2 version 2.20 for HiSeq 3000 and BCL Convert version 1.2.1 for NextSeq 2000. Reads were trimmed using the BBDuk tool (BBMap version 38.84 (US Department of Energy Joint Genome Institute, https://jgi.doe.gov )) and aligned to reference sequence (Wuhan-Hu-1, GenBank accession no. NC_045512.2) using the BWA-MEM algorithm version 0.7.17-r1188 ( https://github.com/lh3/bwa ). The reads were coordinate sorted using SAMtools version 1.10, and the primer sequences were removed using Primerclip version 0.3.8 [ 25 , 26 ]. SAMtools were also used to convert SAM files to BAM files and to coordinate and sort the BAM files [ 25 ]. The sequencing reads were submitted to NCBI and are present under Bioproject: PRJNA938474 [ 27 ]. Identification of SARS-CoV-2 variants using the Multilocus Sequence Typing (MLST) approach The Genome Analysis Toolkit (GATK) was utilized to identify mutations in the sequence reads by comparing the reads against the Wuhan-Hu-1 reference genome [ 28 ]. Within the GATK toolkit, HaplotypeCaller was used to call variants on a per-sample basis, followed by joint genotyping across all samples using CombineGVCFs and GenotypeGVCFs, and final conversion of the resulting VCF file into a tabular format using VariantsToTable for downstream MLST-based analysis [ 11 , 29 ]. The Integrative Genomics Viewer (IGV) was used to manually inspect sequence alignments and confirm mutation calls, ensuring accuracy in mutation detection [ 30 ]. The MLST approach was applied to classify the mutations into known SARS-CoV-2 variants. To screen specific mutations unique to each variant, the mutations found in the wastewater samples were compared against a comprehensive database of mutations from clinical specimens and published variant data. The final identification of a variant in a wastewater sample was based on two criteria: 1) a minimum threshold of 5% of sequence reads with at least 6 total reads covering the mutation site, and 2) at least 2 different sites carrying a variant mutation must be present [ 11 , 16 ]. Identification of SARS-CoV-2 variants and subvariants using the Freyja approach The same sequence reads that were classified by the MLST approach were also classified using the Freyja approach. Applying both approaches to identical sequencing reads ensured that any differences in the variants and relative abundances reported were attributable to the bioinformatic approach rather than to variation in sample processing or sequencing. The Freyja approach identified SARS-CoV-2 variants by linking variants to site-specific SNPs, which act as unique mutation signatures, derived from a global phylogenetic tree created by UShER [ 18 ]. The Freyja approach utilized a depth-weighted least absolute deviation regression approach to deconvolve the relative abundances of identified variants [ 20 ]. This method emphasized data from regions of the genome with higher sequencing depth to provide more reliable mutation frequency estimates [ 17 ]. The analysis was constrained so that the sum of all lineage abundances equals one and each abundance value must be non-negative [ 23 , 31 , 32 ]. After identification and relative abundance calculations were completed, the variants and subvariants were classified as SARS-CoV-2 variants and subvariants according to Pango nomenclature [ 33 ]. Variant calling and demixing were performed using Freyja version 1.4.5 with the UShER barcode library dated 31 July 2023. Comparison of variant relative abundances identified with the MLST and Freyja approaches For consistency in labeling data across genomic surveillance studies and to facilitate comparisons between the MLST and Freyja protocols, an in-house Python notebook ( https://github.com/AnirudhB7/Freyja_Correlation_Study ) was used to group the variants and subvariants, as defined by the World Health Organization nomenclature, using a Pango lineage to WHO variant mapping retrieved in January 2024 from the Nextclade SARS-CoV-2 dataset ( https://nextstrain.org/nextclade/nextstrain/sars-cov-2/wuhan-hu-1/orfs ) [ 34 , 35 ]. The mapping file used is available in the code repository. For instance, variants and subvariants identified as BA.2 and BA.2.86, using Pango nomenclature, were grouped under the Omicron variant as named using WHO classification system. Once grouped, the weekly (based on epi weeks) state-wide variant relative abundances using either GISAID ( https://www.gisaid.org/ ) ( Equation 1 ) or wastewater ( Equation 2 ) data were calculated [ 24 ]. (1) = Statewide relative abundance of a variant or subvariant in clinical samples, = Number of clinical samples belonging to a specific variant (or subvariant) reported for each epiweek, = Total number of clinical sequence samples reported for that epiweek (2) = Statewide relative abundance of a variant or subvariant in wastewater, for each variant in each location, = Flow rate of that location on a given day, reported for each variant on a given day = , = Total gene copies for all variants recorded on a given day = , = all variants in the sample, = all locations included in a particular epiweek, = wastewater SARS-CoV-2 variants concentration at location j, = wastewater flow rate for location j. The relative abundances of the SARS-CoV-2 variants from the MLST and Freyja approaches were organized into spreadsheet-like data frames using Python’s Pandas package ( https://pandas.pydata.org ). Additionally, the variant relative abundances identified in clinical and wastewater samples by the Freyja approach were also organized into a separate data frame. This organization facilitated a clear comparison between the two datasets and streamlined subsequent statistical analyses. All statistical analyses were conducted using Python’s SciPy package ( https://scipy.org ). Normality of the data was assessed via the Shapiro-Wilk test using the ‘shapiro’ function from the SciPy package. With the normality test indicating that the data did not follow normal distribution, Spearman’s rank correlation analyses were conducted on the data sets using the ‘spearmanr’ function in the SciPy package. Since weekly observations within an epidemic wave are serially dependent, so the number of effectively independent observations is smaller than the number of weeks sampled and conventional significance tests for rank correlation are not appropriate. Nominal p-values were therefore not used. Uncertainty in each correlation coefficient was instead quantified using a block bootstrap, an approach developed for serially correlated data [ 36 ], implemented with the CircularBlockBootstrap function of the Python arch package ( https://arch.readthedocs.io/en/latest/bootstrap/timeseries-bootstraps.html ). Blocks of four consecutive weeks were used, following the n^(1/3) rate for block length selection with n = 55 weekly observations [ 36 ]. Each 95% confidence interval is based on 2,000 replicates with a fixed random seed. Confidence intervals that exclude zero indicate agreement between the two measures, with the lower bound giving a conservative estimate of its strength. Intervals that include zero indicate that the data are consistent with no association and are not interpreted as evidence of agreement. Time-series plots comparing the MLST and Freyja approaches as well as the variant trends in wastewater and clinical samples were built using Python’s matplotlib package ( https://matplotlib.org ). Correlation panels include an identity line (y = x) as a reference for agreement between the two measures, with both axes on a common scale within each panel. Correlation of Delta ‘AY’ lineage variant and subvariant relative abundances identified in wastewater and clinical cases The relative abundance data of Delta ‘AY’ variants ( e.g. , AY.26) and subvariants ( e.g. , AY.26.1) found in wastewater were identified using the Freyja approach and clustered according to Pango nomenclature. These classifications were mapped onto Nextstrain nomenclature system and clades were identified [ 37 ]. Metadata for all publicly available Delta ‘AY’ variants and subvariants were retrieved from the Nextstrain database, which provided information about their Pango nomenclature and the corresponding clade categories assigned to them [ 35 ]. Subvariants within each Delta clade were categorized based on the number of periods in their Pango nomenclature. For example, SARS-CoV-2 variants AY.16 and AY.16.1, as identified on the NextStrain database, are part of clade 21A. In this instance, the relative abundances of AY.16 and AY.16.1 in wastewater samples were aggregated under the Delta_21A category. For finer resolution, AY.26 was distinguished from its sub lineage AY.26.1 based on the PANGO lineage hierarchy and AY.16 was specifically included in Delta_21A_Level_1 and AY.16.1 in Delta_21A_Level_2. This classification approach is consistently applied to clades 21I and 21J ( Fig 2 ). Download: PNG larger image TIFF original image Fig 2. Classification of Delta variants and subvariants categorized using Nextstrain clade nomenclature. https://doi.org/10.1371/journal.pone.0357591.g002 Similarly, the subvariants belonging to each Delta clade in GISAID clinical samples were categorized by downloading data from the Nextstrain database, which contained data related to the variants/subvariant classifications ( e.g. , their Pango nomenclature), their WHO lineage ( e.g. , Delta) and their corresponding clade name ( e.g. , 21A, 21I and 21J). The GISAID clinical samples were filtered by variant to isolate Delta-containing samples. The identified Delta variants/su
## Related Clinical Research

- [Hypothesis: Airborne Transmission of COVID-19 from Asymptomatic Individuals by Speaking](https://medichelpline.com/clinical-feed/pubmed-41320299.md) (DOI: 10.7883/yoken.JJID.2025.169)
- [Macrolide-Resistant Bordetella pertussis Detected in Peru During 2025 Outbreak](https://medichelpline.com/clinical-feed/cdc-emerging-infectious-diseases-journal-2-emergence-of-macrolide-resistant-bordetella-pertussis-peru-2025.md)
- [Pennsylvania reports third measles-related death amid ongoing outbreak; other health headlines](https://medichelpline.com/clinical-feed/stat-news-2-pennsylvania-records-third-measles-related-death-amid-outbreak.md)
- [Factors Associated with Influenza Vaccine Uptake Among Georgian Healthcare Workers, 2021–2024](https://medichelpline.com/clinical-feed/plos-one-23-factors-associated-with-influenza-vaccine-uptake-among-healthcare-workers-in.md)
- [Increased risk of acute gastroenteritis after COVID-19 hospitalisation: a retrospective England an](https://medichelpline.com/clinical-feed/medrxiv-3-risk-of-acute-gastroenteritis-following-covid-19-exposure-a-retrospective.md)

## Navigation
- [← Back to Infectious Disease Feed](https://medichelpline.com/clinical-feed/infectious-disease.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.