This pilot study examined whether circulating microbial DNA (cmDNA) in plasma can serve as a cancer biomarker, using a tightly controlled design to separate true biological signal from technical and analytic artefacts. The authors profiled plasma cell-free DNA (cfDNA) and buffy-coat genomic DNA (gDNA) from two patients with metastatic castration-resistant prostate cancer and two healthy volunteers. Parallel mock blood-draw and reagent controls were included for each condition, and experiments were performed both with and without host-DNA depletion.
Specimen-matched negative controls and reagent-only controls were explicitly incorporated to assess contamination and method-driven signal in this low-biomass context.
Sequencing reads were classified using the k-mer–based pipeline Kraken2 with Bracken abundance estimation, and independently with the marker-gene classifier MetaPhlAn. To probe the contribution of sequence composition and classifier/database architecture to apparent taxonomic signal, the authors applied two informatics controls:
Both shuffled and synthetic reads were passed through the same classifiers, allowing direct comparison between real biological sequences and composition-preserving or composition-matched nulls.
Kraken2 reported several thousand genera across 40 samples. Principal-coordinate analysis (PCoA) showed samples clustering by specimen type, and pooled genus counts from the study correlated strongly with a published cancer-microbiome catalog (The Cancer Genome Atlas lung adenocarcinoma, TCGA-LUAD), with Spearman ρ = 0.81 across 282 shared genera. These patterns, taken at face value, could be interpreted as biologically meaningful cmDNA structure that aligns with cancer-associated microbial profiles reported in public datasets.
However, the same classifier behavior and cross-cohort agreement were observed when classifying per-base shuffled reads: shuffled reads, which lack any biological sequence information but retain length and GC composition, still produced abundant classifications, clustered by specimen type, and correlated with the TCGA catalog (ρ ≈ 0.7). Purely synthetic reads matched to an aggregate GC target reproduced much of the cross-cohort agreement as well (synthetic TCGA-LUAD ρ = 0.61 versus 0.81 for real reads). The synthetic read–based cross-cohort signal was significant in 27 of 33 TCGA cancer cohorts examined.
Two systematic dependencies explained a large fraction of the apparent taxonomic signal:
Genus-level counts scaled tightly with each genus's k-mer representation in the Kraken2 reference database. Regression showed strong dependence on database k-mer content for real reads (r2 = 0.74) and for shuffled reads (r2 = 0.85).
GC composition of reads influenced classification outcomes. The per-base shuffling control preserved GC content but eliminated biological sequence; nonetheless classifications persisted, indicating the classifiers and databases assign taxa based on compositional features correlated with GC and k-mer prevalence rather than genuine microbial origin.
A direct demonstration of false positives came from replicate shuffles of a low-GC versus a high-GC plasma sample: although the true biological difference is zero after shuffling, these replicates produced spurious significant differences in about 46% of genera.
To separate plausible biological signal from the composition-driven floor, the authors regressed observed genus counts against the shuffled baseline and assessed residuals. After this adjustment, 23 genera remained above the artifact floor at a 5% false discovery rate (FDR). Nearly all of these genera were either known kit contaminants, viruses enriched in control samples, or extremely low-abundance taxa, highlighting residual contamination and classification noise even after regression.
Applying a four-criterion validity filter (details not fully reproduced in this summary) reduced thousands of Kraken2-reported genera to a single defensible candidate: Klebsiella. The authors emphasize that their small sample set cannot prove absence of authentic cmDNA, but that only one genus passed conservative validity criteria in this controlled pilot.
The study concludes that much of the apparent cmDNA structure reported by short-read, k-mer–based taxonomic pipelines is explained by base composition (GC) and reference-database k-mer architecture, rather than authentic circulating microbial biology. Short-read k-mer pipelines on their own cannot reliably separate composition-driven artefacts from true microbial signal in low-biomass samples.
To limit false discovery when studying low-biomass metagenomic samples, the authors recommend the following measures, all supported by their controlled evaluations:
The authors report no competing interests and provide a code repository linked in the original article for reproducibility. They note that, despite the controlled design, their small sample size cannot exclude the existence of authentic cmDNA signals in larger or different cohorts, but it does show how readily composition and database architecture can produce misleading results.
Researchers planning low-biomass metagenomic studies for biomarker discovery should expect substantial risk of false-positive taxonomic signal driven by read composition and classifier/database biases. Robust experimental controls and compositional nulls should be standard practice before claiming biologically meaningful circulating microbial signatures in cancer or other diseases.