Rare viruses, vaccines, and more in over 1 trillion metagenomic reads

Wastewater sequencing as a tool for sensitively detecting viral sequences

Detection
Author

Chris Doering

Published

July 8, 2026

Thanks to Will Bradshaw, Kelly Chafin, Jo Faraguna, Jeff Kaufman, Ben Mueller, and Alessandro Zulli for their review on this post.

As SecureBio Detection has expanded and the CASPER wastewater sequencing network has grown alongside us, we are commonly asked if we’ve ever had any interesting detections. The short answer is yes! When you sequence sewage at a clip of 80+ billion metagenomic reads a week, rare, strange, and sometimes concerning sequences are bound to turn up. While we won’t share the full details of our capabilities or everything we’ve seen, the cases below show some of what our pipelines surface in practice. Most unusual signals are not emergencies that require response, but each one demonstrates that deep sequencing can surface rare and engineered sequences, distinguish benign sources from concerning ones, and flag the cases that warrant deeper investigation.

Finding Vaccines with Chimera Detection

One of our major internal flagging pipelines that processes incoming sequencing data is called Chimera Detection (CD). With CD, we search for instances of reads that match a virus attached to not-the-same-virus as a signal for potential engineered threats. A category of engineered viruses that this pipeline semi-regularly identifies which aren’t themselves concerning are vaccine candidates. Vaccine sequences can end up in our wastewater data for a number of benign reasons including proper disposal of experiments that doesn’t completely eliminate sequenceable nucleic acids or shedding by recently vaccinated individuals. A few months ago CD identified reads partially matching cytomegalovirus, also known as CMV.

For illustrative purposes, in the above figure we’re showing just a small subset of the vaccine-matching reads. All of the reads flagged by CD aligned to the vaccine sequence in virus/not virus border regions. Outside of the border regions, reads mapping to the vaccine sequence look completely like either (a) the closest related wild-type CMV strain or (b) non-viral molecular biology components. If a read landed outside of one of these border regions it wouldn’t look chimeric and thus wasn’t flagged by CD. Just over 7% of reads matching CMV from the sewershed in question were flagged by CD.

As discussed in a recent post and as demonstrated here, CD requires reads covering very specific parts of a virus or construct’s genome for the flagging system to be activated. However, CD’s ability to identify chimeric sequences, which we could then attribute to a vaccine sequence, was helpful to our overall response. Without searching for signs of engineering, we would not have been able to attribute to vaccine production what otherwise would have looked like a sudden, concerning spike in CMV cases.

While it’s not necessarily true that all reads coming from this sewershed are attributable to vaccine efforts, over the same period of time the vaccine-containing sewershed had 23x more reads matching CMV than all of our other sewersheds combined. This suggests that the vast majority of reads we observed were due to vaccine manufacturing, not infected individuals.

Virally-derived Elements Flagged by Chimera Detection

Another category of benign positives that CD regularly turns up is lab-related discards. Many constructs used in modern molecular biology research are virally-derived; lentiviral vectors are used to express genes in mammalian cells and IRES elements (many plucked from viruses) facilitate expression of proteins from synthetic constructs. Especially in areas with high concentrations of biotech and life science research, there are lots of opportunities for benign but virus-like sequences to end up in the wastewater. Even when samples are properly disposed of, there’s the possibility for enough fragments of nucleic acids to make it all the way through the sewer system to be detected with our deep sequencing techniques. We had a case of this happening recently with an expression vector that used a virally-derived IRES element.

The left hand side of the consensus sequence was a perfect match to the viral IRES element while the right hand side didn’t produce any significant hits among viral sequences. However, the read consensus was a high quality match to a patent sequence held by a biotechnology company with operations in one of our sewersheds. What our reads revealed was not an engineered virus nor even a vaccine candidate but benign lab discard of a protein expression construct.

Targeting Sources of Viral Fragments with Clades of Concern

Our other major flagging pipeline, which we call Clades of Concern (CoC), has a much simpler brief than Chimera Detection. There are some pathogens whose presence alone is enough to be concerning. CoC tracks a number of species including those on the CDC Biological Agents list and HHS and USDA Select Agents lists. Keeping tabs on when we see these various pathogens in our sewersheds has turned up some interesting reads. For instance, from one of our sites we observed reads matching to an agricultural pathogen not known to be present in the US.

All but one read was an exact match to a strain first isolated in the 1980s which, given the expected mutation rate of this virus, would be very surprising to observe for something actively circulating. While our sample had a number of reads mapping to this pathogen, the reads were all concentrated within a relatively small section of the genome; something we also would not expect if we were detecting whole live virus. Additionally, the sewershed in question both contains very limited agricultural land and a manufacturing site for a biotechnology company that would have reason to be working with the pathogen. With careful further monitoring in subsequent samples we identified a few chimeric reads between our identified pathogen and a vector construct, further confirming the benign source of the reads. This same fact pattern has occurred on a few occasions with other viruses that we could subsequently determine came from a benign source.

In all of the cases above, our goal is not to keep tabs on the goings on of biotech companies in our sampling areas. Our goal is to monitor for pandemic threats, particularly those that would evade other detection systems and where we can provide a unique deterrence benefit. Observing low-abundance or engineered constructs is validating, but it’s a byproduct and not our main focus.

Some detections can’t be directly attributed to a benign source and in those instances we alert the responsible entity for further investigation or response action. For example, earlier this year we sequenced a read matching an agricultural pathogen with no obvious lab origin or other explanation. We alerted appropriate partners for their awareness and, luckily, additional sequencing at the same site produced no further hits. Separately, we identified reads matching dengue virus in a sewershed where local transmission isn’t expected, most likely from a recently-returned traveler. Because of the specificity of metagenomic sequencing, when cases like these arise we can quickly pass on alerts to our partners with details beyond just the presence/absence of a virus.

With our deep, untargeted sequencing we are able to identify viruses of varying concern often with enough information to distinguish between benign cases and active infections. Together with our partners, we can alert relevant authorities in public health, the federal government, and other stakeholders to produce actionable results. Our sequencing efforts have already produced meaningful results and as our capabilities continue to grow we’re developing an even more robust early-warning system to safeguard against emerging biological threats.