Database of virus genomes from ultra-deep sequencing of wastewater (WVDB)
Authors/Creators
- Kantor, Rose (Data curator)1
- Shakya, Migun (Data curator)2
- Ruth, Nelson (Data curator)2
- Rothman, Jason (Data collector)3
- Rushford, Clayton (Data collector)4
- Gregory, Devon (Project member)4
- Epstein, Aidan (Project member)1
- Kaufman, Jeff (Sponsor)5
- Allen, Jonathan (Supervisor)1
- Chain, Patrick (Supervisor)2
- O'Connor, David (Project member)6
- Johnson, Marc (Data collector)4
Description
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost.
The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Files
Files
(192.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:a64991d6c97e684521d1d07f4a16d592
|
82.9 MB | Download |
|
md5:5d29f3ad0770c46cec73b0ce54447e88
|
929.1 kB | Download |
|
md5:c2c6fe51c91065a40f88ca84c5749846
|
98.1 MB | Download |
|
md5:ff9ff5f3cc341af99a21301541f30c2d
|
10.8 MB | Download |
Additional details
Dates
- Available
-
2026-04-30