Published August 24, 2022 | Version 1.0.0

Test data for RonaQC - mapped SARS-CoV-2 reads

Authors/Creators

  • 1. Quadram Institute Bioscience

Description

This dataset includes test data for RonaQC

RonaQC accepts mapped SARS-CoV-2 reads (BAM format), generated from the SARS-CoV-2 bioinformatic pipelines like ARTIC, and any control samples from the respective sequencing run (negative/positive) as input. It will then assess the levels of cross contamination and primer contamination in the samples, and determine if the samples are reliable for detecting SARS-CoV-2, phylogenetic analysis, and/or submission to public databases.


The dataset includes SARS-CoV-2 sequenced reads compiled by CDCgov/datasets-sars-cov-2 [1]. 

These were reads were processed using the ncov2019-artic-nf pipelines, which is a Nextflow pipeline for running the ARTIC network's fieldbioinformatics tools, with a focus on ncov2019. 


This dataset includes: 

  • FailedQC - A cohort of 24 samples failed basic QC metrics, covering 8 possible failure scenarios, Illumina platform, amplicon-based approach    
  • VOCRepresentatives - A cohort of 16 samples from 10 representative CDC defined VOI/VOC lineages as of 06/15/2021, Illumina platform, amplicon-based approach    
  • Test - Smaller test samples, including sequenced negative controls of varying quality

[1]  Timme, Ruth E., et al. "Benchmark datasets for phylogenomic pipeline validation, applications for foodborne pathogen surveillance." PeerJ 5 (2017): e3893. 

Files

FailedQC.zip

Files (1.2 GB)

Name Size
md5:37fef47374c56ec5af7913d553e08333
371.8 MB Preview Download
md5:e0e05efab83d7413d3ed0eabe4e490af
1.5 kB Preview Download
md5:65811200f6c2393eda8e7b98a597898e
84.7 MB Preview Download
md5:c9d2be6e2c80266aa9acc39b79e10be3
753.6 MB Preview Download