Published January 16, 2025 | Version v2

SCoV2-VAR: A light-weighted, customizable, and open-source database of 12 million SARS-CoV-2 genomes

Description

Explosive accumulation of SARS-CoV-2 variants is posing a challenge to monitoring virus mutation and other data dealing, particularly based on centralized databases. The present study aimed to establish a light-weighted, customizable, and open-source database for SARS-CoV-2 genomes and annotations, without any access limit. The database, named SCoV2-VAR, was constructed, based on the variations (VAR) of the full-length SARS-CoV-2 (SCoV2) data uploaded on websites. All sequence samples were subject to quality control, single nucleotide polymorphism (SNP) annotation, format conversion, and final compression before appending to SCoV2-VAR. The final version of SCoV2-VAR (up to Feb 2024) contained more than 12 million SARS-CoV-2 records, with full genome and annotations. SCoV2-VAR was extremely light-weighted, with a storage size of 937 Mb for all 12 million sequences, post a 1: 596 compression. SCoV2-VAR is capable of timely updating, quickly querying, and customizable outputting SARS-CoV-2 sequences and their annotations. Additionally, the present study provided an overview of all 12 million SARS-CoV-2 samples, for both sequences and annotations.


Notes

Here, we do not provide the annotation information and complete sequence information from GISAID. If needed, you can retrieve it on GISAID based on the sequence ID number. The link is: https://gisaid.org/

Files

Files (508.2 MB)

Name Size
md5:772e41e7019895e55ac5118bc58c7fba
508.2 MB Download

Additional details

Software

Repository URL
https://github.com/Jamalijama/SCoV2-VAR
Development Status
Active