Published June 12, 2021 | Version v1

Sars-CoV-2 structures -- sequence-to-alignments derived from PDB and from PSSH2, plus dark regions

  • 1. HSWT
  • 2. Garvan Institute of Medical Research

Description

Aquaria Coverage map

In "SARS-CoV-2 structural coverage map reveals viral protein interactions, hijacking, and mimicry" we introduce a novel concept to visually organize a complex dataset of a large numbers of models: a one-stop visualization summarizing what is known - and not known - about the 3D structure of the viral proteome. This tailored visualization — called the SARS-CoV-2 structural coverage map — helps researchers find structural models related to specific research questions and can be viewed in the Aquaria-COVID resource.
Aquaria_COVID_Coverage_Map.csv summarises the most basic information of this map, specifying dark and non-dark regions, as well as number of residues predicted to be disordered in these regions - predicted by Meta-Disorder (Schlessinger et al, 2006).

The PSSH2 data set

The sequence-to-structure alignments were generated using a modified version of the Aquaria sequence-to-structure processing pipeline (O’Donoghue et al, 2015), making up a subset of the PSSH2 database.
PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences.
seq_to_struc_alignemnts_PSSH2.csv.gz contains a subset of the usual PSSH2 database, including only the proteins relevant to visualise Sars-CoV-2 structures. Protein sequences and PDB structures are identified by the MD5 hashes of their respective sequences. 
PDB_chain_identifier_mappings.csv and swissprot_identifier_mappings.csv detail which entries in Swissprot and PDB chain are referred to by the MD5 hashes in the PSSH2 data set.

Calculating PSSH2

The main bunch of Swissprot and PDB data was downloaded in October 2020, but incremental updates, especially as related to Covid19 were added until April 2021.
Generating PSSH2: We used Uniclust30 from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30% (http://gwdu111.gwdg.de/~compbiol/uniclust/2020_03/UniRef30_2020_03_hhsuite.tar.gz). The HHblits code and the code for running the calculations was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively) at the respective time of calculation in the timeframe until April 2021. 

PDB based sequence-to-structure alignments

In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These alignments are summarised in PDB_chain_alignments_pssh2Format.csv.

Files

Aquaria_COVID_Coverage_Map.csv

Files (368.1 kB)

Name Size Download all
md5:8d0fc489a10a1e6748be050d467693fd
961 Bytes Preview Download
md5:f11dc0ef4f61532cbb70662fef6124e9
110.0 kB Preview Download
md5:7756ce73a161678efad4ebc4271cb3e7
202.9 kB Preview Download
md5:f65123942d1f3d3d8063d3b00f2b3e8c
53.4 kB Download
md5:42e42b6786ed8b753338a112c70440d0
783 Bytes Preview Download

Additional details

Related works

Continues
Journal article: 10.1093/nar/gkg110 (DOI)
Is documented by
Preprint: 10.1101/2020.07.16.207308v5 (DOI)
References
Journal article: 10.1038/nmeth.3258 (DOI)