Published February 9, 2022 | Version 2021-11

PSSH2 - database of protein sequence-to-structure homologies (including Sars-CoV-2 structures)

  • 1. HSWT
  • 2. Garvan Institute of Medical Research

Description

Protein sequence and structure data

This data set contains data from Uniprot (in the files called protein_sequence, protein_synonyms, protein_names, organism_synonyms) and PDB (in the files called PDB and PDB_chain) as used by the Aquaria web resource at the time of download (2022-02-08).

 

The PSSH2 data set

PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences.

 

Calculating PSSH2

The Swissprot and PDB data was downloaded in November 2021.
Generating PSSH2: We used UniRef30_2021_03 (originally called UniRef30_2021_06) from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30%. The HHblits code and the code for running the calculations was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively) at the respective time of calculation in the timeframe until December 2021. 
 

PDB based sequence-to-structure alignments

In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These are contained in the PDBchain data.

This data covers sequences and PDB structures in the timeframe until February 2022. 

 

Evaluating PSSH2

The resulting alignment data was analysed using CATH domain assignments downloaded from /cath/releases/all-releases/v4_2_0/cath-classification-data/ to define correct hits and false hits: 

  • The set of query sequences is defined by the CATH non-redundant S40_overlap_60 dataset (ftp://orengoftp.biochem.ucl.ac.uk/cath/releases/all-releases/v4_2_0/non-redundant-data-sets/)
  • The set of all expected hits are all pdb structures containing a domain with the same CATH code if contained in the set of processed sequences (-> all) or only if also contained in the set of non redundant sequences (-> nr40).
  • The set of true positives is defined by sharing the same CATH code up to the level of homology ("CATH") or up to the level of topology ("CAT").

The data was evaluated with respect to false discovery rate (FDR) and recall (true positive rate TPR) by cumulatively considering all hits with an E-value below the threshold ("C") or in bins with an E-value between the threshold and one tenth of the threshold ("B"). This evaluation was carried out for the data obtained in November 2021 (202111) as well as previous data from October 2020 (202010), February 2020 (202002) and September 2017 (201709). The results are  collected in PSSH CATH validation.csv

 

Known errors

Due to processing error, the profile of pdb structure 5fia A / B (sequence md5 052667679fc644184f40063c7602c9e1) is incomplete in the pdb_full hhblits database which led to further errors in generating sequence based alignments for sequences for 1vtm P (sequence md5 c844aff103449363cb8489c78c58ebf1) and 434t A / B (sequence md5 d67aa1c3a36492c719cb48b5e7ecc624).

 

Files

PSSH_CATH_validation.202111.csv

Files (9.2 GB)

Name Size
md5:717ab4b49f2dabac86afeebc60eaf336
9.3 MB Download
md5:289087c55f5d8e8fa187f79bf0fa744e
438.3 kB Download
md5:21e0b4fcad4d63ef0a15ef7e4c2aac75
18.5 MB Download
md5:c4f0603494989714871366f54cd55dd4
195.6 MB Download
md5:751ecf596ef6c7e1dd104644a6bf6e05
534 Bytes Download
md5:bd7599085a1f4568994dc3e4366ecc56
199.4 MB Download
md5:e7e426799afc80633b4ad07d344fccde
40.5 MB Download
md5:cd25c96bc7c1a216d1ca8029d5fb2e9e
8.7 GB Download
md5:dfcceb05f0c60555d93e9d70bd84a6d2
11.8 kB Preview Download

Additional details

Related works

Continues
Journal article: 10.1093/nar/gkg110 (DOI)
Is documented by
Journal article: 10.1101/2020.07.16.207308v5 (DOI)
References
Journal article: 10.1038/nmeth.3258 (DOI)