There is a newer version of the record available.

Published July 19, 2021 | Version 1.0

High quality protein residues: top2018 all-atom-filtered residues

  • 1. Duke University

Description

Introduction
--------------------------------------------------------------------------------
This directory contains files from the top2018 dataset by the Richardson Lab at Duke University.

These are high-quality residues from high-quality, low redundancy protein chains in the PDB.


Usage recommendations
--------------------------------------------------------------------------------
Protein residues that fail the filtering criteria described below have been removed from the files.  As a result, these files can be considered pre-filtered and will return only results for residues of good model quality with supporting experimental data.  All protein atoms have been considered in filtering; these files should be usable for any protein question.  If your work is strictly limited to mainchain atoms (plus CB), there is a separate version that has been filtered on only mainchain atoms.

The top2018 contains several different levels of homology clustering to ensure nonredundant datasets.  The 70% homology level is a reliable default.  These chains are listed in top2018_chains_hom70_allfilter_60pct_complete.txt and found in top2018_pdbs_all_filteredhom70.tar.gz

Files are organized in subdirectories based on the first two letters of their PDB ids.  The included python script sample_file_loop.py may aid in accessing the directory structure.

Files already contain hydrogens added by Reduce.  NQH flips have been performed to ensure that these are the best versions of these structures.

top2018_metadata_all_filtered.csv contains information on release data, resolution, and validation scores.

top2018_passrates_all_filtered.csv contains information on how many protein residues from the original chain passed the quality filters.


Homology sets:
--------------------------------------------------------------------------------
Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the best chain was selected for inclusion in the dataset.  This ensures minimal sequence/structural redundancy.

The top2018 is available at several different levels of homology clustering, which may be appropriate to different uses.  Lists of the included chains at each homology level are included in this distribution.

Lower homology numbers mean greater variety and less redundancy, but also fewer total chains in the dataset.

For general use, ***we recommend the 70% homology set*** as a good balance between inclusivity and variety. This list is given in the file top2018_chains_hom70_allfilter_60pct_complete.txt


Usage caveats:
--------------------------------------------------------------------------------
These files are incomplete.  They are single chains from structures that may have had multiple chains.  Residues that fail the filtering criteria have been removed.  Programs with strong requirements for completeness or uninterrupted chains should be used with care.  Chain completeness and fragmentation statistics are available in top2018_passrates_all_filted.csv and in USER records at the end on each .pdb file.

All header information from the original structure has been preserved.  This includes information about chains and residues no longer present in the file.

All ligands and waters associated with the chain have been preserved without filtering.  Robust ligand filtering is beyond the scope of this dataset.  Trust the ligands at your own discretion.


Filtering criteria: Chain level
--------------------------------------------------------------------------------
Chain is protein
Released on or before Dec 31, 2018
Resolution < 2.0
MolProbity Score < 2.0
<3% residues have cbeta deviations
<2% residues have covalent bond length outliers
<2% residues have covalent bond geometry outliers

Using sequence homology clusters provided by the RCSB PDB, for each homology cluster, the chain with the best (lowest) average of Resolution and MolProbity Score was selected.


Filtering criteria: Residue level
--------------------------------------------------------------------------------
Even good structures may contain poorly-resolved regions.  Residue-level filtering helps avoid including these regions in otherwise high-quality data

All atoms in a residue:
Bfactor <= 40
Real-space correlation coefficient (rscc) >= 0.7
2Fo-Fc map value >= 1.2

Additionally, residues are not allowed to have:
Covalent geometry outliers
Steric overlaps or "clashes", as per Probe
Alternate conformations


Chain Completeness criteria
--------------------------------------------------------------------------------
Chains which lost >40% of their residues during filtering were dropped from this dataset.  All chains present here are at least 60% complete.

Filtering doumentation
--------------------------------------------------------------------------------
Each file documents its pruned and incluced residues with USER records.  These include self-documenting USER  DOC lines as follow:
USER  DOC Lines marked with USER  DEL list residues pruned by
USER  DOC quality filtering.
USER  DOC Format is chain:resseq:icode:reason_for_pruning
USER  DOC Reasons for pruning are abbreviated as 1-letter codes: bcmgoa
USER  DOC b=bfactor, c=real space correlation, m=2Fo-Fc mapvalue
USER  DOC g=geometry outlier, o=steric overlap, a=alternate conformations
USER  DOC Lines marked USER  INC list the uninterrupted fragments of structure
USER  DOC still included after pruning by quality filtering
USER  DOC Format is chain1:resseq1:icode1:chain2:resseq2:icode2:fragment_length
USER  DOC where 1 is the first and 2 the last residue of the fragment
USER  DOC Line marked with USER  PCT gives statistics for structure completeness

Files

top2018_chains_all_allfilter_60pct_complete.txt

Files (3.8 GB)

Name Size
md5:e8318de211fd5c68738e4ef04b9c85d2
5.7 kB Download
md5:308e491c0edf570717d1472bd4300d1a
646 Bytes Download
md5:c6022add84879fef1f9bf431ad4ef754
94.4 kB Preview Download
md5:0986229c6deca79bb368b91c4cbb59e7
50.7 kB Preview Download
md5:ac66581585489bc6922cd007870c90c8
72.9 kB Preview Download
md5:8c3dc3708dbeb71ec0875b95e8297608
84.9 kB Preview Download
md5:6ea17692502e513297dbd86a37fafd04
94.2 kB Preview Download
md5:f718fc4c9be14624632c5f3c4caa032d
1.8 MB Preview Download
md5:0454de5d8d3c6ab27ecff104f30d6829
305.5 kB Preview Download
md5:ec1affc0540631a412a896810bb034a4
627.9 MB Download
md5:fb5efbb4b090df29163e3a2d86a311a4
922.9 MB Download
md5:f0ca2f8f34ad4e0aab323e1d3c1d6d25
1.1 GB Download
md5:dc33d0e16d3043073e49453643f2c21a
1.2 GB Download