Sensitivity Datasets - Leveraging Implicit Knowledge in Neural Networks for Functional Dissection and Engineering of Proteins
Authors/Creators
- 1. Synthetic Biology Group, Institute for Pharmacy and Biotechnology (IPMB) and Center for Quantitative Analysis of Molecular and Cellular Biosystems (BioQuant), University of Heidelberg, Heidelberg, 69120, Germany; Digital Health Center, Berlin Institute of Health (BIH) and Charité University Medicine, Berlin, 10117, Germany
- 2. Synthetic Biology Group, Institute for Pharmacy and Biotechnology (IPMB) and Center for Quantitative Analysis of Molecular and Cellular Biosystems (BioQuant), University of Heidelberg, Heidelberg, 69120, Germany
- 3. Molecular Epidemiology Unit, Berlin Institute of Health (BIH) and Charité University Medicine, Berlin, 10117, Germany
- 4. Digital Health Center, Berlin Institute of Health (BIH) and Charité University Medicine, Berlin, 10117, Germany; Health Data Science Unit, University Hospital Heidelberg, Heidelberg, 69120, Germany
Description
Leveraging Implicit Knowledge in Neural Networks for Functional Dissection and Engineering of Proteins
The Sensitivity datasets cover more than 800 proteins and are structured as follows. The sensitivity values are the mean of four DeeProtein replicates.
It is uploaded as tar.gz. and contains one directory.
File names contain the PDB1 identifier and the respective chain identifier.
The sequences and secondary structure information were downloaded from the RCSB Protein Databank and are available here: https://cdn.rcsb.org/etl/kabschSander/ss_dis.txt.gz This URL can be found with some explanation at http://www.rcsb.org/pdb/static.do?p=download/http/index.html
The secondary structure annotation relies on the DSSP Algorithm by Kabsch and Sander2.
The files are tab-separated and contain the following columns:
- Pos Position in the sequence, starting from zero
- AA Amino acid in that position
- sec Secondary structure as annotated in the RCSB Protein Databank
- dis if a region has not been experimentally observed (sometimes explains mismatches with crystal structures)
- GO:_______ Sensitivity for the GO term
References
- The Protein Data Bank H.M. Berman, J. Westbrook, Z. Feng, G. Gilliland, T.N. Bhat, H. Weissig, I.N. Shindyalov, P.E. Bourne (2000) Nucleic Acids Research, 28: 235-242. doi:10.1093/nar/28.1.235
- Kabsch, W. & Sander, C. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22, 2577-2637, doi:10.1002/bip.360221211 (1983).
Notes
Files
Files
(9.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:ef89cd0f2f8e95d7408aaf4cfdd7c63d
|
9.3 MB | Download |
Additional details
Related works
- Is compiled by
- https://github.com/juzb/DeeProtein (URL)
- 10.5281/zenodo.1402827 (DOI)
References
- The Protein Data Bank H.M. Berman, J. Westbrook, Z. Feng, G. Gilliland, T.N. Bhat, H. Weissig, I.N. Shindyalov, P.E. Bourne (2000) Nucleic Acids Research, 28: 235-242. doi:10.1093/nar/28.1.235
- Kabsch, W. & Sander, C. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22, 2577-2637, doi:10.1002/bip.360221211 (1983).