Dataset Open Access

Solvated protein fragments

Unke, Oliver Thorsten; Meuwly, Markus

The solvated protein fragments dataset probes many-body intermolecular interactions between 
"protein fragments" and water molecules, which are important for the description of many 
biologically relevant condensed phase systems. It contains structures for all possible 
"amons" [1] (hydrogen-saturated covalently bonded fragments) of up to eight heavy atoms 
(C, N, O, S) that can be derived from chemical graphs of proteins containing the 20 natural
amino acids connected via peptide bonds or disulfide bridges. For amino acids that can occur 
in different charge states due to (de-)protonation (i.e. carboxylic acids that can be 
negatively charged or amines that can be positively charged), all possible structures with 
up to a total charge of +-2e are included. In total, the dataset provides reference energies, 
forces, and dipole moments for 2731180 structures calculated at the revPBE-D3(BJ)/def2-TZVP 
level of theory [2-5] using the ORCA 4.0.1 code [6,7]. 

For more details, see https://arxiv.org/abs/1902.08408.

[1] Huang, B. and von Lilienfeld, O. A. arXiv:1707.04146 (2017).
[2] Grimme, S.; Antony, J.; Ehrlich, S. and Krieg, H. J. Chem. Phys. 132, 154104 (2010).
[3] Grimme, S.; Ehrlich, S. and Goerigk, L. J. Comput. Chem. 32, 1456-1465 (2011).
[4] Weigend, F. and Ahlrichs, R. Phys. Chem. Chem. Phys. 7, 3297-3305 (2005).
[5] Zhang, Y. and Yang, W. Phys. Rev. Lett. 80, 890 (1998).
[6] Neese, F. Wiley Interdiscip. Rev. Comput. Mol. Sci. 2, 73-78 (2012).
[7] Neese, F. Wiley Interdiscip. Rev. Comput. Mol. Sci. 8, e1327 (2018).

Files (1.4 GB)
Name Size
read_data.py
md5:93ea0c2cc018b558ed679998bbab59f6
760 Bytes Download
README.txt
md5:6e67fb6765efd8cb5b0e7692f1a5fe58
3.5 kB Download
solvated_protein_fragments.npz
md5:6484ec24acd1b3da5defe962b4c4ecf3
1.4 GB Download
79
62
views
downloads
All versions This version
Views 7979
Downloads 6262
Data volume 50.9 GB50.9 GB
Unique views 6666
Unique downloads 3030

Share

Cite as