Consensus machine-learning models for protein-ligand binding affinity estimation

Lunghini F.; Bonanni D.; Pisapia V.; Scaturro M.; Fava A.; Manelfi C.; Palermo G.; Gadioli D.; Beccari A.R.

doi:10.5281/zenodo.5469978

Published August 25, 2021 | Version v2

Dataset Open

Consensus machine-learning models for protein-ligand binding affinity estimation

Motivation: In structure-based virtual screening, machine learning based scoring function gained popularity in the last few years as they outperformed classical scoring function. The protein-ligand system can be encoded by a set of orthogonal descriptor spaces, which are then mined by machine learning algorithms to find a relationship with the binding affinity experimental value.

Results: In this work we propose our modelling approach to derive a new scoring function, derived from a combination of multiple descriptor spaces coupled with machine learning algorithms ensembled in consensus. The SF has been trained on the PDBbind v.2019 data and has been extensively internally and externally validated on a large set of complexes. When benchmarked on the PDBbind core set, it achieved better performance than state-of-the-art counterparts, scoring: R_Pearson= 0.85-0.86 r² = 0.70-0.72 and RMSE = 1.15-1.21. As highlights: (i) an applicability domain definition has been implemented to delimit the SF’s application boundaries, and (ii) a mechanistic interpretation is proposed by investigating the contribution of each protein-ligand atom pairs in the prediction of the binding affinity, which could provide a support in the lead-optimization process.

Availability and implementation: Our scoring function is freely available through the webportal: https://predictor.exscalate.eu/

Files

Curated_dataset_2020.csv

Files (1.6 GB)

Name	Size	Download all
Curated_dataset_2020.csv md5:34bee315c8b69bf9b7592e5488665c23	65.9 MB	Preview Download
PDBbind_single_entries.zip md5:dfb2134c0a3cd0dfe95a7cd82f175978	1.6 GB	Preview Download

	All versions	This version
Views	409	327
Downloads	135	133
Data volume	50.2 GB	50.1 GB

Consensus machine-learning models for protein-ligand binding affinity estimation

Creators

Description

Files

Curated_dataset_2020.csv

Files (1.6 GB)