Published November 2, 2021 | Version v1

Bridging In-Silico and Experimental: Chemoinformatics Investigation for Mass Spectrometry-Based Metabolomics Study of Soybean

Authors/Creators

  • 1. CSIR National Chemical Laboratory

Contributors

  • 1. CSIR National Chemical Laboratory

Description

Soybean (Glycine max L. Merr.) is a globally important legume crop and contains various small organic molecules that are valuable sources for drug development. This study intended to identify, analyze and design a virtual library of prioritized novel and promising drug-like molecules based on the analysis of secondary metabolites of soybean using chemoinformatics and untargeted mass spectrometry (UHPLC-MS/MS) approaches. In this study, we performed chemoinformatics analysis of previously reported and unreported secondary metabolites from four soybean varieties. The secondary metabolites were identified by UHPLC-MS/MS analysis and text mining, and a virtual library of novel molecules was generated. The metabolomics data were analyzed using machine learning-based quantitative and qualitative methods for identifying putative metabolites by spectral matching and multivariate statistical analysis. A representative virtual library of novel molecules was generated and prioritized further by virtual screening methods. We detected 6628 annotated mass features for small molecules that have not been reported in soybean before, in addition to 443 mass features of molecules that were previously reported in the literature. Tandem mass spectrometry (MS/MS) confirmed the presence of 14 new and six previously reported soybean molecules. We found high molecular diversity in seed and leaf tissues of four soybean varieties (NRC-119, JS-335, JS-7105, and JS-9305). We identified 25 common scaffolds and 231 molecules through scaffold-molecule networks between soybean molecules and known drugs. Five representative scaffolds were used to build a focused virtual library of novel molecules (n= 1225), which were virtually screened to obtain potential drug-like candidates (n= 815) for further studies. We developed a novel virtual library of molecules with drug-like and lead-like properties for further drug discovery-related studies. This study suggests that a combinatorial approach employing high-throughput metabolomics and chemoinformatics methods can efficiently identify new drug-like and lead-like candidates from plant metabolites.

Notes

Supplementary Data

Files

Files (58.6 MB)

Name Size Download all
md5:f9fe33aa7eb77f8fa515368abf2b0ca1
5.2 MB Download
md5:42167946e9252937fe3578b424c31e9d
23.2 MB Download
md5:bfe299d20abd8296430e90d1f4ec4c15
20.2 MB Download
md5:85a61509fd016538e903ccd751b1ed5e
8.4 MB Download
md5:3de8e764c2cb45069b7345a91fb6bfe0
28.6 kB Download
md5:96d9efd1532e044589e10d99e74e7df7
1.5 MB Download

Additional details

Related works