Machine Learning-Based Retention Time Prediction of Trimethylsilyl Derivatives of Metabolites

Sara M. de Cripan; Adrià Cereto-Massagué; Pol Herrero; Andrei Barcaru; Núria Canela; Xavier Domingo-Almenara

doi:10.3390/biomedicines10040879

Published April 11, 2022 | Version v1

Journal article Open

Machine Learning-Based Retention Time Prediction of Trimethylsilyl Derivatives of Metabolites

1. Computational Metabolomics for Systems Biology Lab, Omics Sciences Unit, Eurecat—Technology Centre of Catalonia; Centre for Omics Sciences (COS), Eurecat—Technology Centre of Catalonia & Rovira i Virgili University Joint Unit, Unique Scientific and Technical Infrastructures (ICTS); Department of Electrical, Electronic and Control Engineering (DEEEA), Universitat Rovira i Virgili
2. Centre for Omics Sciences (COS), Eurecat—Technology Centre of Catalonia & Rovira i Virgili University Joint Unit
3. Independent Researcher

In gas chromatography–mass spectrometry-based untargeted metabolomics, metabolites are identified by comparing mass spectra and chromatographic retention time with reference databases or standard materials. In that sense, machine learning has been used to predict the retention time of metabolites lacking reference data. However, the retention time prediction of trimethylsilyl derivatives of metabolites, typically analyzed in untargeted metabolomics using gas chromatography, has been poorly explored. Here, we provide a rationalized framework for machine learning-based retention time prediction of trimethylsilyl derivatives of metabolites in gas chromatography. We compared different machine learning paradigms, in addition to exploring the influence of the computational molecular structure representation to train the prediction models: fingerprint class and fingerprint calculation software. Our study challenged predicted retention time when using chemical ionization and electron impact ionization sources in simulated and real cases, demonstrating a good correct identity ranking capability by machine learning, despite observing a limited false identity filtering power in cases where a spectrum or a monoisotopic mass match to multiple candidates. Specifically, machine learning prediction yielded median absolute and relative retention index (relative retention time) errors of 37.1 retention index units and 2%, respectively. In addition, fingerprint class and fingerprint calculation software, as well as the molecular structural similarity between the training and test or real case sets, showed to be critical modulators of the prediction performance. Finally, we leveraged the structural similarity between the training and test or real case set to determine the probability that the prediction error is below a specific threshold. Overall, our study demonstrates that predicted retention time can provide insights into the true structure of unknown metabolites by ranking from the most to the least plausible molecular identity, and sets the guidelines to assess the confidence in metabolite identification using predicted retention time data.

Files

biomedicines-10-00879-v2.pdf

Files (2.2 MB)

Name	Size	Download all
biomedicines-10-00879-v2.pdf md5:14556cebbebcd4c923c61ab92c5f936c	2.2 MB	Preview Download

Additional details

European Commission
GLOMICAVE - Global Omic Data Integration on Animal, Vegetal and Environment Sectors 952908

	All versions	This version
Views	221	220
Downloads	176	175
Data volume	406.2 MB	403.9 MB

Machine Learning-Based Retention Time Prediction of Trimethylsilyl Derivatives of Metabolites

Authors/Creators

Description

Files

biomedicines-10-00879-v2.pdf

Files (2.2 MB)

Additional details

Funding