Published July 25, 2024 | Version v1
Other Open

Leveraging high-throughput data and unsupervised learning to characterize cancer metabolic heterogeneity

  • 1. ROR icon École Polytechnique Fédérale de Lausanne
  • 2. ROR icon University of Milano-Bicocca

Description

Diseases such as obesity, diabetes and cancer are influenced by various factors, including genetics and environmental conditions. Recent research indicates that metabolic alterations play a significant role in the development and progression of these diseases. However, directly measuring metabolic fluxes remains challenging due to technical and financial constraints. To address this problem, genome-scale metabolic models (GEMs) provide a comprehensive computational framework for predicting reaction fluxes. These models use techniques such as linear programming to numerically simulate metabolism. Consequently, integration methods are employed to incorporate high-throughput omics data within metabolic models to predict fluxomics with higher precision. Unfortunately, their evaluation is limited to just a few fluxes that can be directly measured in the laboratory.

This thesis aims to overcome the limitations of benchmarking integration methods by using intracellular metabolomics as a reference. Two integration methods were applied to 513 different cancer cells from the Cancer Cell Line Encyclopedia (CCLE): INTEGRATE, a constraint-based steady-state method, and scFEA, a novel algorithm based on artificial neural networks.
Using unsupervised learning techniques, clusters derived from metabolomics, transcriptomics, and fluxomics were compared to evaluate their concordance in recognizing metabolic phenotypes.

A high level of coherence was found between metabolomics and transcriptomics, identifying similar metabolic subpopulations and suggesting the robustness of metabolomics as a benchmark. The results highlighted the necessity, even for bulk samples, of preprocessing the RNA-seq matrix with a denoising algorithm, specifically MAGIC. MAGIC significantly
decreased matrix sparsity by revealing lost genetic information and increased the concordance level between metabolomics and inferred INTEGRATE fluxomics. Conversely, scFEA inherently managed to reduce noise and achieved similar performance to the INTEGRATE framework. Overall, neither INTEGRATE nor scFEA drastically outperformed the other in comparing their
metabolic clusters with the benchmark. Eventually, clusters identified among metabolomics, transcriptomics and predicted fluxomics can be analyzed by life scientists to extract valuable insights about the biomarkers of the subpopulations.

Files

Milazzo_Luca_Tesi_LMDS_25_07_2024.pdf

Files (14.0 MB)

Name Size Download all
md5:4b7a17b402438c09c2f5cf4e6ea524de
14.0 MB Preview Download