Published September 12, 2025 | Version 2

Code and Data for "Identification of drug candidates against glioblastoma with machine learning and high-throughput screening of heterogeneous cellular models"

Contributors

  • 1. The University of Edinburgh

Description

Code and data for paper "Identification of drug candidates against glioblastoma with machine learning and high-throughput screening of heterogeneous cellular models" by Smer-Barreto, Elliott, Dawson, Lorente-Macías, Furqan, Unciti-Broceta, Oyarzún, Carragher, under review in RSC Digital Discovery, 2025. 

GBM_ML_virtual_screening.py includes python code for the user to perform their own virtual screening. 

1_train_gbm.csv and 2_predict_all.csv are the training and prediction files, respectively, used in the original paper.

 

Code usage

The code needs three elements to operate: a training file, a prediction file, and an iteration number, in that order: 

python GBM_ML_virtual_screening.py <training_file> <prediction_file> <iteration_number> 

Example:

python GBM_ML_virtual_screening.py 1_train_gbm.csv 2_predict_all.csv 10

The training file should contain a column with the names of the compounds (as in column 1 in 1_train_gbm.csv) a target column (as in column 2 in 2_predict_all.csv) and a list of features of choice, such as columns 6 and above 1_train_gbm.csv. 

The prediction file should contain a column with the names of the compounds (as in column 1 in 2_predict_all.csv) and feature columns to match the numerical measures of those contained in the training file, such as columns 4 and above in 2_predict_all.csv.

The iteration number allows the user to specify how many iterations the Monte Carlo algorithm for virtual screening will perform. 

Code modularity

The code contains four functions that can be adapted to the user's needs: 

prepare_training_data: prepares training file for later use in ML pipeline. 

feature_importance: calculates an array of feature importance values given the features provided for training. It outputs the ordered columns' feature importance as a histogram (feature_importance.pdf) and a csv file (rdkit_importance_features.csv). 

prepare_prediction_data: prepares prediction file for virtual screening. 

MonteCarlo_loop: performs Monte Carlo machine learning virtual screen. The function saves the models produced in .sav files and outputs the performance metrics of said models (precision, recall, accuracy, f1 score, FPR, AUC, PR-AUC) in file rsc_gbm_metrics.csv and the virtual screen predictions in file rsc_gbm_predictions.csv. 

Files

1_train_gbm.csv

Files (25.2 MB)

Name Size Download all
md5:08c14cbae5093326d51e1c506a9b2582
5.4 MB Preview Download
md5:c219ba03891cb4e2c9f6b298290c976e
19.8 MB Preview Download
md5:7c394d463fc6d804f9b11c70fb7647a9
7.8 kB Download

Additional details

Related works

Is supplement to
10.1101/2025.03.06.641926 (DOI)

Dates

Submitted
2025-09-12

Software

Programming language
Python

References

  • Website. RDKit: Open-source cheminformatics. Available: https://www.rdkit.org.
  • Louppe, Wehenkel & Sutera. Understanding variable importances in forests of randomized trees. Adv. Neural Inf. Process. Syst.