Code and Data for "Identification of drug candidates against glioblastoma with machine learning and high-throughput screening of heterogeneous cellular models"
Authors/Creators
Description
Code and data for paper "Identification of drug candidates against glioblastoma with machine learning and high-throughput screening of heterogeneous cellular models" by Smer-Barreto, Elliott, Dawson, Lorente-Macías, Furqan, Unciti-Broceta, Oyarzún, Carragher, under review in RSC Digital Discovery, 2025.
GBM_ML_virtual_screening.py includes python code for the user to perform their own virtual screening.
1_train_gbm.csv and 2_predict_all.csv are the training and prediction files, respectively, used in the original paper.
Code usage
The code needs three elements to operate: a training file, a prediction file, and an iteration number, in that order:
python GBM_ML_virtual_screening.py <training_file> <prediction_file> <iteration_number>
Example:
python GBM_ML_virtual_screening.py 1_train_gbm.csv 2_predict_all.csv 10
The training file should contain a column with the names of the compounds (as in column 1 in 1_train_gbm.csv) a target column (as in column 2 in 2_predict_all.csv) and a list of features of choice, such as columns 6 and above 1_train_gbm.csv.
The prediction file should contain a column with the names of the compounds (as in column 1 in 2_predict_all.csv) and feature columns to match the numerical measures of those contained in the training file, such as columns 4 and above in 2_predict_all.csv.
The iteration number allows the user to specify how many iterations the Monte Carlo algorithm for virtual screening will perform.
Code modularity
The code contains four functions that can be adapted to the user's needs:
prepare_training_data: prepares training file for later use in ML pipeline.
feature_importance: calculates an array of feature importance values given the features provided for training. It outputs the ordered columns' feature importance as a histogram (feature_importance.pdf) and a csv file (rdkit_importance_features.csv).
prepare_prediction_data: prepares prediction file for virtual screening.
MonteCarlo_loop: performs Monte Carlo machine learning virtual screen. The function saves the models produced in .sav files and outputs the performance metrics of said models (precision, recall, accuracy, f1 score, FPR, AUC, PR-AUC) in file rsc_gbm_metrics.csv and the virtual screen predictions in file rsc_gbm_predictions.csv.
Files
1_train_gbm.csv
Additional details
Identifiers
Related works
- Is supplement to
- 10.1101/2025.03.06.641926 (DOI)
Dates
- Submitted
-
2025-09-12
Software
- Programming language
- Python
References
- Website. RDKit: Open-source cheminformatics. Available: https://www.rdkit.org.
- Louppe, Wehenkel & Sutera. Understanding variable importances in forests of randomized trees. Adv. Neural Inf. Process. Syst.