Published October 30, 2022 | Version v1

A systematic assessment of deep learning methods for drug response prediction: From in vitro to clinical applications

Description

## GDSC dataset

**GDSC_EXP.csv** GDSC gene expression profiles for 966 cancer cell lines, where each column represents a cell line in the form of its name and tissue collection site, and each row represents a gene in the form of the HGNC symbol.

 

**GDSC_MUT.csv** GDSC gene mutation profiles for 966 cancer cell lines, where each column represents a cell line in the form of its name and tissue collection site, and each row represents a gene in the form of the HGNC symbol. The wild type is coded as 1 and the wild type as 0.

 

**GDSC_CNV.csv** GDSC copy number variation profiles for 966 cancer cell lines, where each column represents a cell line in the form of its name and tissue collection site, and each row represents a gene in the form of the HGNC symbol. The copy-neutral is coded as 0 and the deletion or amplification as 1.

 

**GDSC_DR.csv** GDSC drug response data for 966 cancer cell lines and 282 drugs in the form of the natural logarithm of the IC50 readout. The first column shows the cell line name and tissue collection site, the second column shows the drug name, and the third column shows the drug response readout.

 

**GDSC_DrugAnnotation.csv** GDSC annotations for 282 drugs include drug name, PubChem CID, PubChem canonical SMILES, Rdkit canonical SMILES, Target Pathway, standard deviation, bimodality coefficient and density coverage.

## TCGA dataset

**TCGA_EXP.csv** TCGA gene expression profiles, where each column represents a patient in the form of TCGA patient ID, and each row represents a gene in the form of the HGNC symbol.

 

**TCGA_MUT.csv** TCGA gene mutation profiles, where each column represents a patient in the form of TCGA patient ID, and each row represents a gene in the form of the HGNC symbol. The wild type is coded as 1 and the wild type as 0.

 

**TCGA_CNV.csv** TCGA copy number variation profiles, where each column represents a patient in the form of TCGA patient ID, and each row represents a gene in the form of the HGNC symbol. The copy-neutral is coded as 0 and the deletion or amplification as 1.

 

**TCGA_DR.csv** TCGA clinical response data. The first column shows the TCGA patient ID, the second column shows the drug name, the third column shows the clinical response category, the fourth column shows the cancer type, and the last column shows the clinical label as responder or non-responder.

## PMID17185464 (Bortezomib) dataset

**PMID17185464_EXP.csv** Bortezomib clinical trial gene expression profiles, where each column represents a patient in the form of patient ID, and each row represents a gene in the form of the HGNC symbol.

**PMID17185464_DR.csv** Bortezomib clinical trial clinical response data. The first column shows the TCGA patient ID, the second column shows the drug name, the third column shows the clinical response category, and the last column shows the clinical label as responder or non-responder (NR: Non-responder, R: Responder).

Files

GDSC_CNV.csv

Files (1.4 GB)

Name Size
md5:153e897447b121069009ce70fb63d9d0
43.0 MB Preview Download
md5:219b0569ef9a657b704173f416cff229
14.2 MB Preview Download
md5:40d16a1966d45b87b6d2ead570877135
54.4 kB Preview Download
md5:4cd14f5e00814f1728ed1c673b733691
457.4 MB Preview Download
md5:8fd663e42f37d070704a058e1309a894
39.7 MB Preview Download
md5:655d721fc7bb09abb102d3b5cf619412
4.7 kB Preview Download
md5:41ca695de7ccae4107320c95c4b5c3a7
56.0 MB Preview Download
md5:beaa0e93f5ad5b8281231e0f561fc726
56.4 MB Preview Download
md5:c774b9d0b10872558d4afaaac9647f3e
92.7 kB Preview Download
md5:a1fefba102f3fb8a979bb4e9eb1ccd93
681.0 MB Preview Download
md5:0177d96d062246722720615627d529d2
45.1 MB Preview Download