There is a newer version of the record available.

Published February 20, 2025 | Version v21

Dataset for drevalpy

  • 1. Technische Universität München
  • 2. ROR icon Freie Universität Berlin

Description

All preprocessing details can be found in https://github.com/JudithBernett/preprocess_drp_data.

News:

  • Drug IDs and names have been harmonized over datasets: Every drug is now identified by its pubchem ID. If it had no pubchem ID, the drug name is in the respective field.
  • Fingerprints and DIPK MolGNet features have now been generated for all drugs with SMILES. They are filtered for each dataset such that the respective folder only contains features for drugs contained in the dataset. All drugs from CCLE, CTRPv1, and CTRPv2 had SMILES. For GDSC1 and GDSC2, there are some missing values.
  • The cell line ID is now always called cellosaurus_id, the drug ID pubchem_id
  • The drug features are now identified over the pubchem_id (column names in fingerprint files, suffix in MolGNet files)
  • The old response files have been removed, now, for each dataset, the response file is [dataset].csv which also contains the curve-curated measurements
  • Unnamed: 0 column removed from copy number variation

Files

CCLE.zip

Files (2.1 GB)

Name Size
md5:5a4d3653a310538c24271048043903ee
312.9 MB Preview Download
md5:4d5764a14576f50967eab5df3ba048f2
352.6 MB Preview Download
md5:d2e44eb039658226ae64f2c72d60e8e2
463.2 MB Preview Download
md5:9ca9dd38e20b12b830b96899bbde37b5
486.6 MB Preview Download
md5:88c366e4fdcc7def7363e88f92a731e4
462.8 MB Preview Download
md5:5616c8b88bc135f4696d9b2aef8c738f
42.7 MB Preview Download