Published February 20, 2025
| Version v21
Dataset
Open
Dataset for drevalpy
Authors/Creators
Description
All preprocessing details can be found in https://github.com/JudithBernett/preprocess_drp_data.
News:
- Drug IDs and names have been harmonized over datasets: Every drug is now identified by its pubchem ID. If it had no pubchem ID, the drug name is in the respective field.
- Fingerprints and DIPK MolGNet features have now been generated for all drugs with SMILES. They are filtered for each dataset such that the respective folder only contains features for drugs contained in the dataset. All drugs from CCLE, CTRPv1, and CTRPv2 had SMILES. For GDSC1 and GDSC2, there are some missing values.
- The cell line ID is now always called cellosaurus_id, the drug ID pubchem_id
- The drug features are now identified over the pubchem_id (column names in fingerprint files, suffix in MolGNet files)
- The old response files have been removed, now, for each dataset, the response file is [dataset].csv which also contains the curve-curated measurements
- Unnamed: 0 column removed from copy number variation
Files
CCLE.zip
Files
(2.1 GB)
| Name | Size | |
|---|---|---|
|
md5:5a4d3653a310538c24271048043903ee
|
312.9 MB | Preview Download |
|
md5:4d5764a14576f50967eab5df3ba048f2
|
352.6 MB | Preview Download |
|
md5:d2e44eb039658226ae64f2c72d60e8e2
|
463.2 MB | Preview Download |
|
md5:9ca9dd38e20b12b830b96899bbde37b5
|
486.6 MB | Preview Download |
|
md5:88c366e4fdcc7def7363e88f92a731e4
|
462.8 MB | Preview Download |
|
md5:5616c8b88bc135f4696d9b2aef8c738f
|
42.7 MB | Preview Download |