Published September 6, 2023
| Version v1
Dataset
Open
Molecules used to train or generated by chemical language models
Description
This upload contains training datasets or generated molecules from the paper “Invalid SMILES are helpful, not harmful, for chemical language models.”
The contents of the directories are as follows:
- training_sets: sets of molecules from ChEMBL or GDB-13 used to train chemical language models, represented either as SMILES or SELFIES
- sampled-*: unprocessed samples of 10 million molecules from each model trained on ChEMBL or GDB-13
- prior_inputs: sets of molecules from LOTUS, COCONUT, FooDB and NORMAN, split into ten folds and used to train chemical language models
- priors-*: samples of 100 million molecules from chemical language models trained on each cross-validation fold, with unique molecules represented as canonical SMILES and sorted in descending order by their sampling frequency
Files
prior_inputs.zip
Files
(119.2 GB)
| Name | Size | |
|---|---|---|
|
md5:e491360f9677b2cefe45a3e4428c5a9d
|
276.8 MB | Preview Download |
|
md5:82da4002d7ceb11a6fcda5543232eb19
|
29.9 GB | Preview Download |
|
md5:7e5bb4c823254b0f4f7557062a682fda
|
2.8 GB | Preview Download |
|
md5:ec8337e95327f44401251aa8e6f6fb43
|
24.1 GB | Preview Download |
|
md5:a00cc47baa44d34f591c5490cd51a2f5
|
10.9 GB | Preview Download |
|
md5:5811d855a63e22a3d5df84ec8eddf46a
|
31.1 GB | Preview Download |
|
md5:ec3a97ca80e0349e0020d59c8def715a
|
18.7 GB | Preview Download |
|
md5:709ec6fb69a869152636e3f503b243f9
|
1.4 GB | Preview Download |