Published September 6, 2023 | Version v1

Molecules used to train or generated by chemical language models

Authors/Creators

  • 1. anonymized

Description

This upload contains training datasets or generated molecules from the paper “Invalid SMILES are helpful, not harmful, for chemical language models.”

The contents of the directories are as follows:

  • training_sets: sets of molecules from ChEMBL or GDB-13 used to train chemical language models, represented either as SMILES or SELFIES
  • sampled-*: unprocessed samples of 10 million molecules from each model trained on ChEMBL or GDB-13
  • prior_inputs: sets of molecules from LOTUS, COCONUT, FooDB and NORMAN, split into ten folds and used to train chemical language models
  • priors-*: samples of 100 million molecules from chemical language models trained on each cross-validation fold, with unique molecules represented as canonical SMILES and sorted in descending order by their sampling frequency

Files

prior_inputs.zip

Files (119.2 GB)

Name Size
md5:e491360f9677b2cefe45a3e4428c5a9d
276.8 MB Preview Download
md5:82da4002d7ceb11a6fcda5543232eb19
29.9 GB Preview Download
md5:7e5bb4c823254b0f4f7557062a682fda
2.8 GB Preview Download
md5:ec8337e95327f44401251aa8e6f6fb43
24.1 GB Preview Download
md5:a00cc47baa44d34f591c5490cd51a2f5
10.9 GB Preview Download
md5:5811d855a63e22a3d5df84ec8eddf46a
31.1 GB Preview Download
md5:ec3a97ca80e0349e0020d59c8def715a
18.7 GB Preview Download
md5:709ec6fb69a869152636e3f503b243f9
1.4 GB Preview Download