Performance of chemical structure string representations for chemical image recognition using transformers dataset
Description
The datasets contain string representations used for DECIMER short communication paper.
ChEMBL dataset:
Train and test datasets downloaded from ChEMBL and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.
String representations contain SMILES, DeepSMILES, SELFIES and InChIs.
- Train dataset: 1.5 Mio molecules
- Test dataset: ~100K molecules
Pubchem dataset:
Train and test datasets downloaded from PubChem and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.
String representations contain SMILES, DeepSMILES and SELFIES.
- Train dataset: 3 Mio molecules
- Test dataset: 250K molecules
Files
Files
(707.1 MB)
| Name | Size | |
|---|---|---|
|
md5:b64e1bd8ae00128f335ceefdd602ec9e
|
707.1 MB | Download |