Published April 18, 2023 | Version v1.0.0

PubChem compounds before and after standardization

Description

This repository contains a subset of 200k PubChem compounds before and after standardization.
This dataset is the basis of the PubChem-pretrained model presented in "Standardizing chemical compounds with language models" (available on ChemRxiv, see also the associated GitHub repository).

The associated pretrained model is also provided here, along with the splits used for training.

The data is provided under the CDLA-Sharing-1.0 license.

Provided files:

  • README.md: General README.
  • src_and_tgt_all.csv: All the compounds before and after standardization, in CSV format.
  • src-train.txt: The tokenized compounds before standardization in the train split.
  • tgt-train.txt: The tokenized compounds after standardization in the train split.
  • src-valid.txt: The tokenized compounds before standardization in the validation split.
  • tgt-valid.txt: The tokenized compounds after standardization in the validation split.
  • src-test.txt: The tokenized compounds before standardization in the test split.
  • tgt-test.txt: The tokenized compounds after standardization in the test split.
  • LICENSE.md: The details of the CDLA-Sharing-1.0 license.
  • pretrained_pubchem_step_120000.pt: the pretrained model.

Files

LICENSE.md

Files (107.6 MB)

Name Size Download all
md5:b90e452228bee63aa71c347b8503726e
11.4 kB Preview Download
md5:0e834f6b1153af4e785cd738a1b0c0cb
57.5 MB Download
md5:619392deed865f33c0096d8b5f46b9a7
1.4 kB Preview Download
md5:ca62ec35ff9ac60fa40d2f4ae5bf058a
939.4 kB Preview Download
md5:1d12b8cbc82edf8f053689320efed072
14.8 MB Preview Download
md5:abb8d5d1bed7f3a397faada7a6ea3aa3
622.2 kB Preview Download
md5:9d6f6022c2f5099b47f26e45e6bdce97
17.4 MB Preview Download
md5:74aaf488ff7c427602f5c2116ddb61dd
937.5 kB Preview Download
md5:e4ab4fe922b40ec23004462b73e4febb
14.8 MB Preview Download
md5:40d42d865e487aedfe2560f8d0190e4c
620.7 kB Preview Download

Additional details

Related works