Published May 17, 2024
| Version v1
Dataset
Open
MolClassifier Training and Validation Datasets
Authors/Creators
Description
The dataset contains 18626 chemical images (15720 for training and 2906 for validation) with annotated classes: `Molecular Structure`, `Markush Structure` and `Background`.
Selected chemical images are randomly selected from the outputs of a segmentation module applied to documents from the United States Patent and Trademark Office.
This dataset is part of PatCID: an open-access dataset of chemical structures in patent documents.