Published December 12, 2024 | Version v2

USPTO-LLM: A Large Language Model-Assisted Information-enriched Chemical Reaction Dataset

  • 1. ROR icon Renmin University of China

Description

USPTO-LLM is an information-enriched chemical reaction dataset that provides more side information (reaction conditions and reaction steps division) for developing new reaction prediction and retrosynthesis methods and inspires new problems, such as reaction condition prediction. It comprises over 247K chemical reactions extracted from the patent documents of USPTO (United States Patent and Trademark Office), encompassing abundant information on reaction conditions. 

We employ large language models to expedite the data collection procedures automatically with a reliable quality control process. The extracted chemical reactions are organized as heterogeneous directed graphs, allowing us to formulate a series of prediction tasks, such as reaction prediction, retrosynthesis, and reaction condition prediction, in a unified graph-filling framework.

Files

multi_step.zip

Files (702.1 MB)

Name Size
md5:4c151f3375aeb80f3f1f3ec7abc72cde
13.6 MB Preview Download
md5:c2d3e395e0ca283d880e1987beb0f8ad
2.3 kB Preview Download
md5:8027c066fd65d35b454049c78e6ccbc7
688.5 MB Preview Download

Additional details

Dates

Created
2024-12-12