This dataset is the official test set from the "Chemical Structure Recognition in Images" task held as part of the CLEF-IP 2012 workshop. It contains 992 images of chemical structures extracted from US patents, each paired with a ground-truth MOL file. This Hugging Face version has been augmented with canonical SMILES, InChI, and SELFIES strings to provide a comprehensive resource for evaluating image-to-structure models… See the full description on the dataset page:
https://huggingface.co/datasets/hheiden/CLEF_OCSR_benchmark.