Skip to main content
zenodoopen

Images dataset for Chemical Images Classifier model

<h1>Original paper</h1> <p>The manually curated images dataset is a part of the Supplementary Materials of the paper:&nbsp;<code>A. Krasnov, S. Barnabas, T. B&ouml;hme, S. Boyer, L. Weber, Comparing software tools for optical chemical structure recognition, Digital Discovery (2024). <a href="https://doi.org/10.1039/D3DD00228D">https://doi.org/10.1039/D3DD00228D</a></code><br><br></p> <h1>Images dataset description</h1> <p>The dataset was used to generate the image classifier model. The dataset consists of <strong>16,000</strong> images that were collected from different sources:</p> <p>1)&nbsp;&nbsp;&nbsp; Chemical data images extracted from EP, US, and WO patents by OntoChem GmbH.</p> <p>2)&nbsp;&nbsp;&nbsp; Images from the MolScribe datasets <code><a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480">https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480</a></code></p> <p>3)&nbsp;&nbsp;&nbsp; DECIMER&ndash;hand-drawn molecule images dataset <code>H.O. Brinkhaus, A. Zielesny, C. Steinbeck, K. Rajan, &ldquo;DECIMER - hand-drawn molecule images dataset&rdquo;, 2022, Journal of Cheminformatics, 14, 36.&nbsp;<a href="https://doi.org/10.1186/s13321-022-00620-9">https://doi.org/10.1186/s13321-022-00620-9</a></code></p> <p>4)&nbsp;&nbsp;&nbsp; Images from the Rxnscribe training set <code>Y. Qian, J. Guo, Z. Tu, C.W. Coley, R. Barzilay, &ldquo;RxnScribe: A Sequence Generation Model for Reaction Diagram Parsing&rdquo;, 2023, &nbsp;arXiv:2305.11845v1, <a href="https://doi.org/10.48550/arXiv.2305.11845">https://doi.org/10.48550/arXiv.2305.11845</a>&nbsp;</code></p> <p>5)&nbsp;&nbsp;&nbsp; Formulas images from the im2latex-100k dataset <code>A prebuilt dataset for OpenAI's task for image-2-latex system,&nbsp;<a href="../record/56198#.YJjuCGZKgox">https://zenodo.org/record/56198#.YJjuCGZKgox</a> (accessed 16 Januar 2024)</code></p> <h1>Structure of dataset</h1> <p>The dataset consists of two directories:</p> <p>The "<strong>classified</strong>" directory contains manually labeled images. These images are divided into four distinct categories, with each category including 4000 images:</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; one_molecule</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; several_molecules</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; reactions</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; other</p> <p>In the &ldquo;<strong>for_model</strong>&rdquo; folder, we have split the images for training, validation, and testing in order to create a Chemical Image Classifier model:</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; training: 12,804 images</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; test: 1,604 images</p> <p>●&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; validation: 1,604 images.</p> <p>&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4