Images dataset for Chemical Images Classifier model
<h1>Original paper</h1> <p>The manually curated images dataset is a part of the Supplementary Materials of the paper: <code>A. Krasnov, S. Barnabas, T. Böhme, S. Boyer, L. Weber, Comparing software tools for optical chemical structure recognition, Digital Discovery (2024). <a href="https://doi.org/10.1039/D3DD00228D">https://doi.org/10.1039/D3DD00228D</a></code><br><br></p> <h1>Images dataset description</h1> <p>The dataset was used to generate the image classifier model. The dataset consists of <strong>16,000</strong> images that were collected from different sources:</p> <p>1) Chemical data images extracted from EP, US, and WO patents by OntoChem GmbH.</p> <p>2) Images from the MolScribe datasets <code><a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480">https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480</a></code></p> <p>3) DECIMER–hand-drawn molecule images dataset <code>H.O. Brinkhaus, A. Zielesny, C. Steinbeck, K. Rajan, “DECIMER - hand-drawn molecule images dataset”, 2022, Journal of Cheminformatics, 14, 36. <a href="https://doi.org/10.1186/s13321-022-00620-9">https://doi.org/10.1186/s13321-022-00620-9</a></code></p> <p>4) Images from the Rxnscribe training set <code>Y. Qian, J. Guo, Z. Tu, C.W. Coley, R. Barzilay, “RxnScribe: A Sequence Generation Model for Reaction Diagram Parsing”, 2023, arXiv:2305.11845v1, <a href="https://doi.org/10.48550/arXiv.2305.11845">https://doi.org/10.48550/arXiv.2305.11845</a> </code></p> <p>5) Formulas images from the im2latex-100k dataset <code>A prebuilt dataset for OpenAI's task for image-2-latex system, <a href="../record/56198#.YJjuCGZKgox">https://zenodo.org/record/56198#.YJjuCGZKgox</a> (accessed 16 Januar 2024)</code></p> <h1>Structure of dataset</h1> <p>The dataset consists of two directories:</p> <p>The "<strong>classified</strong>" directory contains manually labeled images. These images are divided into four distinct categories, with each category including 4000 images:</p> <p>● one_molecule</p> <p>● several_molecules</p> <p>● reactions</p> <p>● other</p> <p>In the “<strong>for_model</strong>” folder, we have split the images for training, validation, and testing in order to create a Chemical Image Classifier model:</p> <p>● training: 12,804 images</p> <p>● test: 1,604 images</p> <p>● validation: 1,604 images.</p> <p> </p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4