MER dataset im2latexv2 - Part 1
<h1>Mathematical Expression Recognition Dataset im2latexv2 - Part 1</h1> <p>This repository contains Part 1 of the im2latexv2 dataset presented in the paper <em>MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.</em></p> <p>The dataset is an enhanced version of the im2latex-100k dataset. It uses a novel LaTeX normalization process and 61 rendering environments to make the dataset more realistic.</p> <p>Please also download <a href="../records/11296280">Part 2</a> of the im2latexv2 dataset (doi: 10.5281/zenodo.11296280) and copy the subfolders in the folder of Part 1.</p> <p>To unpack all images, please use the unpack_im2latexv2.py script.<br><br>The CSV files have the following structure:</p> <table> <tbody> <tr> <td>formula</td> <td>images</td> <td> </td> <td> </td> </tr> <tr> <td>tokenized formula (tokens separated by white spaces)</td> <td>path to image with rendering env 1</td> <td>path to image with rendering env 2</td> <td>....</td> </tr> </tbody> </table>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0