Skip to main content
zenodoopen

MER dataset im2latexv2 - Part 1

<h1>Mathematical Expression Recognition Dataset&nbsp;im2latexv2 - Part 1</h1> <p>This repository contains Part 1 of the im2latexv2 dataset presented in the paper <em>MathNet: A Data-Centric Approach for Printed Mathematical Expression Recognition.</em></p> <p>The dataset is an enhanced version of the im2latex-100k dataset. It uses a novel LaTeX normalization process and 61 rendering environments to make the dataset more realistic.</p> <p>Please also download <a href="../records/11296280">Part 2</a> of the im2latexv2 dataset (doi: 10.5281/zenodo.11296280) and copy the subfolders in the folder of Part 1.</p> <p>To unpack all images, please use the unpack_im2latexv2.py script.<br><br>The CSV files have the following structure:</p> <table> <tbody> <tr> <td>formula</td> <td>images</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>tokenized formula (tokens separated by white spaces)</td> <td>path to image with rendering env 1</td> <td>path to image with rendering env 2</td> <td>....</td> </tr> </tbody> </table>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0