Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

237

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

237 results for “OCR”

Learn how ShareScore rates datasets ↗
zenodo48/100

CoAID dataset with multiple extracted features (both sparse and dense) and degraded by OCR

<p>This is the same datasets as:</p> <p>Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). CoAID dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630710</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Event Registry dataset with multiple extracted features (both sparse and dense) and degraded by OCR

<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630367</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry dataset texts with OCR degradations and synthesised segmentation (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6631305</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Event Registry titles dataset with multiple extracted features (both sparse and dense) and degraded by OCR

<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry titles dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630828</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

FibVid dataset with multiple extracted features (both sparse and dense) and degraded by OCR

<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Fibvid dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630409</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). FibVid dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630758</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

ICDAR'15 SMARTPHONE DOCUMENT CAPTURE AND OCR COMPETITION (SmartDoc) - Challenge 1 (original version)

<p><strong>CHALLENGE 1: SMARTPHONE DOCUMENT CAPTURE COMPETITION</strong></p> <p><strong>Smartphones are replacing personal scanners.</strong>&nbsp;They are portable, connected, powerful and affordable. They are on their way to become the new entry point in business processing applications like document archival, ID scanning, check digitization, just to name a few. In order keep our workflows streamlined,&nbsp;<strong>we need to make those new capture device as reliable as batch scanners</strong>.</p> <p>We believe an efficient capture process should be able to:</p> <ol> <li><em>detect and segment</em>&nbsp;the relevant document object during the preview phase;</li> <li><em>assess the quality</em>&nbsp;of the capture conditions and help the user improve them;</li> <li>optionally,&nbsp;<em>trigger the capture</em>&nbsp;at the perfect moment;</li> <li>and&nbsp;<em>produce a high-quality, controlled output</em>&nbsp;based on the high resolution captured image.</li> </ol> <p>This competition is focused on the first step of this process:<strong>&nbsp;</strong><strong>efficiently detect and segment document regions</strong>, as illustrated by following video showing the ideal output for the preview phase of some acquisition session: <a href="https://youtu.be/WNsI0R_rpO0">Click here to watch the video.</a> This video shows the ideal document object detection &lrm;(ie the ground truth, as a red frame)&lrm;.</p> <p>For this challenge, the <strong>input</strong> consists in a set of <strong>videoclips containing a document</strong> from a predefined set, and the <strong>output</strong> should be an <strong>xml file containing the quadrilateral coordinates</strong> in which we can find the document per each frame of the video. Click <a href="https://sites.google.com/site/icdar15smartdoc/challenge-1/challenge1dataset">here</a> for detailed information about the dataset.&nbsp;</p> <p>&nbsp;</p> <p><strong>Licence</strong> for the dataset of challenge 1 (page outline detection in preview frames) :</p> <p>This work is licensed under a <strong>Creative Commons Attribution 4.0 International License</strong> &lt;<a href="https://www.google.com/url?q=http://creativecommons.org/licenses/by/4.0/&amp;sa=D&amp;ust=1524734857667000&amp;usg=AFQjCNEt4YXnUv2nXCFwkeOuBDqxDpvknQ">http://creativecommons.org/licenses/by/4.0/</a>&gt;. Author attribution should be given by citing the following conference paper: Jean-Christophe Burie, Joseph Chazalon, Micka&euml;l Coustaty, S&eacute;bastien Eskenazi, Muhammad Muzzamil Luqman, Maroua Mehri, Nibal Nayef, Jean-Marc OGIER, Sophea Prum and Mar&ccedil;al Rusinol: &ldquo;ICDAR2015 Competition on Smartphone Document Capture and OCR (SmartDoc)&rdquo;, In 13th International Conference on Document Analysis and Recognition (ICDAR), 2015.</p> <p><strong>If you use this dataset, please send us a short email at &lt;icdar.smartdoc (at) gmail.com&gt; to tell us why it was useful to you, and whether you have results or publications we can reference on our website. Thank you!</strong></p>

opencc-by-4.0Aug 2015View details →
zenodo44/100

OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)

<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl &nbsp; &nbsp;&nbsp; entropy rate per document page</p> <p>corpus-language.pkl&nbsp;&nbsp; language per document page</p> <p>corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Selection list of German documents</p> <p>xml2csv_alto.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; fulltext corpus per document page (incl.OCR word confidences)</p> <p>&nbsp;</p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL &rsquo;12, pages 25&ndash;30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Pathway Figure OCR GMT

<p>Sample Pathway Figure OCR results from 17 December 2018, condensed by PMCID into GMT format for use in gene set analysis. Columns are PMCID, URL, Entrez Gene IDs</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

CT-OCR-2022 (scroll001)

<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source&nbsp;paper document, 400 X-ray projections, 2687 CT-reconstructed cross-sections and&nbsp; segmentation markups for 6 model objects.</p> <p>Description of the data for each model object is presented in the table.</p> <table> <tbody><tr> <th>Files</th> <th>Data description</th> </tr> </tbody><tbody> <tr> <td>2022-ICMV-CT-OCR.[package].png</td> <td>sample projection and slice visualization</td> </tr> <tr> <td>[package].proj_src/proj_s_.log</td> <td>X-ray measurement log file</td> </tr> <tr> <td>[package].proj_src/*.tif</td> <td>preprocessed projections before rotation axe correction</td> </tr> <tr> <td>2022-ICMV-CT-OCR.[package].png</td> <td>package single projection and slice vizualization</td> </tr> <tr> <td>[package].proj_src/*.tif</td> <td>preprocessed projections before rotation axe correction</td> </tr> <tr> <td>[package].proj_norm/proj_s_.log</td> <td>X-ray measurement and geometry correction log file</td> </tr> <tr> <td>[package].proj_norm/*.tif</td> <td>preprocessed projections after rotation axe correction</td> </tr> <tr> <td>[package].rec_XXXX/metadata.json</td> <td>reconstruction metadata</td> </tr> <tr> <td>[package].rec_XXXX/*.tif</td> <td>CT-reconstructed volume, slices with size XXXX&times;XXXX</td> </tr> <tr> <td>[package].seg_XXXX/*.tif.seg.png</td> <td>segmentation markup</td> </tr> <tr> <td>[package].blank.png</td> <td>sample croped from pdf</td> </tr> <tr> <td>[package].scan.png</td> <td>sample croped from scanned image</td> </tr> </tbody> </table> <p>Due to the large amount of data, folders were packed into multi-volume zip-archives. Dataset published in Zenodo service in several linked repositories.</p> <p>scroll01 - <a href="https://doi.org/10.5281/zenodo.7123495">10.5281/zenodo.7123495</a><br>scroll02 - <a href="https://doi.org/10.5281/zenodo.7157600">10.5281/zenodo.7157600</a><br>scroll03 - <a href="https://doi.org/10.5281/zenodo.7157610">10.5281/zenodo.7157610</a><br>scroll04 - <a href="https://doi.org/10.5281/zenodo.7161350">10.5281/zenodo.7161350</a><br>folded01 - <a href="https://doi.org/10.5281/zenodo.7162001">10.5281/zenodo.7162001</a>,&nbsp;<a href="https://doi.org/10.5281/zenodo.7164152">10.5281/zenodo.7164152</a><br>folded02 - <a href="https://doi.org/10.5281/zenodo.7267141">10.5281/zenodo.7267141</a>, <a href="https://doi.org/10.5281/zenodo.7272064">10.5281/zenodo.7272064</a></p> <p>Any questions, complaints, etc. can be directed to: polevoy@smartengines.com (Dmitry Polevoy)</p> <p><strong>Share and Cite:</strong></p> <p>D. V. Polevoy, P. A. Kulagin, A. S. Ingacheva, Zh. V. Soldatova, M. V. Chukalina, D. P. Nikolaev, V. V. Arlazarov, "From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligence," Proc. SPIE 12701, Fifteenth International Conference on Machine Vision (ICMV 2022), 127010P (7 June 2023); https://doi.org/10.1117/12.2680132</p> <p>in BibTex format:</p> <p>@inproceedings{10.1117/12.2680132,<br>author = {D. V. Polevoy and P. A. Kulagin and A. S. Ingacheva and Zh. V. Soldatova and M. V. Chukalina and D. P. Nikolaev and V. V. Arlazarov},<br>title = {{From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligence}},<br>volume = {12701},<br>booktitle = {Fifteenth International Conference on Machine Vision (ICMV 2022)},<br>editor = {Wolfgang Osten and Dmitry P. Nikolaev and Jianhong (Jessica) Zhou},<br>organization = {International Society for Optics and Photonics},<br>publisher = {SPIE},<br>pages = {127010P},<br>keywords = {virtual unrolling, virtual unwrapping, digital unfolding, computational flattening, computed tomography, non-destructive analysis, open dataset},<br>year = {2023},<br>doi = {10.1117/12.2680132},<br>URL = {https://doi.org/10.1117/12.2680132}<br>}</p> <p><strong>See also</strong></p> <p>P. A. Kulagin, D. V. Polevoy, M. V. Chukalina, D. P. Nikolaev and V. V. Arlazarov, &ldquo;Fully automatic virtual unwrapping method for documents imaged by X-ray tomography,&rdquo; Proc. ICDAR 2024, to be published.</p> <p><a href="https://github.com/SmartEngines/virtual-unwrapping-article-code">https://github.com/SmartEngines/virtual-unwrapping-article-code</a></p>

opencc-by-2.5Sep 2022View details →
zenodo44/100

Drinov Orthography for Post-OCR Correction dataset

<p>The Drinov Orthography for Post-OCR Correction (DOPOC) dataset was created by annotating a historical newspaper collection provided by the <a href="https://digital.libplovdiv.com/en">National Library "Ivan Vazov"</a> (NLIV) in Plovdiv, Bulgaria. We consider printed versions of these documents, which we manually annotate and align at the character level in the same format as the one from the ICDAR 2019 post-OCR correction competition.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CT-OCR-2022 (folded002, part2)

<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source&nbsp;paper document, X-ray projections, CT-reconstructed cross-sections and&nbsp; segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference&nbsp;to the dataset&nbsp;see at&nbsp;<a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>

opencc-by-2.5Oct 2022View details →
zenodo44/100

CT-OCR-2022 (scroll002)

<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source&nbsp;paper document, X-ray projections, CT-reconstructed cross-sections and&nbsp; segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference&nbsp;to the dataset&nbsp;see at&nbsp;<a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>

opencc-by-2.5Oct 2022View details →
zenodo44/100

CT-OCR-2022 (folded002, part1)

<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source&nbsp;paper document, X-ray projections, CT-reconstructed cross-sections and&nbsp; segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference&nbsp;to the dataset&nbsp;see at&nbsp;<a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>

opencc-by-2.5Oct 2022View details →
zenodo44/100

CT-OCR-2022 (folded001, part2)

<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source&nbsp;paper document, X-ray projections, CT-reconstructed cross-sections and&nbsp; segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference&nbsp;to the dataset&nbsp;see at&nbsp;<a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>

opencc-by-2.5Oct 2022View details →
zenodo40/100

OCRed text of the Allgemeine musikalische Zeitung (ALTO format) from 1798-1848 and 1863-1965

<p><span>Dataset as used in the LREC 2020 publication &#39;<em>Allgemeine Musikalische Zeitung</em></span><span> as a Searchable Online Corpus&#39;</span>.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

NewsEye / READ OCR training dataset from French Newspapers (18th, 19th, early 20th C.)

<p>The dataset comprises French newspaper pages from 18th, 19th and early 20th century with carefully corrected text. The page images were provided by the&nbsp;<a href="https://www.bnf.fr/en">French National Library</a> and comprise 127 pages (training set) and 8 pages (validation set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

CIS OCR Workshop v1.0: OCR and postcorrection of early printings for digital humanities

<p>The 2-day CIS OCR Workshop on &quot;OCR and postcorrection of early printings for digital humanities&quot; originally held at LMU, Munich 14/15 September 2015 (see http://www.cis.lmu.de/ocrworkshop).</p> <p>Release date: 2016-02-25</p> <p><br /> CIS OCR Workshop by Uwe Springmann, Florian Fink is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p>

opencc-by-nc-sa-4.0Feb 2016View details →
zenodo40/100

Castren 1844: Elementa grammaticae Syrjaenae, OCR Ground Truth

<p>Matthias Alexander Castr&eacute;n published&nbsp;Elementa grammaticae Syrjaenae in 1844. This dataset contains two scans of this work, the higher quality version originating from the Internet Archive:&nbsp;</p> <p>https://archive.org/details/elementagrammati00cast/mode/2up</p> <p>All pages are layout detected with Transkribus, and there are 26 proofread pages. Three pages contain table layouts.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Tesseract OCR of IIT-CDIP Dataset

<p>This is&nbsp;Tesseract generated&nbsp;<strong>transcriptions (no&nbsp;images)</strong>&nbsp;of (most of) the IIT-CDIP dataset. To download the images of the IIT-CDIP dataset go to&nbsp;<a href="https://data.nist.gov/od/id/mds2-2531">https://data.nist.gov/od/id/mds2-2531</a>&nbsp;</p> <p>The directory struture of this dataset is the same as the IIT-CDIP dataset (although has everything in one tar, with &quot;a.a&quot;, &quot;a.b&quot;, ... directories)&nbsp;and can thus be combine with the image&nbsp;IIT-CDIP dataset&nbsp;using rsync or similar tool. This dataset&nbsp;contains a &quot;X.layout.json&quot; for each &quot;X.png&quot; in the IIT-CDIP dataset (doesn&#39;t have sections &#39;a&#39;, &#39;w&#39;, &#39;x&#39;, &#39;y&#39;, and &#39;z&#39;).</p> <p>The jsons contain block/paragraph, line and word bounding boxes, with&nbsp;transcriptions for the words following the Tesseract format. The line and word annotations are directly taken from Tesseract. The block and paragraph output of Tesseract was discarded. The images were then run through both the&nbsp;Publaynet&nbsp;and PrimaNet&nbsp;models available on LayoutParser (<a href="https://layout-parser.github.io/">https://layout-parser.github.io/</a>). The combine output of these models became the block/paragraph annotations (we kept the Tesseract output format, but each block has 1 paragraph of exactly the same shape).</p> <p><strong>Important:</strong> There is also a &quot;rotation&quot; value in the json (0, 90, 180, or 270) indicating the json may be&nbsp;for a rotated version of the IIT-CDIP image by the given amount (attempted to rotated documents to upright position to get better OCR results).</p> <p>These are the annotations used to pre-train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>).</p> <p>These annotations will be worse than&nbsp;those that would be obtained&nbsp;using&nbsp;a commercial OCR system&nbsp;(like those used to pre-train LayoutLMv2/v3).</p> <p>The code used to produce these&nbsp;annotations is available here:&nbsp;<a href="https://github.com/herobd/ocr">https://github.com/herobd/ocr</a></p>

opencc-by-4.0May 2022View details →
zenodo40/100

OCR model for Pracalit for Sanskrit and Newar MSS 16th to 19th C., Ground Truth

<p>Ground truth data (png and xml files) for a an OCR model. Will be continually updated.</p> <p>Originally trained on Transkribus with a&nbsp;PyLaia model created from ground truth data based on transcripts into Pracalit Unicode of four Nepalese manuscripts. The manuscripts used to create this model are Staatsbibliothek zu Berlin&#39;s Hitopadeśa (MIK I 4851) (mixed Newar and Sanskrit dating to 1561) and Vetālapa&ntilde;caviṃśati (HS. Or. 6414) (Newar dating to 1675) as well as Cambridge Digital Library&#39;s Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) (Sanskrit, 18th century) and the Royal Asiatic Society Online Collection&#39;s Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) (Newar and Sanskrit dating to c. 1800).</p> <p>The training was done on 441 pages and validation on 242 pages.</p> <p>This model does not recognise spacing, except for large gaps (i.e. for pictures or string holes). Newar word divider markers may not be represented or may be transcribed as virama. In general, the model is made for MSS with scriptio continua and will transcribe into scriptio continua into Pracalit Unicode.</p> <p>Transcription was performed by Dr Alexander O&#39;Neill (SOAS University of London). Transcription of the Vetālapa&ntilde;caviṃśati (HS. Or. 6414) and Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) was aided by unpublished materials provided by Dr Felix Otter (Philipps-Universit&auml;t Marburg), as well as the published transcription in Shakya, Min Bahadur, and Shanta Harsha Bajracharya, eds. &quot;Svayambhū Purāṇa.&quot; Lalitpur: Nagarjuna Institute of Exact Methods, 2001. The transcription of Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) was aided by the transcription provided by the Digital Sanskrit Buddhist Canon Project based on Lokesh Chandra, &quot;Guṇakāraṇḍavyūhasūtram,&quot; New Delhi: International Academy of Indian Culture, 1999.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin

<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>

opencc-by-4.0Aug 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record