Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2 results for “historical corpora”

Learn how ShareScore rates datasets ↗
zenodo40/100

Annotated Corpora of Historical Catalan (HisCat) - Llibre dels Fets

<p>This repository is part of the Annotated Corpora of Historical Catalan (HisCat). It contains the first POS-tagged text that is partially manually corrected and used to train Old Catalan POS taggers, described in the following paper:</p> <p>Meelen, Marieke &amp; Pujol i Campeny, Afra, (2021) &#39;Old Catalan Morphosyntax: developing an annotated corpus&#39; in <em>Journal of Open Humanities Data</em>.</p> <p>This POS-tagged text is the 13th century <em>Llibre dels Fets</em>, a historical chronicle. The version of the text used for this project is</p> <p>Bruguera, J. (1991). <em>El Llibre dels Fets del Rei en Jaume</em>. Barcelona: Barcino.</p> <p>as prepared for the <em>Corpus Informatitzat del Catal&agrave; Antic</em></p> <p>Torruella, J., P&eacute;rez Saldanya, M., &amp; Martines, J. (2009). <em>Corpus Informatitzat del Catal&agrave; Antic</em>. URL: <a href="http://cica.cat/">http://cica.cat/</a>.</p> <p>The subcorpus counts with 164,096 POS-annotated tokens (165,538 tokens including punctuation and folio markers), of which 60,000 have been manually corrected. This subcorpus contains a total of and 4,506 main clauses. POS tagging of this text was done with the Memory-Based Tagger by TiMBL (<a href="https://languagemachines.github.io/mbt/">https://languagemachines.github.io/mbt/</a>). The code accompanying the paper can be found on GitHub: <a href="https://github.com/lothelanor/catalancorpora">https://github.com/lothelanor/catalancorpora</a>). In addition to memory-based tagging, have tried neural-based tagging with TARGER (<a href="https://github.com/achernodub/targer">https://github.com/achernodub/targer</a>) for which we created word embeddings that can be found on <a href="https://doi.org/10.5281/zenodo.5615556">Zenodo</a>. Results for memory-based tagging were better, however, which is why this version is uploaded here.</p> <pre>&nbsp;</pre>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Metadata for Historical Corpora. Realization of the Metamodel for Corpus Metadata with the help of TEI Customization

<p>TEI ODD Customization for the documentation of historical corpora:</p> <p>The TEI ODD customizations map the Metamodel for Corpus Metadata (MCM) to a TEI p5 header structure for each of the objects of the classes &#39;Corpus&#39;, &#39;Document&#39; and &#39;Preparation&#39;. The MCM is realized with a subset of the TEI p5 guidelines.</p> <p>Each ODD contains further information and explanations regarding the MCM and the customization of the TEI. Additionally, for each ODD, an HTML documentation is provided.</p>

opencc-by-4.0Feb 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record