Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

4 results for “linguistic corpus”

Learn how ShareScore rates datasets ↗
zenodo36/100

VariaNTS corpus: A spoken Dutch corpus containing talker and linguistic variability

<p>The VariaNTS (Variatie in Nederlandse Taal en Sprekers) corpus is a Dutch spoken corpus that was developed to maximize both linguistic and talker variability. It contains 1000 stimulus materials from 11 linguistic subcategories, recorded by 8 male and 8 female native speakers of standard Dutch. The corpus contains audio recordings, orthographic transcriptions, stimulus-specific details such as word frequencies, neighborhood densities and phonotactic probabilities, and talker details. The VariaNTS corpus aims to provide new materials to be used for broad assessment of speech perception and word recognition in Dutch clinical and academic settings.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Corpus Linguistic Analysis of the BMSatire Descriptions corpus

<p>Data and documentation related to the analysis of the BMSatire Descriptions corpus, produced as part of the British Academy funded project &#39;<a href="https://curatorialvoice.github.io/">Curatorial Voice: legacy descriptions of art objects and their contemporary uses</a>&#39;.</p> <p>All data are derived from text written by M. Dorothy George and published between 1935 and 1954 as volumes 5 to 11 of the&nbsp;<em><a href="https://en.wikipedia.org/wiki/Catalogue_of_Political_and_Personal_Satires_Preserved_in_the_Department_of_Prints_and_Drawings_in_the_British_Museum">Catalogue of Political and Personal Satires Preserved in the Department of Prints and Drawings in the British Museum</a></em>. This text is published in lightly edited form by the British Museum via ResearchSpace as linked open data at&nbsp;<a href="https://public.researchspace.org/sparql">https://public.researchspace.org/sparql</a>. The data, text and images available via this service are published under a&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)</a>&nbsp;license (Research Space,&nbsp;<a href="https://public.researchspace.org/resource/Termsofuse">2016</a>; accessed&nbsp;<a href="https://www.zotero.org/jwbaker/items/itemKey/9LY9J5XY">10 September 2018</a>).</p> <p>Code is licensed under a&nbsp;<a href="https://github.com/CuratorialVoice/code/blob/master/LICENSE">GNU General Public License v3.0</a>.</p>

opencc-by-nc-sa-4.0Jun 2019View details →
zenodo32/100

Linguistic Features and Philosophical Functions of a Unique Rhyming Stylistic Genre: A Corpus-based Exploration of the Rhyming Translation of Dizi Gui

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo28/100

GENERAL CONSIDERATIONS ON CORPUS LINGUISTICS

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record