Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7 results for “language identification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Image-based Many-language Programming Language Identification - Replication Package

<p>This dataset contains the data, software, and instructions&nbsp;needed to replicate the findings of the paper:</p> <p>Francesca Del Bonifro, Maurizio Gabbrielli, Antonio Lategano, and Stefano Zacchiroli.&nbsp;Image-based Many-language<br> Programming Language Identification. <a href="https://peerj.com/computer-science/"><em>PeerJ Computer Science</em></a>, 2021 (to appear).&nbsp;DOI:&nbsp;<a href="https://dx.doi.org/10.7717/peerj-cs.631">10.7717/peerj-cs.631</a></p> <p>After retrieving the full dataset, extract the replication-package.zip archive&nbsp;and follow the instructions described in the README.md&nbsp;file.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)

<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab:&nbsp;<a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is&nbsp;<a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Datasets for "Foreign Language Usage and National and European Identification in the Netherland"

<p>Two datasets that correspond with the article &quot;Foreign Language Usage and National and European Identification in the Netherlands&quot;, currently accepted for publication and Journal of Language and Social Psychology (November 2020).&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo28/100

Youtube-Dataset for Language Identification in Speech Signals

<p><strong>Youtube-Dataset for Language Identification in Speech Signals</strong></p> <p>- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de</p> <p><strong>Reference</strong></p> <p>In case you use this dataset for your research, please cite</p> <p>Alexandra Draghici, Jakob Abe&szlig;er &amp; Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020</p> <p><strong>Dataset</strong></p> <p>The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo.</p> <p>- 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders &quot;0&quot; - &quot;5&quot; encode the language id:<br> &nbsp; 0 - English<br> &nbsp; 1 - French<br> &nbsp; 2 - German<br> &nbsp; 3 - Greek<br> &nbsp; 4 - Italian<br> &nbsp; 5 - Spanish</p> <p><strong>Audio Processing</strong></p> <p>- mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br> &nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

WiLI-2018 - Wikipedia Language Identification database

<p>WiLI-2018, the Wikipedia language identification benchmark dataset, contains 235000 paragraphs&nbsp;of 235&nbsp;languages. The dataset is balanced and a train-test split is provided.<br> <br> See &quot;The WiLI benchmark dataset for written language identification&quot;&nbsp;paper (soon on arXiv) for more information.</p>

openodc-odblJan 2018View details →
zenodo24/100

Multilingual test set for language identification and speech recognition from European Parliament recordings

<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings.&nbsp;</p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes.&nbsp;</p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>

opencc-zeroJul 2024View details →
ClinicalTrials.gov24/100

Genetic Biomarkers of Child Language Development in Taiwan: an Identification and Validation Study

ClinicalTrials.gov study NCT05504564. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record