Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.7.1
Dataset results
7 results for “language identification”
Image-based Many-language Programming Language Identification - Replication Package
<p>This dataset contains the data, software, and instructions needed to replicate the findings of the paper:</p> <p>Francesca Del Bonifro, Maurizio Gabbrielli, Antonio Lategano, and Stefano Zacchiroli. Image-based Many-language<br> Programming Language Identification. <a href="https://peerj.com/computer-science/"><em>PeerJ Computer Science</em></a>, 2021 (to appear). DOI: <a href="https://dx.doi.org/10.7717/peerj-cs.631">10.7717/peerj-cs.631</a></p> <p>After retrieving the full dataset, extract the replication-package.zip archive and follow the instructions described in the README.md file.</p>
SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)
<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab: <a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is <a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>
Datasets for "Foreign Language Usage and National and European Identification in the Netherland"
<p>Two datasets that correspond with the article "Foreign Language Usage and National and European Identification in the Netherlands", currently accepted for publication and Journal of Language and Social Psychology (November 2020). </p>
Youtube-Dataset for Language Identification in Speech Signals
<p><strong>Youtube-Dataset for Language Identification in Speech Signals</strong></p> <p>- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de</p> <p><strong>Reference</strong></p> <p>In case you use this dataset for your research, please cite</p> <p>Alexandra Draghici, Jakob Abeßer & Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020</p> <p><strong>Dataset</strong></p> <p>The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo.</p> <p>- 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders "0" - "5" encode the language id:<br> 0 - English<br> 1 - French<br> 2 - German<br> 3 - Greek<br> 4 - Italian<br> 5 - Spanish</p> <p><strong>Audio Processing</strong></p> <p>- mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br> </p>
WiLI-2018 - Wikipedia Language Identification database
<p>WiLI-2018, the Wikipedia language identification benchmark dataset, contains 235000 paragraphs of 235 languages. The dataset is balanced and a train-test split is provided.<br> <br> See "The WiLI benchmark dataset for written language identification" paper (soon on arXiv) for more information.</p>
Multilingual test set for language identification and speech recognition from European Parliament recordings
<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings. </p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes. </p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>
Genetic Biomarkers of Child Language Development in Taiwan: an Identification and Validation Study
ClinicalTrials.gov study NCT05504564. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.