Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7 results for “code-switching”

Learn how ShareScore rates datasets ↗
zenodo40/100

Corpus of English and Nigerian Pidgin Code-switching (CENCOS)

<p>This dataset was compiled from a fieldwork in Nigeria in 2019. It features naturally occurring spoken conversations from educated speakers of English and Nigerian Pidgin in Nigeria, with very few conversations involving uneducated speakers. Nigeria is a multi-lingual nation with over 500 languages that are not mutually intelligible. English and Nigerian Pidgin serve as lingua francas used to bridge linguistic gaps between speakers whose languages are mutually unintelligible. English is used in both formal and informal settings, but Nigerian Pidgin is used only in informal settings. Nigerian Pidgin was formerly regarded as the language of the uneducated in Nigeria.&nbsp; Over time, it has developed into a language spoken not only by the uneducated, but also by the educated in Nigeria. The compilation of this corpus is an effort to understand how the educated speakers with the knowledge of both languages are able to use them in interactions. This corpus contains both sound and text files, but the sound files are not included here for data protection reasons. The sound files are manually transcribed into texts, amounting to over 100, 000 word tokens. &nbsp;It contains no annotation other than the speakers. An excel sheet containing speakers&rsquo; basic information like gender, age, ethnic group and education status is included. With these social factors, this corpus is useful for any form of investigation on the use of English and Nigerian Pidgin in Nigeria. &nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

TunSwitch: Code-Switched Tunisian Arabic Speech Dataset

<p>We developed a tool for collecting Tunisian dialect data, prompting users to record themselves reading provided phrases. We sourced sentences from Tunisiya.&nbsp;These sentences are consequently removed from the LM training corpus. 89 persons have participated leading to the collection of 2631 distinct phrases. This set will be called TunSwitch TO, ``TO&quot; standing for Tunisian Only, as these sentences do not have non-Tunisian words.&nbsp;</p> <p>In response to the limited availability of paired Text-Speech Tunisian datasets with &nbsp;code-switching, we have built a &nbsp;corpus through meticulous manual annotation. Whenever encountered, French and English &nbsp;words are enclosed &nbsp;within &quot;&lt;&gt;&quot;&nbsp;tags, and left Tunisian words without any enclosing tags. While these tags have not been used in the proposed models, they allow to have language-usage statistics &nbsp;and may be useful for further approaches handling code-switching. The resulting set is released as TunSwitch CS, ``CS&quot; standing for Code-Switched.</p> <p>The TunSwitch CS dataset samples come from a set of radio shows and podcasts, representing diverse topics and a large number of unique speakers. The audio are first segmented into chunks, prioritizing word integrity using the WebRTC-VAD algorithm for silence detection. Afterward, we used a Pyannote overlap detection model to remove overlapping speech sections. Then, a music detection model is employed to eliminate music-containing chunks that could disrupt ASR model accuracy.&nbsp;<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

TunSwitch: Code-Switched Tunisian Arabic Speech Dataset

<p>This folder contains the data used to develop and test the Tunisian Arabic Automatic Speech Recognition model developed in the following paper :</p> <p>A. A. Ben Abdallah*, A. Kabboudi, A. Kanoun, and S. Zaiem*, &ldquo;Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition&rdquo;, Submitted to ICASSP 2024, vol. * : These two authors have contributed equally. 2023.</p> <p><br> It contains 4 zipped folders containing audio data :<br> - TunSwitchCS.zip : containing annotated code-switched data.<br> - TunSwitchTO.zip : containing annotated Tunisian-Only data.<br> - weakly_labeled_tn.zip : containing weakly-labeled (or unlabeled) audio data. Audios may contain code-switching, but the current weak labels do not.<br> - test_wavs.zip : contains annotated testing data, divided between a code-switched part and a tunisian-only part.</p> <p><br> It also contains textual data, used for language modelling, contained in TextData.zip. Finally it also contains a language-detailed annotation of TunSwitchCS in the&nbsp; language_annotation.zip file&nbsp;.</p> <p>More details about the data are available in the paper. The current table are in a SpeechBrain-friendly format, the column path is irrelevant and has to be changed according to your local setting. Please use the provided train-dev-test splits if you work with this dataset.</p> <p>Please cite the aforementioned paper if you use or refer to this dataset. You can find models trained and tested on this dataset <a href="https://huggingface.co/SalahZa">Here</a>. Space demos are also available.&nbsp;</p> <p>If you use or refer to this dataset, please cite :&nbsp;</p> <p>```</p> <p>@misc{abdallah2023leveraging,<br> &nbsp; &nbsp; &nbsp; title={Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition},&nbsp;<br> &nbsp; &nbsp; &nbsp; author={Ahmed Amine Ben Abdallah and Ata Kabboudi and Amir Kanoun and Salah Zaiem},<br> &nbsp; &nbsp; &nbsp; year={2023},<br> &nbsp; &nbsp; &nbsp; eprint={2309.11327},<br> &nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br> &nbsp; &nbsp; &nbsp; primaryClass={eess.AS}<br> }</p> <p>```</p> <p><br> &nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

Code-Switching Speech Corpus

<p><strong>German-English Code-Switching speech dataset</strong></p> <p>We provide means to resegment a subset of the German **Spoken Wikipedia Corpus** (SWC) enabling a particular focus on code-switching.&nbsp; This results in the German-English code-switching corpus, a 34h transcribed speech corpus of read Wikipedia articles which can be used as a benchmark for research on code-switching.&nbsp; The articles are read by a large and diverse group of people. The SWC is perhaps the largest corpus of freely-available aligned speech for German.&nbsp; It contains 1014 spoken articles read by more than 350 identified speakers comprising 386h of speech. This corpus is available at http://nats.gitlab.io/swc.</p> <p>In SWC, since most of the articles are long, the recordings submitted by the volunteers are also long (&sim;54min) on average.&nbsp; These audio files are manually annotated at word-level and also segment level in XML format.&nbsp; We use a language identification tool to detect code-switching in the transcription of the audio files with consecutive indices. To extract intra-sentential code-switching segments, we ensure that the detected code-switching is preceded and followed by German words or sentences. The final set consists of 34h of speech data and 12,437 code-switching segments (in Kaldi ASR toolkit data format).</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>@article{baumann2019spoken,<br> &nbsp; title={The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening},<br> &nbsp; author={Baumann, Timo and K{\&quot;o}hn, Arne and Hennig, Felix},<br> &nbsp; journal={Language Resources and Evaluation},<br> &nbsp; volume={53},<br> &nbsp; number={2},<br> &nbsp; pages={303--329},<br> &nbsp; year={2019},<br> &nbsp; publisher={Springer}<br> }</p> <p>@article{grave2018learning,<br> &nbsp; title={Learning word vectors for 157 languages},<br> &nbsp; author={Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas},<br> &nbsp; journal={arXiv preprint arXiv:1802.06893},<br> &nbsp; year={2018}<br> }</p> <p>&nbsp;</p>

opencc-by-sa-3.0Jan 2021View details →
zenodo16/100

SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET

<p>This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech type,the target and language.</p> <p>The dataset can be accessed upon request from the authors via email addresses:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-nodhianbo@gmail.com</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-endeshanelly@gmail.com</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;-mscci00083@student.maseno.ac.ke</p>

restrictedcc-by-4.0Nov 2024View details →
zenodo16/100

The TongueSwitcher Corpus of German-English Code-Switching

<p>This is the TongueSwitcher Corpus of German-English code-switching tweets. Included are the train and dev sets with automatic word (and subword for mixed words) language identification, alongside the human-annotated corpus with test and interlingual homograph sets.</p> <p>BibTeX entry and citation info</p> <p>@inproceedings{sterner2023tongueswitcher,<br>&nbsp; &nbsp;author &nbsp; &nbsp; = {Igor Sterner and Simone Teufel},<br>&nbsp; &nbsp;title &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= {TongueSwitcher: Fine-Grained Identification of German-English Code-Switching},<br>&nbsp; &nbsp;booktitle &nbsp;= {Sixth Workshop on Computational Approaches to Linguistic Code-Switching},<br>&nbsp; &nbsp;publisher = {Empirical Methods in Natural Language Processing},<br>&nbsp; &nbsp;year &nbsp; &nbsp; &nbsp; &nbsp;= {2023},<br>}</p>

restrictedcc-by-4.0Oct 2023View details →
zenodo4/100

cantonese loanwords and code-switching

<p>VIDEO</p>

restrictedNov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record