Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.7.1
Dataset results
7 results for “code-switching”
Corpus of English and Nigerian Pidgin Code-switching (CENCOS)
<p>This dataset was compiled from a fieldwork in Nigeria in 2019. It features naturally occurring spoken conversations from educated speakers of English and Nigerian Pidgin in Nigeria, with very few conversations involving uneducated speakers. Nigeria is a multi-lingual nation with over 500 languages that are not mutually intelligible. English and Nigerian Pidgin serve as lingua francas used to bridge linguistic gaps between speakers whose languages are mutually unintelligible. English is used in both formal and informal settings, but Nigerian Pidgin is used only in informal settings. Nigerian Pidgin was formerly regarded as the language of the uneducated in Nigeria. Over time, it has developed into a language spoken not only by the uneducated, but also by the educated in Nigeria. The compilation of this corpus is an effort to understand how the educated speakers with the knowledge of both languages are able to use them in interactions. This corpus contains both sound and text files, but the sound files are not included here for data protection reasons. The sound files are manually transcribed into texts, amounting to over 100, 000 word tokens. It contains no annotation other than the speakers. An excel sheet containing speakers’ basic information like gender, age, ethnic group and education status is included. With these social factors, this corpus is useful for any form of investigation on the use of English and Nigerian Pidgin in Nigeria. </p>
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>We developed a tool for collecting Tunisian dialect data, prompting users to record themselves reading provided phrases. We sourced sentences from Tunisiya. These sentences are consequently removed from the LM training corpus. 89 persons have participated leading to the collection of 2631 distinct phrases. This set will be called TunSwitch TO, ``TO" standing for Tunisian Only, as these sentences do not have non-Tunisian words. </p> <p>In response to the limited availability of paired Text-Speech Tunisian datasets with code-switching, we have built a corpus through meticulous manual annotation. Whenever encountered, French and English words are enclosed within "<>" tags, and left Tunisian words without any enclosing tags. While these tags have not been used in the proposed models, they allow to have language-usage statistics and may be useful for further approaches handling code-switching. The resulting set is released as TunSwitch CS, ``CS" standing for Code-Switched.</p> <p>The TunSwitch CS dataset samples come from a set of radio shows and podcasts, representing diverse topics and a large number of unique speakers. The audio are first segmented into chunks, prioritizing word integrity using the WebRTC-VAD algorithm for silence detection. Afterward, we used a Pyannote overlap detection model to remove overlapping speech sections. Then, a music detection model is employed to eliminate music-containing chunks that could disrupt ASR model accuracy. <br> </p> <p> </p>
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>This folder contains the data used to develop and test the Tunisian Arabic Automatic Speech Recognition model developed in the following paper :</p> <p>A. A. Ben Abdallah*, A. Kabboudi, A. Kanoun, and S. Zaiem*, “Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition”, Submitted to ICASSP 2024, vol. * : These two authors have contributed equally. 2023.</p> <p><br> It contains 4 zipped folders containing audio data :<br> - TunSwitchCS.zip : containing annotated code-switched data.<br> - TunSwitchTO.zip : containing annotated Tunisian-Only data.<br> - weakly_labeled_tn.zip : containing weakly-labeled (or unlabeled) audio data. Audios may contain code-switching, but the current weak labels do not.<br> - test_wavs.zip : contains annotated testing data, divided between a code-switched part and a tunisian-only part.</p> <p><br> It also contains textual data, used for language modelling, contained in TextData.zip. Finally it also contains a language-detailed annotation of TunSwitchCS in the language_annotation.zip file .</p> <p>More details about the data are available in the paper. The current table are in a SpeechBrain-friendly format, the column path is irrelevant and has to be changed according to your local setting. Please use the provided train-dev-test splits if you work with this dataset.</p> <p>Please cite the aforementioned paper if you use or refer to this dataset. You can find models trained and tested on this dataset <a href="https://huggingface.co/SalahZa">Here</a>. Space demos are also available. </p> <p>If you use or refer to this dataset, please cite : </p> <p>```</p> <p>@misc{abdallah2023leveraging,<br> title={Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition}, <br> author={Ahmed Amine Ben Abdallah and Ata Kabboudi and Amir Kanoun and Salah Zaiem},<br> year={2023},<br> eprint={2309.11327},<br> archivePrefix={arXiv},<br> primaryClass={eess.AS}<br> }</p> <p>```</p> <p><br> </p>
Code-Switching Speech Corpus
<p><strong>German-English Code-Switching speech dataset</strong></p> <p>We provide means to resegment a subset of the German **Spoken Wikipedia Corpus** (SWC) enabling a particular focus on code-switching. This results in the German-English code-switching corpus, a 34h transcribed speech corpus of read Wikipedia articles which can be used as a benchmark for research on code-switching. The articles are read by a large and diverse group of people. The SWC is perhaps the largest corpus of freely-available aligned speech for German. It contains 1014 spoken articles read by more than 350 identified speakers comprising 386h of speech. This corpus is available at http://nats.gitlab.io/swc.</p> <p>In SWC, since most of the articles are long, the recordings submitted by the volunteers are also long (∼54min) on average. These audio files are manually annotated at word-level and also segment level in XML format. We use a language identification tool to detect code-switching in the transcription of the audio files with consecutive indices. To extract intra-sentential code-switching segments, we ensure that the detected code-switching is preceded and followed by German words or sentences. The final set consists of 34h of speech data and 12,437 code-switching segments (in Kaldi ASR toolkit data format).</p> <p> </p> <p><strong>Citation</strong></p> <p>@article{baumann2019spoken,<br> title={The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening},<br> author={Baumann, Timo and K{\"o}hn, Arne and Hennig, Felix},<br> journal={Language Resources and Evaluation},<br> volume={53},<br> number={2},<br> pages={303--329},<br> year={2019},<br> publisher={Springer}<br> }</p> <p>@article{grave2018learning,<br> title={Learning word vectors for 157 languages},<br> author={Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas},<br> journal={arXiv preprint arXiv:1802.06893},<br> year={2018}<br> }</p> <p> </p>
SWAHILI AND CODE-SWITCHED ENGLISH-SWAHILI POLITICAL HATE SPEECH DETECTION TEXTUAL DATASET
<p>This dataset consist of swahili and code switched English-Swahili tweets labeled for hate speech type,the target and language.</p> <p>The dataset can be accessed upon request from the authors via email addresses:</p> <p> -nodhianbo@gmail.com</p> <p> -endeshanelly@gmail.com</p> <p> -mscci00083@student.maseno.ac.ke</p>
The TongueSwitcher Corpus of German-English Code-Switching
<p>This is the TongueSwitcher Corpus of German-English code-switching tweets. Included are the train and dev sets with automatic word (and subword for mixed words) language identification, alongside the human-annotated corpus with test and interlingual homograph sets.</p> <p>BibTeX entry and citation info</p> <p>@inproceedings{sterner2023tongueswitcher,<br> author = {Igor Sterner and Simone Teufel},<br> title = {TongueSwitcher: Fine-Grained Identification of German-English Code-Switching},<br> booktitle = {Sixth Workshop on Computational Approaches to Linguistic Code-Switching},<br> publisher = {Empirical Methods in Natural Language Processing},<br> year = {2023},<br>}</p>
cantonese loanwords and code-switching
<p>VIDEO</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.