Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>We developed a tool for collecting Tunisian dialect data, prompting users to record themselves reading provided phrases. We sourced sentences from Tunisiya. These sentences are consequently removed from the LM training corpus. 89 persons have participated leading to the collection of 2631 distinct phrases. This set will be called TunSwitch TO, ``TO" standing for Tunisian Only, as these sentences do not have non-Tunisian words. </p> <p>In response to the limited availability of paired Text-Speech Tunisian datasets with code-switching, we have built a corpus through meticulous manual annotation. Whenever encountered, French and English words are enclosed within "<>" tags, and left Tunisian words without any enclosing tags. While these tags have not been used in the proposed models, they allow to have language-usage statistics and may be useful for further approaches handling code-switching. The resulting set is released as TunSwitch CS, ``CS" standing for Code-Switched.</p> <p>The TunSwitch CS dataset samples come from a set of radio shows and podcasts, representing diverse topics and a large number of unique speakers. The audio are first segmented into chunks, prioritizing word integrity using the WebRTC-VAD algorithm for silence detection. Afterward, we used a Pyannote overlap detection model to remove overlapping speech sections. Then, a music detection model is employed to eliminate music-containing chunks that could disrupt ASR model accuracy. <br> </p> <p> </p>
TunSwitch: Code-Switched Tunisian Arabic Speech Dataset
<p>This folder contains the data used to develop and test the Tunisian Arabic Automatic Speech Recognition model developed in the following paper :</p> <p>A. A. Ben Abdallah*, A. Kabboudi, A. Kanoun, and S. Zaiem*, “Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition”, Submitted to ICASSP 2024, vol. * : These two authors have contributed equally. 2023.</p> <p><br> It contains 4 zipped folders containing audio data :<br> - TunSwitchCS.zip : containing annotated code-switched data.<br> - TunSwitchTO.zip : containing annotated Tunisian-Only data.<br> - weakly_labeled_tn.zip : containing weakly-labeled (or unlabeled) audio data. Audios may contain code-switching, but the current weak labels do not.<br> - test_wavs.zip : contains annotated testing data, divided between a code-switched part and a tunisian-only part.</p> <p><br> It also contains textual data, used for language modelling, contained in TextData.zip. Finally it also contains a language-detailed annotation of TunSwitchCS in the language_annotation.zip file .</p> <p>More details about the data are available in the paper. The current table are in a SpeechBrain-friendly format, the column path is irrelevant and has to be changed according to your local setting. Please use the provided train-dev-test splits if you work with this dataset.</p> <p>Please cite the aforementioned paper if you use or refer to this dataset. You can find models trained and tested on this dataset <a href="https://huggingface.co/SalahZa">Here</a>. Space demos are also available. </p> <p>If you use or refer to this dataset, please cite : </p> <p>```</p> <p>@misc{abdallah2023leveraging,<br> title={Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition}, <br> author={Ahmed Amine Ben Abdallah and Ata Kabboudi and Amir Kanoun and Salah Zaiem},<br> year={2023},<br> eprint={2309.11327},<br> archivePrefix={arXiv},<br> primaryClass={eess.AS}<br> }</p> <p>```</p> <p><br> </p>
ODSS: An Open Dataset of Synthetic Speech
<p>ODSS is a multilingual, multispeaker dataset of synthetic and natural speech, designed to foster research and benchmarking of novel studies on synthetic speech detection. </p> <p>ODSS comprises audio utterances generated from text by state-of-the-art synthesis methods, paired with their corresponding natural counterparts. The synthetic audio data includes several languages, with an equal representation of genders.</p> <p>Natural and synthetic speech audio files within ODSS are released under the CC-BY-SA 4.0 license: Usage, extension and redistribution by the research community are strongly encouraged.</p>
Effectiveness of Ultrasound-Aided Articulation Therapy for Children with Speech Sound Disorders
ClinicalTrials.gov study NCT06831396. IPD Sharing: Not stated. Countries: 1. Publications: 42.
Efficacy of Ultrasound Biofeedback in Brazilian Childhood Apraxia of Speech
ClinicalTrials.gov study NCT07087249. IPD Sharing: NO. Countries: 1. Publications: 2.
Comparing Traditional and Biofeedback Telepractice Treatment for Residual Speech Errors
ClinicalTrials.gov study NCT04625062. IPD Sharing: NO. Countries: 1. Publications: 10.
Correcting Residual Errors With Spectral, Ultrasound, Traditional Speech Therapy
ClinicalTrials.gov study NCT03737318. IPD Sharing: NO. Countries: 1. Publications: 19.
Speech-Evoked Auditory Potentials: Multisite Pediatric Evaluation
ClinicalTrials.gov study NCT07392164. IPD Sharing: YES. Countries: 2. Publications: 0.
Smart Computing Models, Sensors, and Early Diagnostic Speech and Language Deficiencies Indicators in Child Communication
ClinicalTrials.gov study NCT06633874. IPD Sharing: NO. Countries: 1. Publications: 1.
Video Assisted Speech Technology to Enhance Motor Planning for Speech
ClinicalTrials.gov study NCT04764539. IPD Sharing: NO. Countries: 1. Publications: 7.
Using a Speech-Generating Device to Support Communication in Childhood Dementia
ClinicalTrials.gov study NCT07039084. IPD Sharing: YES. Countries: 1. Publications: 4.
Apraxia of Speech: Comparison of EPG Treatment (Tx) and Sound Production Treatment (SPT)
ClinicalTrials.gov study NCT02554513. IPD Sharing: NO. Countries: 1. Publications: 1.
Predictors of Speech Ability in Down Syndrome
ClinicalTrials.gov study NCT05016037. IPD Sharing: NO. Countries: 1. Publications: 3.
Delineation of Sensorimotor Subtypes Underlying Residual Speech Errors
ClinicalTrials.gov study NCT03736213. IPD Sharing: NO. Countries: 1. Publications: 20.
Hearing Study: Sensitivity to Features of Speech Sounds
ClinicalTrials.gov study NCT03666676. IPD Sharing: NO. Countries: 1. Publications: 1.
Treating Childhood Apraxia of Speech
ClinicalTrials.gov study NCT03238677. IPD Sharing: NO. Countries: 1. Publications: 2.
Wide-Bandwidth Open Canal Hearing Aid For Better Multitalker Speech Understanding
ClinicalTrials.gov study NCT00582946. IPD Sharing: NO. Countries: 1. Publications: 1.
Deep Brain Stimulation Motor Ventral Thalamus (VOP/VIM) for Restoration of Speech and Upper-limb Function in People With Subcortical Stroke
ClinicalTrials.gov study NCT06303869. IPD Sharing: YES. Countries: 1. Publications: 7.
Bihemispheric Transcranial Direct Current Stimulation* on Speech Fluency
ClinicalTrials.gov study NCT06278233. IPD Sharing: YES. Countries: 1. Publications: 3.
Treatments of Acquired Apraxia of Speech
ClinicalTrials.gov study NCT01483807. IPD Sharing: UNDECIDED. Countries: 1. Publications: 8.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.