Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo36/100

TunSwitch: Code-Switched Tunisian Arabic Speech Dataset

<p>We developed a tool for collecting Tunisian dialect data, prompting users to record themselves reading provided phrases. We sourced sentences from Tunisiya.&nbsp;These sentences are consequently removed from the LM training corpus. 89 persons have participated leading to the collection of 2631 distinct phrases. This set will be called TunSwitch TO, ``TO&quot; standing for Tunisian Only, as these sentences do not have non-Tunisian words.&nbsp;</p> <p>In response to the limited availability of paired Text-Speech Tunisian datasets with &nbsp;code-switching, we have built a &nbsp;corpus through meticulous manual annotation. Whenever encountered, French and English &nbsp;words are enclosed &nbsp;within &quot;&lt;&gt;&quot;&nbsp;tags, and left Tunisian words without any enclosing tags. While these tags have not been used in the proposed models, they allow to have language-usage statistics &nbsp;and may be useful for further approaches handling code-switching. The resulting set is released as TunSwitch CS, ``CS&quot; standing for Code-Switched.</p> <p>The TunSwitch CS dataset samples come from a set of radio shows and podcasts, representing diverse topics and a large number of unique speakers. The audio are first segmented into chunks, prioritizing word integrity using the WebRTC-VAD algorithm for silence detection. Afterward, we used a Pyannote overlap detection model to remove overlapping speech sections. Then, a music detection model is employed to eliminate music-containing chunks that could disrupt ASR model accuracy.&nbsp;<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

TunSwitch: Code-Switched Tunisian Arabic Speech Dataset

<p>This folder contains the data used to develop and test the Tunisian Arabic Automatic Speech Recognition model developed in the following paper :</p> <p>A. A. Ben Abdallah*, A. Kabboudi, A. Kanoun, and S. Zaiem*, &ldquo;Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition&rdquo;, Submitted to ICASSP 2024, vol. * : These two authors have contributed equally. 2023.</p> <p><br> It contains 4 zipped folders containing audio data :<br> - TunSwitchCS.zip : containing annotated code-switched data.<br> - TunSwitchTO.zip : containing annotated Tunisian-Only data.<br> - weakly_labeled_tn.zip : containing weakly-labeled (or unlabeled) audio data. Audios may contain code-switching, but the current weak labels do not.<br> - test_wavs.zip : contains annotated testing data, divided between a code-switched part and a tunisian-only part.</p> <p><br> It also contains textual data, used for language modelling, contained in TextData.zip. Finally it also contains a language-detailed annotation of TunSwitchCS in the&nbsp; language_annotation.zip file&nbsp;.</p> <p>More details about the data are available in the paper. The current table are in a SpeechBrain-friendly format, the column path is irrelevant and has to be changed according to your local setting. Please use the provided train-dev-test splits if you work with this dataset.</p> <p>Please cite the aforementioned paper if you use or refer to this dataset. You can find models trained and tested on this dataset <a href="https://huggingface.co/SalahZa">Here</a>. Space demos are also available.&nbsp;</p> <p>If you use or refer to this dataset, please cite :&nbsp;</p> <p>```</p> <p>@misc{abdallah2023leveraging,<br> &nbsp; &nbsp; &nbsp; title={Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition},&nbsp;<br> &nbsp; &nbsp; &nbsp; author={Ahmed Amine Ben Abdallah and Ata Kabboudi and Amir Kanoun and Salah Zaiem},<br> &nbsp; &nbsp; &nbsp; year={2023},<br> &nbsp; &nbsp; &nbsp; eprint={2309.11327},<br> &nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br> &nbsp; &nbsp; &nbsp; primaryClass={eess.AS}<br> }</p> <p>```</p> <p><br> &nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

ODSS: An Open Dataset of Synthetic Speech

<p>ODSS is a multilingual, multispeaker dataset of synthetic and natural speech, designed to foster research and benchmarking of novel studies on synthetic speech detection.&nbsp;</p> <p>ODSS comprises audio utterances generated&nbsp;from text&nbsp;by state-of-the-art synthesis methods, paired with their corresponding natural counterparts. The synthetic audio data includes several languages, with an equal representation of genders.</p> <p>Natural and synthetic speech audio files within ODSS are released under the CC-BY-SA 4.0 license:&nbsp;Usage, extension and redistribution by the research community are strongly encouraged.</p>

openSep 2023View details →
ClinicalTrials.gov36/100

Effectiveness of Ultrasound-Aided Articulation Therapy for Children with Speech Sound Disorders

ClinicalTrials.gov study NCT06831396. IPD Sharing: Not stated. Countries: 1. Publications: 42.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Efficacy of Ultrasound Biofeedback in Brazilian Childhood Apraxia of Speech

ClinicalTrials.gov study NCT07087249. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Comparing Traditional and Biofeedback Telepractice Treatment for Residual Speech Errors

ClinicalTrials.gov study NCT04625062. IPD Sharing: NO. Countries: 1. Publications: 10.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Correcting Residual Errors With Spectral, Ultrasound, Traditional Speech Therapy

ClinicalTrials.gov study NCT03737318. IPD Sharing: NO. Countries: 1. Publications: 19.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Speech-Evoked Auditory Potentials: Multisite Pediatric Evaluation

ClinicalTrials.gov study NCT07392164. IPD Sharing: YES. Countries: 2. Publications: 0.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Smart Computing Models, Sensors, and Early Diagnostic Speech and Language Deficiencies Indicators in Child Communication

ClinicalTrials.gov study NCT06633874. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Video Assisted Speech Technology to Enhance Motor Planning for Speech

ClinicalTrials.gov study NCT04764539. IPD Sharing: NO. Countries: 1. Publications: 7.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Using a Speech-Generating Device to Support Communication in Childhood Dementia

ClinicalTrials.gov study NCT07039084. IPD Sharing: YES. Countries: 1. Publications: 4.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Apraxia of Speech: Comparison of EPG Treatment (Tx) and Sound Production Treatment (SPT)

ClinicalTrials.gov study NCT02554513. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Predictors of Speech Ability in Down Syndrome

ClinicalTrials.gov study NCT05016037. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Delineation of Sensorimotor Subtypes Underlying Residual Speech Errors

ClinicalTrials.gov study NCT03736213. IPD Sharing: NO. Countries: 1. Publications: 20.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Hearing Study: Sensitivity to Features of Speech Sounds

ClinicalTrials.gov study NCT03666676. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Treating Childhood Apraxia of Speech

ClinicalTrials.gov study NCT03238677. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Wide-Bandwidth Open Canal Hearing Aid For Better Multitalker Speech Understanding

ClinicalTrials.gov study NCT00582946. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Deep Brain Stimulation Motor Ventral Thalamus (VOP/VIM) for Restoration of Speech and Upper-limb Function in People With Subcortical Stroke

ClinicalTrials.gov study NCT06303869. IPD Sharing: YES. Countries: 1. Publications: 7.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Bihemispheric Transcranial Direct Current Stimulation* on Speech Fluency

ClinicalTrials.gov study NCT06278233. IPD Sharing: YES. Countries: 1. Publications: 3.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Treatments of Acquired Apraxia of Speech

ClinicalTrials.gov study NCT01483807. IPD Sharing: UNDECIDED. Countries: 1. Publications: 8.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record