Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo28/100

parliamentr: speeches from european parliaments in a standardized, machine-readable format

<p>Data accompagnying the R package parliamentr. Here on zenodo is the dataset, and the R package holds the code used to scrape &amp; clean the data, as well as code to download and use this dataset.</p> <p>Currently under development&nbsp;</p>

opencc-by-sa-4.0Mar 2020View details →
zenodo28/100

Urdu-Sindhi Speech Emotion Corpus

<p>The <strong>Urdu-Sindhi Speech Emotion Corpus</strong> is a dataset collected at Mehran University of Engineering &amp; Technology, Pakistan by a research team led by Dr. Zafi Sherhan Syed and Dr. Sajjad Ali Memon. The dataset consists on&nbsp;1,435 audio recordings in total for seven types of emotions which include&nbsp;anger, disgust, happiness, neutral, sarcasm, sadness, and surprise in two low-resource languages of South Asia,&nbsp;that is Urdu and Sindhi.</p> <p>Due to ethical restrictions we cannot release audio recordings at the moment and instead release five feature sets from the OpenSmile toolkit.</p> <p>Please note that we&nbsp;will endeavour&nbsp;to compute any features for you locally and send those features back to you. If this interests you, please contact Dr. Zafi Syed at zafisherhan.shah@faculty.muet.edu.pk</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

Music and Medicine: applications in Neurology, Neuropsychiatry and Speech Therapy

<p>Scientists are showing growing interest on Music Therapy. However, there is still a need to raise awareness among the public on the potential clinical applications of the use of Music.&nbsp;</p> <p>This popularised video&nbsp;is part of a series included in the Certificate in the Fundamentals of Music Therapy CPD activity held at Weill Cornell Medicine Qatar.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Youtube-Dataset for Language Identification in Speech Signals

<p><strong>Youtube-Dataset for Language Identification in Speech Signals</strong></p> <p>- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de</p> <p><strong>Reference</strong></p> <p>In case you use this dataset for your research, please cite</p> <p>Alexandra Draghici, Jakob Abe&szlig;er &amp; Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020</p> <p><strong>Dataset</strong></p> <p>The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo.</p> <p>- 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders &quot;0&quot; - &quot;5&quot; encode the language id:<br> &nbsp; 0 - English<br> &nbsp; 1 - French<br> &nbsp; 2 - German<br> &nbsp; 3 - Greek<br> &nbsp; 4 - Italian<br> &nbsp; 5 - Spanish</p> <p><strong>Audio Processing</strong></p> <p>- mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br> &nbsp;</p>

opencc-by-4.0Jul 2020View details →
dryad28/100

Evolution of the speech‐ready brain: The voice/jaw connection in the human motor cortex

<p>A prominent model of the origins of speech, known as the "frame/content" theory, posits that oscillatory lowering and raising of the jaw provided an evolutionary scaffold for the development of syllable structure in speech. Because such oscillations are non‐vocal in most non‐human primates, the evolution of speech required the addition of vocalization onto this scaffold in order to turn such jaw oscillations into vocalized syllables. In the present functional MRI study, we demonstrate overlapping somatotopic representations between the larynx and the jaw muscles in the human primary motor cortex. This proximity between the larynx and jaw in the brain might support the coupling between vocalization and jaw oscillations to generate syllable structure. This model suggests that humans inherited voluntary control of jaw oscillations from ancestral species, but added voluntary control of vocalization onto this via the evolution of a new brain area that came to be situated near the jaw region in the human motor cortex.</p>

opencc-zeroAug 2020View details →
zenodo28/100

Developing an English course for beginners with the topic Part of Speech using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Code-Switching Speech Corpus

<p><strong>German-English Code-Switching speech dataset</strong></p> <p>We provide means to resegment a subset of the German **Spoken Wikipedia Corpus** (SWC) enabling a particular focus on code-switching.&nbsp; This results in the German-English code-switching corpus, a 34h transcribed speech corpus of read Wikipedia articles which can be used as a benchmark for research on code-switching.&nbsp; The articles are read by a large and diverse group of people. The SWC is perhaps the largest corpus of freely-available aligned speech for German.&nbsp; It contains 1014 spoken articles read by more than 350 identified speakers comprising 386h of speech. This corpus is available at http://nats.gitlab.io/swc.</p> <p>In SWC, since most of the articles are long, the recordings submitted by the volunteers are also long (&sim;54min) on average.&nbsp; These audio files are manually annotated at word-level and also segment level in XML format.&nbsp; We use a language identification tool to detect code-switching in the transcription of the audio files with consecutive indices. To extract intra-sentential code-switching segments, we ensure that the detected code-switching is preceded and followed by German words or sentences. The final set consists of 34h of speech data and 12,437 code-switching segments (in Kaldi ASR toolkit data format).</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>@article{baumann2019spoken,<br> &nbsp; title={The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening},<br> &nbsp; author={Baumann, Timo and K{\&quot;o}hn, Arne and Hennig, Felix},<br> &nbsp; journal={Language Resources and Evaluation},<br> &nbsp; volume={53},<br> &nbsp; number={2},<br> &nbsp; pages={303--329},<br> &nbsp; year={2019},<br> &nbsp; publisher={Springer}<br> }</p> <p>@article{grave2018learning,<br> &nbsp; title={Learning word vectors for 157 languages},<br> &nbsp; author={Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas},<br> &nbsp; journal={arXiv preprint arXiv:1802.06893},<br> &nbsp; year={2018}<br> }</p> <p>&nbsp;</p>

opencc-by-sa-3.0Jan 2021View details →
zenodo28/100

Developing English courses for beginners with Part of Speech material by using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Developing English courses for beginners with Part of Speech material by using Google Classroom

<p>Self-Paced Learning&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Developing English courses for beginners with Part of Speech material by using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Developing English courses for beginners with Part of Speech material by using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →
zenodo28/100

Developing English courses for beginners with Part of Speech material by using Google Classroom

<p>Self-Paced Learning</p>

opencc-by-4.0Jan 2021View details →
dryad28/100

Data from: On the physical origin of linguistic laws and lognormality in speech

Physical manifestations of linguistic units include sources of variability due to factors of speech production which are by definition excluded from counts of linguistic symbols. In this work we examine whether linguistic laws hold with respect to the physical manifestations of linguistic units in spoken English. The data we analyze comes from a phonetically transcribed database of acoustic recordings of spontaneous speech known as the Buckeye Speech corpus. First, we verify with unprecedented accuracy that acoustically transcribed durations of linguistic units at several scales comply with a lognormal distribution, and we quantitatively justify this 'lognormality law' using a stochastic generative model. Second, we explore the four classical linguistic laws (Zipf's law, Herdan's law, Brevity law, and Menzerath-Altmann's law) in oral communication, both in physical units and in symbolic units measured in the speech transcriptions, and find that the validity of these laws is typically stronger when using physical units than in their symbolic counterpart. Additional results include (i) coining a Herdan's law in physical units, (ii) a precise mathematical formulation of Brevity law, which we show to be connected to optimal compression principles in information theory and allows to formulate and validate yet another law which we call the size-rank law, or (ii) a mathematical derivation of Menzerath-Altmann's law which also highlights an additional regime where the law is inverted. Altogether, these results support the hypothesis that statistical laws in language have a physical origin.

opencc-zeroJul 2019View details →
dryad28/100

Data from: Long-term use benefits of personal frequency-modulated systems for speech in noise perception in patients with stroke with auditory processing deficits: a non-randomised controlled trial study

Objectives: Approximately one in five stroke survivors suffer from difficulties with speech reception in noise, despite normal audiometry. These deficits are treatable with personal Frequency Modulated systems (FMs). This study aimed to evaluate long term benefits in speech reception in noise, after daily 10 week use of personal FMs, in non-aphasic stroke patients with auditory processing deficits. Design: This was a prospective non randomised controlled trial study. Patients were allocated to an intervention care group or standard care subjects group according to their willingness to use the intervention or not. Setting: Tertiary care setting. Participants: Nine non-aphasic subjects with ischemic stroke, normal/near normal audiometry, and auditory processing deficits and with reported difficulties understanding speech in background noise were recruited in the subacute stroke stage (3-12 months after stroke). Interventions: Four patients (intervention care subjects) used the FMs in their daily life over 10 weeks. Five patients (standard care subjects) received standard care. Primary outcome measures: All subjects were tested at baseline (visit 1) and 10 weeks later (visit 2) on a sentences in noise test with the FMs (aided) and without the FMs (unaided). Results: Speech reception thresholds showed clinically and statistically significant improvements in intervention but not in standard care subjects at 10 weeks in both aided and unaided conditions. Conclusions: 10 week use of FM systems by adult stroke patients may lead to benefits in unaided speech in noise perception. Our findings may indicate auditory plasticity type changes and require further investigation.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Individual differences in selective attention predict speech identification at a cocktail party

Listeners with normal hearing show considerable individual differences in speech understanding when competing speakers are present, as in a crowded restaurant. Here, we show that one source of this variance are individual differences in the ability to focus selective attention on a target stimulus in the presence of distractors. In 50 young normal-hearing listeners, the performance in tasks measuring auditory and visual selective attention was associated with sentence identification in the presence of spatially separated competing speakers. Together, the measures of selective attention explained a similar proportion of variance as the binaural sensitivity for the acoustic temporal fine structure. Working memory span, age, and audiometric thresholds showed no significant association with speech understanding. These results suggest that a reduced ability to focus attention on a target is one reason why some listeners with normal hearing sensitivity have difficulty communicating in situations with background noise.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Investigation of the effect of cochlear implant electrode length on speech comprehension in quiet and noise compared with the results with users of electro-acoustic-stimulation, a retrospective analysis

Objectives: This investigation evaluated the effect of cochlear implant (CI) electrode length on speech comprehension in quiet and noise and compare the results with those of EAS users. Methods: 91 adults with some degree of residual hearing were implanted with a FLEX20, FLEX24, or FLEX28 electrode. Some subjects were postoperative electric-acoustic-stimulation (EAS) users; the other subjects were in the groups of electric stimulation-only (ES-only). Speech perception was tested in quiet and noise at 3 and 6 months of ES or EAS use. Speech comprehension results were analyzed and correlated to electrode length. Results: While the FLEX20 ES and FLEX24 ES groups were still in their learning phase between the 3 to 6 months interval, the FLEX28 ES group was already reaching a performance plateau at the three months appointment yielding remarkably high test scores. EAS subjects using FLEX20 or FLEX24 electrodes outscored ES-only subjects with the same short electrodes on all 3 tests at each interval, reaching significance with FLEX20 ES and FLEX24 ES subjects on all 3 tests at the 3-months interval and on 2 tests at the 6- months interval. Amongst ES-only subjects at the 3- months interval, FLEX28 ES subjects significantly outscored FLEX20 ES subjects on all 3 tests and the FLEX24 ES subjects on 2 tests. At the-6 months interval, FLEX28 ES subjects still exceeded the other ES-only subjects although the difference did not reach significance. Conclusions: Among ES-only users, the FLEX28 ES users had the best speech comprehension scores, at the 3- months appointment and tendentially at the 6 months appointment. EAS users showed significantly better speech comprehension results compared to ES-only users with the same short electrodes.

opencc-zeroDec 2016View details →
zenodo28/100

FORMATION OF WRITTEN AND SPEECH COMPETENCE IN ENGLISH AMONG UZBEK STUDENTS

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

THE ROLE OF PUNCTUATION MARKS IN POETIC SPEECH

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Golden Ratio Speech Codec (img)

<p>Original speech signal and Golden ratio speech codec resultant speech signal.</p>

opencc-by-nc-nd-4.0Jun 2016View details →
zenodo28/100

THE METHODOLOGICAL DESCRIPTION OF IMPROVING THE SPEECH COMPETENCE OF FUTURE FOREIGN LANGUAGE TEACHERS

<p>This article provides a methodological description of how to improve the speech competence of future foreign language teachers. Effective communication skills are crucial for language instructors, and speech competence plays a vital role in facilitating comprehension, modeling correct pronunciation, and creating an engaging learning environment. The article emphasizes the importance of pronunciation, intonation, fluency, and communicative skills in language teaching. It explores various methodologies for enhancing speech competence, including language immersion, pronunciation practice, technology integration, and cultural competence development. The article also highlights the significance of continuous professional development, peer collaboration, and assessment and feedback in the improvement process. By following these methodological approaches, future foreign language teachers can enhance their speech competence, leading to more effective language instruction and enhanced learning outcomes for their students.</p>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record