Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

859

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

859 results for “Speeches”

Learn how ShareScore rates datasets ↗
zenodo36/100

CUCO Database: A voice and speech corpus of patients who underwent upper airway surgery in pre- and post-operative states

<p>The data set comprises 3,800 speech audio files of 3 types of upper respiratory tract surgeries and 1 control set. The dataset has an average of <strong>35.51 +- 5.91</strong> audio recordings per patient. It provides valuable resources to the scientific community to systematically investigate the objective effects of upper respiratory tract surgery on voice and speech. &nbsp;</p> <p>This data set is a complete corpus comprising data from 107 Spanish Castilian speakers. This corpus encompasses voice and speech recordings from both control speakers and patients who underwent upper airway surgical procedures in pre- and post-operative stages. The surgeries in focus include <strong>Tonsillectomy</strong>, <strong>Functional Endoscopic Sinus Surgery</strong>, and <strong>Septoplasty</strong>, all consistently performed by a single surgeon.</p> <p>This corpus has been the basis for different previous studies to evaluate changes in voice and its quality due to surgery. The results do not suggest significant changes in the most relevant acoustic parameters studied for the voice, which is consistent with the initial hypothesis. However, the analysis of speech recordings remains open, with a special focus on the nasalised segments, which are expected to change due to surgical intervention.</p> <p>This data set also opens the way to study the effect of upper airway surgery on the performance of speaker recognition and identification methods, as well as to be used to test anti-spoofing methodologies to make them more robust.&nbsp;</p> <p>Please, if you use this database, cite this open-access paper where the data acquisition is explained and detailed:</p> <p>Hern&aacute;ndez-Garc&iacute;a, E., Guerrero-L&oacute;pez, A., Arias-Londo&ntilde;o, J. D., &amp; Godino-Llorente, J. I. (2024). A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states.&nbsp;<em>Scientific Data</em>,&nbsp;<em>11</em>(1), 746.</p>

opencc-by-nc-nd-4.0May 2024View details →
zenodo36/100

AISHELL-Stammertalk 中文口吃数据库 A Mandarin stuttered speech dataset

<p>Dataset official website: <a href="https://aishelltech.com/aishell_6A" target="_blank" rel="noopener">https://aishelltech.com/aishell_6A</a><br><br>This Zenodo page contains dataset samples. To access and download the full dataset, please send an application here&nbsp;<a href="https://opendata.aishelltech.com/stammertalk" target="_blank" rel="noopener">https://opendata.aishelltech.com/stammertalk</a></p> <p>The AISHELL-Stammertalk datasets consists of recordings from 70 native mardarin AWS (Adults who stutter), including 46 males and 24 females. The total duration is 48.8 hours. Each participant engaged in a recording session lasting up to one hour, comprising two parts: conversation and voice command reading. Conversations were conducted through online interviews using platforms like Zoom or Tencent Meet, aiming to capture spontaneous speech on diverse topics. The interviewer, one of the two authors, posed questions based on a prepared list, with the flexibility to introduce impromptu questions as needed.</p> <p>In the voice command reading part, participants were tasked with reading a set of 200 commands, categorized into car navigation and smart home device interaction. To ensure variety, a new set of 200 commands was introduced for every 25 participants, resulting in a dataset featuring a total of 600 unique commands. Participants were encouraged to employ the Voluntary Stuttering technique, deliberately introducing stuttering.</p> <p>Five types of stuttering were specified by the annotation guidelines, including:<br><strong>[]</strong>: Word/phrase repetition. Designated for marking entire repeated character or phrase.<br><strong>/b</strong>: block. Gasps for air or stuttered pauses.<br><strong>/p</strong>: prolongation. Elongated phoneme.<br><strong>/r</strong>: sound repetition. Repeated phoneme that do not constitute an entire character.<br><strong>/i</strong>: interjections. Filler characters due to stuttering e.g., &lsquo;嗯&rsquo;, &lsquo;啊&rsquo;, or &lsquo;呃&rsquo;. Notably, naturally occurring interjections that don't disrupt the speech flow are excluded.</p>

openJul 2024View details →
zenodo36/100

A Comprehensive Central Kurdish Sound Dataset for Robust Automatic Speech Recognition (Part 1).

<p>Exploring the intricacies of Speech Recognition Technology (SRT), our dataset encompasses a wide range of age demographics, spanning from adolescents to individuals in their fifties. This diverse dataset comprises a substantial collection of raw data, amounting to 1,739,089 entries. Within this dataset, a meticulous curation process has yielded a total of 1,683 hours of data, providing a thorough examination of language acquisition patterns across different age cohorts within the Central Kurdish linguistic domain.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

The benefit of combining a deep neural network architecture with ideal ratio mask estimation in computational speech segregation to improve speech intelligibility

<p>Contains all the data:</p> <p>Bentsen, T., T.May, A. A. Kresnner, and T. Dau. The benefit of combining<br> a deep neural network architecture with ideal ratio mask estimation<br> in computational speech segregation to improve speech intelligibility.<br> PLOS ONE., in review.</p> <p>There are two folders:</p> <ol> <li><strong>WRSs:</strong> the Word Recognition Scores (WRSs) from the listener study. The matrix has dimensions 9 conditions x 20 subjects. Data is ordered corresponding to the following condition order:<br> &#39;UP&#39;, &#39;GMM&#39;, &#39;GMM (3 subbands)&#39;, &#39;GMM (7 subbands)&#39;, &#39;GMM (11 subbands)&#39;, &#39;DNN (IBM)&#39;; &#39;DNN (IBM, 40 ms)&#39;; &#39;DNN (IRM)&#39;; &#39;DNN (IRM, 40 ms)&#39;</li> <li><strong>Masks:</strong> <ul> <li><strong>GMM-IBMs:&nbsp;</strong>IBMs and estimated IBMs for the models&nbsp;&#39;GMM&#39;, &#39;GMM (3 subbands)&#39;, &#39;GMM (7 subbands)&#39;, &#39;GMM (11 subbands)&#39;</li> <li><strong>DNN-IBMs:</strong>&nbsp;IBMs and estimated IBMs for the models&nbsp;&#39;DNN (IBM)&#39;; &#39;DNN (IBM, 40 ms)&#39;</li> <li><strong>DNN-IRMs</strong>: IRMs and estimated IRMs for the models&nbsp; &#39;DNN (IRM)&#39;; &#39;DNN (IRM, 40 ms)&#39;</li> </ul> </li> </ol>

opencc-by-4.0Mar 2018View details →
zenodo36/100

VoiceHome-2 corpus : A corpus dedicated to distant-microphone speech processing in domestic environments

<p><strong>Purpose: </strong></p> <p>This corpus includes reverberated, noisy speech signals spoken by 12 native French talkers in 4 houses (3 rooms per house) and recorded by an 8-microphone device at various angles and distances and in various noise conditions.</p> <p>This corpus stands apart from other corpora in the field by the number of rooms and homes considered by the diversity of acoustic conditions recorded and by the facts that it is publicly available at no cost.</p> <p><strong>Other materials:</strong></p> <ul> <li>Article : N. Bertin, E. Camberlein, R. Lebarbenchon, E. Vincent, S. Sivasankaran, I. Illina and F. Bimbot: <a href="https://hal.inria.fr/hal-01923108"><strong>VoiceHome-2, an extended corpus for multichannel speech processing in real homes</strong></a>, <em>Speech Communication</em>, Elsevier : North-Holland, 2019, 106, pp.68-78. <a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">&lang;10.1016/j.specom.2018.11.002&rang;</a>.</li> <li>Code baseline to reproduce article&#39;s results: <ul> <li><a href="https://hal.inria.fr/hal-02963528">Localization and speech enhancement</a></li> <li><a href="https://doi.org/10.5281/zenodo.4079314">Acoustic models</a> and <a href="https://hal.inria.fr/hal-02963802">recognition scripts</a> for automatic speech recognition</li> </ul> </li> <li>Related software to execute the baseline: <ul> <li><a href="https://gitlab.inria.fr/bass-db/mbss_locate">MBSS Locate (v2.0)</a></li> <li><a href="https://gitlab.inria.fr/bass-db/fasst">FASST</a></li> </ul> </li> </ul> <p><strong>Documentation:</strong></p> <p>The corpus documentation is both available into the archive and hereafter by clicking on voiceHome-2_corpus_v1.0_documentation.pdf .</p> <p><strong>Terms of use</strong></p> <p>You may exploit the corpus for a non-commercial scientific purpose provided you mention it in any written work or software you derive from its use. Within a published article, paper or report, the corpus must appear in the bibliographical references.</p> <p><strong>Speaker records diffusion consent</strong></p> <p>All participants have given an informed and signed consent about public diffusion of recorded sentences.</p> <p><strong>Contact:</strong></p> <p>nancy [dot] bertin [at] irisa [dot] fr</p>

opencc-by-nc-sa-4.0May 2018View details →
zenodo36/100

voiceHome corpus: A corpus dedicated to distant-microphone speech processing in domestic environments

<p><strong>Purpose: </strong></p> <p>This corpus includes reverberated, noisy speech signals spoken by native French talkers in a lounge and recorded by an 8-microphone device at various angles and distances and in various noise conditions.</p> <p>Room impulse responses and noise-only signals recorded in various real rooms and homes and baseline speaker localization and enhancement software are also provided.</p> <p>This corpus stands apart from other corpora in the field by the number of rooms and homes considered and by the fact that it is publicly available at no cost.</p> <p>&nbsp;</p> <p><strong>Other materials:</strong></p> <ul> <li>Article: N. Bertin, E. Camberlein, E. Vincent, R. Lebarbenchon, S. Peillon, E. Lamand&eacute;, S. Sivasankaran, F. Bimbot, I. Illina, A. Tom, S. Fleury and E. Jamet: <a href="https://hal.inria.fr/hal-01343060"><strong>A French corpus for distant-microphone speech processing in real homes</strong></a>, Interspeech2016, Sep 2016, San Francisco, United States, 2016.</li> <li>Related software to reproduce article&#39;s results: <ul> <li><a href="https://gitlab.inria.fr/bass-db/mbss_locate">Multi-Channel BSS Locate (v1.3)</a></li> <li><a href="https://gitlab.inria.fr/bass-db/fasst">FASST (v2.2.1)</a></li> </ul> </li> </ul> <p><strong>Documentation (in french):</strong></p> <p>The corpus documentation is both available into the archive and hereafter by clicking on voiceHome_corpus_french_documentation_v1.2.pdf .</p> <p><strong>Terms of use</strong></p> <p>You may exploit the corpus for a non-commercial scientific purpose provided you mention it in any written work or software you derive from its use. Within a published article, paper or report, the corpus must appear in the bibliographical references.</p> <p><strong>Speaker records diffusion consent</strong></p> <p>All participants have given an informed and signed consent about public diffusion of recorded sentences.</p> <p><strong>New corpus version available : voiceHome-2 corpus</strong></p> <p>A new version of the corpus is available : <a href="https://doi.org/10.5281/zenodo.1252143"><strong>voiceHome-2 corpus web page</strong></a></p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Jul 2018View details →
zenodo36/100

RAW DATA form - Speech Auditory Brainstem Responses: Effects of Background, Stimulus Duration, Consonant-Vowel, and Number of Epochs

<p><strong>Speech Auditory Brainstem Responses: Effects of Background, Stimulus Duration, Consonant-Vowel, and Number of Epochs</strong></p> <p>Ghada BinKhamis, Agn&egrave;s L&eacute;ger, Steven L. Bell, Garreth Prendergast, Martin O&rsquo;Driscoll, and Karolina Kluk</p> <p><strong>doi: 10.1097/AUD.0000000000000648</strong></p> <p><em>(<strong>Please site above article)</strong></em></p> <p>&nbsp;</p> <p><strong>Description of raw EEG (speech-ABR) data main folder, subfolders, and raw EEG files</strong></p> <p><strong>Folder Information</strong></p> <p><strong>Main folder:</strong></p> <ul> <li>Contains 144 subfolders with raw data from 12 participants</li> </ul> <p><strong>Subfolder names:</strong></p> <ul> <li>Each subfolder starts with the participant code <ul> <li>S01, S02, S03, S04, S05, S06, S07, S08, S09, S10, S11, S12</li> </ul> </li> </ul> <ul> <li>Next is the stimulus duration: <ul> <li>40ms, 50ms, 170ms</li> </ul> </li> <li>Next is the CV used to evoke speech-ABRs <ul> <li>ba, da, ga</li> </ul> </li> <li>And finally the background condition&nbsp; <ul> <li>quiet, noise</li> </ul> </li> </ul> <p><strong>Example subfolder names:</strong></p> <ul> <li><em>S01 40ms da noise:</em>Participant number 1, speech-ABRs in response to the 40ms [da] in background noise</li> <li><em>S07 170ms ga quiet:</em>Participant number 7, speech-ABRs in response to the 170ms [ga] in quiet</li> </ul> <p><strong>Each participant has 12 subfolders:</strong></p> <ol> <li>S__ 40ms da quiet&nbsp;</li> <li>S__ 40ms da noise</li> <li>S__ 50ms da quiet</li> <li>S__ 50ms da noise</li> <li>S__ 50ms ba quiet</li> <li>S__ 50ms ba noise</li> <li>S__ 50ms ga quiet</li> <li>S__ 50ms ga noise</li> <li>S__ 170ms da quiet</li> <li>S__ 170ms da noise</li> <li>S__ 170ms ba quiet</li> <li>S__ 170ms ga quiet</li> </ol> <p><strong>Each participant subfolder contains four &lsquo;.mat&rsquo; files, &lsquo;.mat&rsquo; file names:</strong></p> <ul> <li>Each &lsquo;.mat&rsquo; file starts with the participant code <ul> <li>S01, S02, S03, S04, S05, S06, S07, S08, S09, S10, S11, S12&nbsp;</li> </ul> </li> <li>Next is the stimulus duration: <ul> <li>40ms, 50ms, 170ms</li> </ul> </li> <li>Next is the CV used to evoke speech-ABRs <ul> <li>ba, da, ga</li> </ul> </li> <li>Next is &lsquo;noise&rsquo; if background condition was noise</li> <li>Next is the stimulus polarity <ul> <li>Pos for positive/standard</li> <li>Neg for negative (reversed polarity stimulus)</li> </ul> </li> <li>And finally is the recording number for that polarity <ul> <li>R1 is the first recording</li> <li>R2 is the second recording</li> </ul> </li> </ul> <p><strong>Example &lsquo;.mat&rsquo; file name:</strong></p> <ul> <li><em>S04 50 ba Neg R1:</em>Participant number 4, speech-ABR in response to the 50ms [ba] in quiet, reversed polarity stimulus, recording number one&nbsp;</li> <li><em>S02 40 da noise Pos R2:</em>Participant number 2, speech-ABR in response to the 40ms [da] in background noise, standard/positive stimulus, recording number two</li> </ul> <p>&nbsp;</p> <p><strong>File Information:</strong></p> <p><strong>Description of &lsquo;.mat&rsquo; files that can be accessed and processed using MATLAB (MathWorks):</strong></p> <p>Each &lsquo;.mat&rsquo; file is a structure that contains the following fields:</p> <ul> <li>The first nine fields are informational, for example: <ul> <li>xunits: &lsquo;s&rsquo; indicates that the recording time window is in seconds, conversion to milliseconds would be required to plot the data in milliseconds</li> <li>start: &lsquo;0&rsquo; indicates that both stimulus and recording start at 0 seconds</li> <li>points:&nbsp;<strong>1800</strong>is the number of sample points for speech-ABRs to the 40ms da, this number will be&nbsp;<strong>2200</strong>for the speech-ABRs to the 50ms stimuli (ba, da, ga), and&nbsp;<strong>4600</strong>for the speech-ABRs to the 170ms stimuli (ba, da, ga)</li> <li>chans: 2 is the number of channels (channel 2 is the ipsilateral channel)</li> <li>frames: 3000 is the number of epochs</li> </ul> </li> <li>The last filed&nbsp;<strong>&lsquo;values&rsquo;</strong>is what contains the raw EEG data (1800x2x3000) <ul> <li><strong>1800&nbsp;</strong>is the number of samples</li> <li><strong>2&nbsp;</strong>is the number of channels (channel one is recorded from the left ear lobe (A1) and channel two is from the right ear lobe (A2))</li> <li><strong>3000&nbsp;</strong>is the number of epochs</li> <li>The field&nbsp;<strong>&lsquo;values&rsquo;&nbsp;</strong>for speech-ABRs to the 50ms stimuli is&nbsp;<strong>2200x2x3000&nbsp;</strong>and for speech-ABRs to the 170ms stimuli is&nbsp;<strong>4600x2x3000</strong>.</li> </ul> </li> <li>Stimulus starts at 0 seconds per epoch, pre-stimulus baseline may be extracted from the end of each epoch (i.e. before the next stimulus).</li> </ul> <p><strong>Data is recorded in Volts and will need to be converted to Micro Volts</strong></p> <p><strong>Date of data collection</strong>: May to November 2016</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

A Canadian French Emotional Speech Dataset

<p>The Canadian French Emotional (CaFE) speech dataset contains six different sentences, pronounced by six male and six female actors, in six basic emotions plus one neutral emotion. The six basic emotions are acted in two different intensities: mild (&quot;Faible&quot;) and strong (&quot;Fort&quot;).</p> <p>This dataset is freely available under a Creative Commons license (CC BY-NC-SA 4.0).</p> <p>The dataset was digitally recorded at a high-resolution (192 kHz sampling rate, 24 bits per sample).</p> <p>A filtered and downsampled version at a lower resolution (48 kHz sampling rate, 16 bits per sample) is also available.</p> <p>This work is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p> <p>To view a copy of this license, visit <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons</a></p>

opencc-by-nc-sa-4.0Apr 2018View details →
zenodo36/100

A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus

<p>We produce Sejong-style morphological analysis and part-of-speech tagging results which have been the de facto&nbsp;standard for Korean language processing by using UDPipe (http://ufal.mff.cuni.cz/udpipe)&nbsp;</p> <p>&nbsp;</p> <p>udpipe --tokenize --tag sjmorph.model input &gt; output</p> <p>see&nbsp;https://github.com/jungyeul/sjmorph</p>

opencc-by-4.0May 2019View details →
zenodo36/100

Dataset for "What is Gab? A Bastion of Free Speech or an Alt-Right Echo Chamber?"

<p>This dataset was used for this project: &quot;What is Gab? A Bastion of Free Speech or an Alt-Right Echo Chamber?&quot;. Savvas Zannettou, Barry Bradlyn, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringhini, Haewoon Kwak, Jeremy Blackburn. Workshop on Computational Methods in CyberSafety, Online Harassment and Misinformation, 2018.&nbsp;DOI:&nbsp;<a href="https://arxiv.org/ct?url=http%3A%2F%2Fdx.doi.org%2F10%252E1145%2F3184558%252E3191531&amp;v=345a781d">10.1145/3184558.3191531</a>.</p> <p>In addition, this project has received funding from the European Union&rsquo;s Horizon 2020 Research and Innovation program under the Marie Skłodowska-Curie ENCASE project (Grant Agreement No. 691025). The work reflects only the authors&rsquo; views; the Agency and the Commission are not responsible for any use that may be made of the information it contains.</p> <p>Using Gab&rsquo;s API, we crawl the social network using a snowball methodology. Specifically, we obtain data for the most popular users as returned by Gab&rsquo;s API and iteratively collect data from all their followers as well as their followings. Subsequently, for all users in our dataset we collect all of their the posts. Overall, we collect 22,112,812 posts from 336,752 users, between August 2016 and January 2018.&nbsp;This dataset is a .json file and each line has one .json&nbsp;object.</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

Repository of speech features from speakers with and without Parkinson's Disease. Neurovoz - Rasta PLP - V2 - Scientific Reports Publication: Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease

<p>This repository contains the Rasta-PLP features of six different speech recordings (sentences) from Neurovoz corpus (47 parkinsonian and 32 control speakers whose mother tongue is Spanish Castillian.)<br> Number of PLP coefficients: [6, 8, 10, 12, 14, 16, 18, 20].<br> Delta coefficients: Yes<br> Delta Delta coefficients: Yes<br> Sampling rate: 16 kHz<br> Frame size: 15 ms<br> Frame overlapping: 50%</p> <p>This subset of the Neurovoz corpus was recorded between 2015 and 2017 by Universidad Polit&eacute;cncia de Madrid and Hospital General Universitario Gregorio Mara&ntilde;&oacute;n.</p> <p>This version includes the same files as the previous version and information about UPDRS, H&amp;Y, years since diagnosis and age of each participant.</p> <p>The sentences were:</p> <p>BARBAS: &quot;Cuando las barbas de tu vecino veas pelar, pon las tuyas a remojar&quot;</p> <p>CALLE: &quot;De la calle vendr&aacute; quien de tu casa te echar&aacute;&quot;</p> <p>DIABLO: &quot; Cuando el diablo no sabe qu&eacute; hacer, con el rabo mata moscas &quot;</p> <p>PETACA BLANCA: &quot; La petaca blanca es m&iacute;a&quot;</p> <p>PIDIO: &quot;No pidas a quien pidi&oacute; ni sirvas a quien sirvi&oacute;&quot;</p> <p>SOMBRA: &quot; El que a buen &aacute;rbol se arrima, buena sombra le cobija &quot;</p> <p>&nbsp;</p> <p>How to cite:<br> [1] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Grandas-Perez, F., Shattuck-Hufnagel, S. Yag&uuml;e-Jimenez, V., and Dehak, N. (2019).&nbsp;Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson&rsquo;s disease.Scientific reports&nbsp;9,&nbsp;19066.</p> <p><br> [2] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Villalba, J., Rusz,&nbsp;J.,&nbsp;Shattuck-Hufnagel, S. and Dehak, N. (2019).&nbsp;A forced Gaussians based methodology for the differential evaluation of Parkinson&#39;s Disease by means of speech processing. Biomedical Signal Processing and Control, 48, 205-220.</p> <p>BibTeX:</p> <pre><code>@article{moro2019phonetic, title={Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge A. and Godino-Llorente, Juan I. and Grandas-Perez, Francisco and Shattuck-Hufnagel, Stefanie and Yague-Jimenez, Virginia and Dehak, Najim}, journal={Scientific Reports}, volume={9}, pages={19066}, year={2019}, publisher={Nature Research Publishing} } @article{moro2019forced, title={A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge Andres and Godino-Llorente, Juan Ignacio and Dehak, Najim}, journal={Biomedical Signal Processing and Control}, pages={205--220}, volume={48}, year={2019}, publisher={Elsevier} } </code></pre> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

NeuroVox: Bilingual Brain-to-Speech Translation and Neural Activity Detection

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo36/100

Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,&rdquo; arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong>&nbsp;<strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>,&nbsp;</li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications.&nbsp;</p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset&nbsp;</strong><strong>are&nbsp;</strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian.&nbsp;</p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Toward Visual Pronunciation Learning: A Speech-to-Articulatory Animation Pipeline Leveraging wav2vec 2.0 and rtMRI Landmarks

<p>Each Video Includes 5 sections:</p> <ul> <li>Top left-most: Phoneme Transcription from Original dataset of USC-TIMIT Dataset.</li> <li>Top middle-left: Generated Articulatory Animation from speech input from this paper.</li> <li>Top middle-right: Ground Truth from refined contour dataset from Refined rtMRI Landmark-Based Vocal Tract Contour Labels.</li> <li>Top right-most: rtMRI data from USC-TIMIT Dataset.</li> <li>Below: Sentences and words</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Indonesian Foreign Policy towards Iran: Shia Hate speech on X social media with LSTM and SVM Analysis

<p>This table and figure are integral components of research on Shia hate speech on social media X, a critical issue in the context of identity politics in Indonesia and globally. This study is of paramount importance as it delves into the identity politics often exploited by politicians in Indonesia and around the world. The Shia community's support for President Jokowi in the first and second stages of the Election was met with hate speech from the opposition group. The study further investigates whether this Shia hate speech is linked to the government's policy towards Iran, a country known for its Shia ideology. The study is presented in three parts:<br>1. Table detailing the sentiment analysis process and results, which were conducted using advanced machine learning techniques such as SVM and LSTM. This approach significantly enhances the accuracy and reliability of the study's findings.<br>2. Figure in the form of a graph related to the study results and the results of comments from the Indonesian public about Shia.<br>3. This research is backed by a comprehensive dataset comprising public comments from Indonesia on Shia. This extensive data collection ensures the study's conclusions are thorough and reliable.</p> <p>4. Processing of machine learning</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Helsinki Speech Challenge 2024 open audio dataset

<h1>Dataset</h1> <p>This training dataset is originally designed for the Helsinki Speech Challenge 2024 (HSC2024). While it was created with this challenge in mind, its applications extend far beyond, making it a valuable resource for developing and testing audio algorithms across diverse uses.</p> <p>Our dataset features clean speech samples generated by OpenAI's text-to-speech model, paired with corresponding recorded signals. These recorded signals are purposefully distorted by real-world effects such as filtering and reverb, offering a realistic testing ground for your audio processing algorithms.</p> <p>To ensure ease of use, the audio samples are organized into 10 separate zip files, each dedicated to a specific task and level of complexity. Most zip files include two folders clean and recorded: one containing clean audio data and the other housing the corresponding recorded (distorted) data. However, note that folders Task_3_Level_1 and Task_3_Level_2 only include recorded data, as their clean counterparts are identical to those in Task_2_Level_2 and Task_2_Level_3, respectively. Additionally, each zip file includes a .txt file with the original text samples associated with the audio clips.<br><br><strong>Update 29th October:&nbsp;</strong>We have now also added the test data that was used for evaluation in the challenge with similar structure to the rest of the data, contained in the zip files starting with "Test".</p> <p>Additionally, there is folder Impulse_Responses, which contains a clean sine sweep signal and short and a long white noise signal with recorded counterparts.&nbsp;<br>To help you get started, we've also provided an example folder containing samples from each task and level, giving you a comprehensive overview of the dataset's scope and variety.<br><br>The dataset also contains a python script evaluate.py. This can be used to evaluate the quality of audio files using the Mozilla Deepspeech speech recognition model. For more details on this script, see the more detailed description of the data challenge either on the website below or on arXiv <a href="https://arxiv.org/abs/2406.04123">https://arxiv.org/abs/2406.04123</a>.<br><br>Important dates:</p> <ul> <li>Data Challenge Launch: 10. June 2024.</li> <li>Sign-up deadline: 1. September 2024 (if you missed this deadline and wish to participate in the challenge, please send us an email).</li> <li>Submission deadline: 6. October 2024. We realize this deadline is a bit optimistic, but we humbly ask participants to try to make this deadline.</li> <li>Results are published: 4. November.</li> <li>Inverse days: 10.-13. December in Oulu, Finland.</li> </ul> <p><strong>Here is a link to the official webpage of the HSC2024: </strong><a href="https://blogs.helsinki.fi/helsinki-speech-challenge/" target="_blank" rel="noopener">https://blogs.helsinki.fi/helsinki-speech-challenge/</a><br><br><strong>Contact Email: </strong>hsc2024@helsinki.fi</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Data in support of the article "Orthogonal neural codes for speech in the infant brain"

<p>Each subject has a dedicated .zip directory containing:</p> <p>[A] the EEG data in the original .raw format (recorded and converted from .mff files with the NetStation 5.3 software);</p> <p>[B] an event file with 6 columns reporting in order:</p> <ul> <li><strong>the onset time of the events (i.e. syllables) in samples</strong></li> <li>zeros (only useful when working with MNE Python)</li> <li>event codes (useful when working with MNE Python)</li> <li><strong>the id of the event</strong><strong> (i.e. syllable: e.g. &lsquo;bi_f&rsquo; where the last letter corresponds to speaker&rsquo;s gender)</strong></li> <li>the name&nbsp;of the precise&nbsp;speech token presented, specifying&nbsp;its duration (in ms)</li> <li>inter-stimulus-intervals (ISI, i.e. time lag in ms&nbsp;from syllable offset to next onset)</li> </ul> <p><em>Additional material</em></p> <p>The .txt file contains the <em>xyz</em> coordinates of the prototype EEG net employed. Being a custom and unique net, please note that such coordinates are approximate (and not perfectly symmetrical). In order to use them&nbsp;it is necessary to drop the channels E125 to E128 (EOG), not employed for the experiment.</p> <p>In scripts.zip there is a demonstration of&nbsp;how to use the event file (with MNE Python) to construct epochs and Python codes for&nbsp;the decoding analyses reported in the paper.</p> <p><em>Useful to know:</em></p> <p>For the original publication, EEG data was pre-processed with a preliminary&nbsp;version of&nbsp;<a href="https://www.biorxiv.org/content/10.1101/2021.05.21.445085v1">APICE</a>&nbsp;(see the Methods section of the paper for more details). A more advanced version of the pre-processing pipeline is now available <a href="https://github.com/neurokidslab/eeg_preprocessing">here</a>.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Noise data for the study of consonant-in-noise discrimination using an auditory model with different speech-based decision devices

<p>The data stored in this repository correspond to two sets of 5000 speech-shaped noises (SSN) that were used in the conference paper titled &quot;Consonant-in-noise discrimination using an auditory model with different speech-based decision devices&quot; by the same authors, presented in the DAGA conference in Vienna, Austria,&nbsp;on 17/08/2021.&nbsp;</p> <p>The two zip files (<strong>osses2021c_S01</strong> and <strong>osses2021c_S02</strong> for participants S01 and S02, respectively) have following structure:</p> <ul> <li><strong>NoiseStim-SSN</strong>: Folder containing the 5000 noises</li> <li><strong>Results</strong>: Results of the listening experiment collected using the fastACI toolbox (https://github.com/aosses-tue/fastACI).</li> </ul> <p>To obtain similar results for other participants the same experiment has to be run using the fastACI toolbox. For instance, to collect&nbsp; new data for participant &#39;S03&#39;, you have to input the following command in MATLAB:</p> <pre><code class="language-bash">fastACI_experiment('speechACI_varnet2013','S03','SSN');</code></pre>

opencc-by-4.0Aug 2021View details →
zenodo36/100

emoUERJ: an emotional speech database in Portuguese

<p><strong>GOAL</strong></p> <p>Since language is a key issue in speech emotion recognition (SER) and there are few databases in Portuguese, this database was developed at the State University of Rio de Janeiro aiming to the development of specific SER models for this language.</p> <p>&nbsp;</p> <p><strong>DATABASE DESIGN</strong></p> <p>Ten sentences were made available to eight actors, equally divided between genders, and they were free to choose the phrases for record audios in four emotions target: happiness, anger, sadness or neutral. The following phrases were used:</p> <ul> <li>N&atilde;o importa quem est&aacute; certo. (It doesn&#39;t matter who is right.)</li> <li>Voc&ecirc; perde tempo demais com a Internet. (You waste too much time on the Internet.)</li> <li>A garrafa est&aacute; na geladeria. (The bottle is in the fridge.)</li> <li>Eu estou me sentindo doente hoje. (I&#39;m feeling sick today)</li> <li>Eu estou um pouco atrasado. (I&#39;m a little late)</li> <li>Nos fins de semana, eu sempre ia para a casa dele(a). (On weekends, I&nbsp;always used to go to his/her house)</li> <li>De quem s&atilde;o essas malas que est&atilde;o debaixo da mesa? (Whose bags are under the table?)</li> <li>Ele volta na quarta-feira. (He comes back on wednesday)</li> <li>J&aacute; chega! Eu vou tomar um banho e ir para a cama. (Enough! I&#39;m going to take a shower and go to bed)</li> <li>Voc&ecirc; poderia arrumar a mesa, por favor? (Could you set the table, please?)</li> </ul> <p>The result of this process was 377&nbsp;audios distributed as follows</p> <ul> <li>happiness: 91</li> <li>anger: 94</li> <li>sadness: 100</li> <li>neutral: 92</li> </ul> <p><strong>FILE IDENTIFICATION</strong></p> <p>Each database file corresponds to a phrase recorded by an actor expressing one of the four emotions and was named as follows:</p> <ul> <li>Position 1: actor&#39;s gender (&#39;m&#39; for man or &#39;w&#39; for woman)</li> <li>Positions 2 and 3: actor&#39;s id&nbsp;(from 01 to 04)</li> <li>Position 4: emotion (h: happiness, a: anger, s: sadness, n: neutral)</li> <li>Positions 5 and 6: recording identification</li> </ul> <p>For example, the file &#39;w04a11&#39; was the eleventh audio recorded by actress 04 interpreting the anger emotion.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

DDS (Device-Degraded Speech) Dataset - DAPS portion

<p>DDS (Device-Degraded Speech) dataset provides aligned parallel recordings of high-quality speech (recorded in professional studios) and a large number of versions of low-quality speech, producing approximately 2,000 hours speech data.&nbsp;</p> <p>DDS is built on top of two datasets: DAPS and VCTK. We play clean speech recordings (4 hours from DAPS and 8 hours from VCTK) and re-record waveforms in nine environments (two offices, two conference rooms, three studios, one living room, one waiting room) on three different devices (one MEMS and two condenser microphones), producing 27 different recording conditions. Moreover, each version of condition consists of multiple recordings recorded at 6 different microphone positions to simulate various signal-to-noise ratio (SNR) and reverberation levels.&nbsp;</p> <p><strong>Arxiv:&nbsp;</strong>https://arxiv.org/abs/2109.07931</p> <p>&nbsp;</p> <p><strong>The whole dataset is split into 3 repositories (one part for DAPS portion, two parts for VCTK portion). This repository contains DAPS portion of DDS.</strong></p> <p><strong>For all repository links of DDS v0.8:</strong></p> <ul> <li><strong>DAPS portion:</strong>&nbsp;https://zenodo.org/record/5464104</li> <li><strong>VCTK portion part1:</strong>&nbsp;https://zenodo.org/record/5499506</li> <li><strong>VCTK portion part2:</strong>&nbsp;https://zenodo.org/record/5501697</li> </ul>

opencc-by-nc-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record