Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Speech Corpus of Interpreted Premier Press Conferences (SCIPPC)
<p>SCIPPC v1.0 is a parallel corpus of consecutive interpreting between Mandarin Chinese and English and vice versa in two Chinese premiers’ press conferences in March 2003–2007 and 2013–2017. The conferences were held after sessions of the National People’s Congress and the Chinese People’s Political Consultative Conference. They were moderated by spokespersons of the Congress and the Chinese Ministry of Foreign Affairs and attended by journalists, who asked the premiers questions.</p> <p>SCIPPC v1.0 includes source speeches by approximately 170 speakers and interpretations by six different staff interpreters of the Chinese Ministry of Foreign Affairs, who worked into their B language. It contains 192,209 tokens (source: 108,296, target: 83,913; Chinese: 112,528, English: 79,681) and 19 h 43 min 5 s of video recordings. It is fully transcribed and aligned at the recording–transcript and source–target transcript levels.</p>
Ressources for End-to-End French Text-to-Speech Blizzard challenge
<p>Here are 289 chapters of 5 audiobooks from Librivox (51:12) read by Nadine Eckert-Boulet (NEB):</p> <ol> <li>Madame Bovary (MB) by Gustave Flaubert (FL) - 3 volumes, 35 chapters<br>(original <a href="https://librivox.org/madame-bovary-french-by-gustave-flaubert">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/14155/pg14155.txt">text</a>)</li> <li>Les mystères de Paris (LMP) by Eugene Sue (ES) - 4 volumes, 83 chapters (original <a href="https://librivox.org/les-mysteres-de-paris-tome-1-by-eugene-sue">wavs1</a>, <a href="https://librivox.org/les-mysteres-de-paris-tome-2-by-eugene-sue/">wavs2</a>,<a href="https://librivox.org/les-mysteres-de-paris-tome-3-by-eugene-sue"> wavs3</a>; <a href="https://www.gutenberg.org/cache/epub/18921/pg18921.txt">text1</a>, <a href="https://www.gutenberg.org/cache/epub/18922/pg18922.txt">text2</a>, <a href="https://www.gutenberg.org/cache/epub/18923/pg18923.txt">text3</a>)</li> <li>Les tribulations d'un chinois en Chine (TCC) by Jules Verne (JV) - 1 volume, 22 chapters (original <a href="https://librivox.org/les-tribulations-dun-chinois-en-chine-by-jules-verne">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/14162/pg14162.txt">text</a>)</li> <li>La fille du pirate (LFDP) by Henri Émile Chevalier (EC) - 7 volumes, 121 chapters (original <a href="https://librivox.org/la-fille-du-pirate-by-henri-emile-chevalier">wavs</a>, <a href="https://www.gutenberg.org/cache/epub/18403/pg18403.txt">text)</a></li> <li>La vampire (VAMP) by Paul Féval (PF) - 1 volume, 28 chapters (original <a href="https://librivox.org/la-vampire-by-feval-paul-henry-corentin">wavs</a>, <a href="https://www.gutenberg.org/cache/epub/10053/pg10053.txt">text</a>)</li> </ol> <p>and</p> <p>2515 utterances (2:03) read by another female French speaker Aurélie Derbier (AD):</p> <ol> <li>1608 utterances extracted from various books (DIVERS_BOOK_AD*)</li> <li>907 transcripts of the sessions of the French parliament (DIVERS_PARL_01*)</li> </ol> <p>We recently added three speakers from Librivox/Litteratureaudio:</p> <ol> <li>Ezwa (EZWA): L'épouvante by Maurice Level (original <a href="https://librivox.org/lepouvante-by-maurice-level-1010/">wavs</a>; <a href="https://www.gutenberg.org/cache/epub/17794/pg17794.txt">text</a>) - 11 chapters - 4869 utterances> 03:16</li> <li>Pauline Latournerie (PL): Le pédagogue n'aime pas les enfants by Henri Roorda (<a href="https://librivox.org/le-pedagogue-naime-pas-les-enfants-by-henri-roorda/">original wavs</a>; <a href="https://ebooks-bnr.com/ebooks/pdf4/roorda_le_pedagogue_n_aime_pas_les_enfants.pdf">text</a>) - 6 chapters - 1320 utterances> 01:17</li> <li>Jean-Luc Fischer (JLF): L’Affaire Charles Dexter Ward by Howard Phillips Lovecraft (<a href="https://www.litteratureaudio.com/livre-audio-gratuit-mp3/howard-phillips-lovecraft-laffaire-charles-dexter-ward.html">original wavs</a>; <a href="https://www.litteratureaudio.com/textes/H_P_Lovecraft_L_Affaire_CDW.pdf">text</a>) - 16 chapters - 1823 utterances> 02:37</li> </ol> <p>Each .wav file (sampled at 22050Hz) corresponds to one entire chapter. The format of the filenames is:<br>{author's acronym}_{book's acronym}_{reader's acronym}_{volume's number}_{chapter's number}</p> <p>The NEB_train.csv file gives text and phonetic alignments (essentially for MB and LMP) for utterances in 4 fields separated by '|':<br>{filename}|{start_ms}|{end_ms}|{text or phonetic content}. Most utterances are separated by at least a pause of 400ms. The intervals [start_ms:end_ms] comprise leading and trailing silences of 130ms (since wavs are entire chapters, these silences are "true" ambient silences). Same for AD_train.csv.</p> <p>When phonetic alignment has been performed, 2 additional fields have been added: {aligned phones}|{durations in ms}. Each input character or phone has a corresponding aligned phone and a duration. Note that all aligned utterances start and end with an aligned phone of 130ms. The set of aligned phones comprises:</p> <ul> <li>The set of input phones</li> <li>The silence: '__'</li> <li>The symbol '_' for silent characters, e.g. "chat" is aligned with 's^ _ a _'</li> <li>29 combined aligned phones ('a&i', 'a&j', 'b&q', 'd&q','d&z', 'd&z^', 'f&q', 'g&q', 'g&z', 'j&i', 'j&u', 'j&q', 'i&j', 'k&q', 'k&s', 'k&s&q', 'l&q', 'm&q', 'n&q', 'r&w', 'r&q', 's&q', 't&q', 't&s', 't&s^', 'w&a', 'z&q', 'p&q') that align to only one character, e.g. "expatrier" is aligned with 'e^ k&s p a t r i&j e _'</li> </ul> <p>Text is in UTF8. '«»','¬', '~','""','()','[]' are respectively used for speaking quotes, turn switches, three dots, quoted expression, aside quotes, notes. Because of rare occurrences, 'ö' has been transcribed as 'oe'. Paragraphs (two consecutive carriage returns in the original text) are cued by a special character '§'. It usually ends an utterance but could be used within an utterance if its associated pause is too short.</p> <p>When available, phonetic content is given per word in curly brackets '{}'. We use 39 phonetic symbols:</p> <ul> <li><strong>oral vowels</strong>: a (f<strong><em>a</em></strong>), e (f<em><strong>ée</strong></em>), e^ (f<em><strong>ait</strong></em>), x (f<em><strong>eu</strong></em>), x^ (c<em><strong>oeu</strong></em>r), i (r<em><strong>iz</strong></em>), y (f<em><strong>ut</strong></em>), u (f<em><strong>ou</strong></em>), o (f<em><strong>aux</strong></em>), o^ (p<strong><em>o</em></strong>rc)</li> <li><strong>schwa</strong>: q (gag<strong><em>e</em></strong>)</li> <li><strong>nasal vowels</strong>: a~ (r<strong><em>an</em></strong>g), e~ (f<em><strong>in</strong></em>), x~ (<strong><em>un</em></strong>), o~ (r<em><strong>on</strong></em>d)</li> <li><strong>semi-vowels</strong>: h (h<em><strong>u</strong></em>it), w (<strong><em>ou</em></strong>ate), j (h<em><strong>i</strong></em>er)</li> <li><strong>consonants</strong>: p (<em><strong>p</strong></em>as), t (<strong><em>t</em></strong>as), k (<em><strong>c</strong></em>as), b (<strong><em>b</em></strong>as), d (<em><strong>d</strong></em>os), g (<em><strong>g</strong></em>ars), f (<em><strong>f</strong></em>aux), s (<strong><em>s</em></strong>ot) , s^ (<strong><em>ch</em></strong>at), v (<strong><em>v</em></strong>u), z (<strong><em>z</em></strong>ut), z^ (<em><strong>j</strong></em>us), r (<strong><em>r</em></strong>iz), l (<em><strong>l</strong></em>a), m (<strong><em>m</em></strong>a), n (<strong><em>n</em></strong>on), n~ (oi<strong><em>gn</em></strong>on), ng (campi<em><strong>ng</strong></em>)</li> </ul> <p> </p>
AVbook, a high-frame-rate corpus of narrative audiovisual speech for investigating multimodal speech perception
<p><strong>Please cite</strong><br> Varano E, Guilleminot P, Reichenbach T. <em>AVbook, a high-frame-rate corpus of narrative audiovisual speech for investigating multimodal speech perception</em>. J Acoust Soc Am. 2023 May 1;153(5):3130. doi: 10.1121/10.0019460. PMID: 37249407.<br> <br> Seeing a speaker's face can help substantially in understanding them, in particular in challenging listening conditions. Research into the neurobiological mechanisms behind the audiovisual integration has recently begun to employ continuous natural speech. However, these efforts are impeded by a lack of high-quality audiovisual recordings of a speaker narrating a longer text. Here we seek to close this gap by developing AVbook, an audiovisual speech corpus designed for cognitive neuroscience studies and audiovisual speech recognition. The corpus consists of 3.6 hours of audiovisual recordings of two speakers, one male and one female, reading 59 passages from a narrative English text. The recordings were acquired at a high frame rate of 119.88 frames per second. The corpus includes a sets of multiple-choice questions to test attention to the different passages. We verified the efficacy of these questions in a pilot study. A short written summary is also provided for each recording. To enable audiovisual synchronization when presenting the stimuli, four videos of an electronic clapperboard were recorded with the corpus. The corpus is available for download to support research into the neurobiology of audiovisual speech processing as well as the development of computer algorithms for audiovisual speech recognition.</p>
BinauRec: A dataset to test the influence of the use of room impulse responses on binaural speech enhancement
<p>BinauRec is a dataset for binaural speech enhancement. It is composed of real recordings, measured and simulated room impulse responses for the same audio scenes. Measurements are realized using behind-the-ears hearing aid shells, with and without a dummy head.</p>
Persian Speech to Test dataset
<p>The Persian Speech to Text dataset is a collection of audio files and their corresponding transcripts, provided in CSV file format. The dataset is intended for use in training machine learning models for the task of transcribing audio files in the Persian language into text. The dataset includes 60GB of data, consisting of audio files in the WAV format and their transcripts. Each CSV file corresponds to a single ZIP or RAR file, and the name of each CSV file is the same as the corresponding ZIP or RAR file. The CSV files contain the following columns:</p> <ul> <li>wav_filename: The name of the WAV file within the ZIP or RAR file.</li> <li>wav_filesize: The size of each audio file.</li> <li>transcript: The text transcription of the audio file.</li> <li>confidence_level: A measure of the accuracy of the transcription.</li> </ul> <p>This dataset is the largest open source dataset of its kind, and it is a valuable resource for researchers and developers working on natural language processing tasks involving the Persian language. The open source nature of the dataset means that it is freely available to be used and modified by anyone, making it an important resource for advancing research and development in the field.</p>
ESAA: an EEG-Speech auditory attention detection database
<p>We build a database for AAD research, which consists of competing speech stimuli and associated human neural responses, i.e, electroencephalography (EEG) recordings, namely EEG-Speech AAD (ESAA) database. This is the first AAD database with speech stimuli in a tonal language (Mandarin). Moreover, we develop an AAD baseline as a reference model for decoding which speech stream a listening subject is attending to (speaker attention detection), and a baseline for decoding which spatial locus a listening subject is attending to (speaker locus attention detection) on the ESAA database.</p> <p>We release the source code and the database for use in research purpose.</p> <p>This database consists of response data for 17 normal-hearing subjects (S1-S17). It includes:</p> <p>- 64-channel EEG data: responses to two-speaker speech stimuli<br> - Auditory stimuli data (clean): Chinese short stories narrated by a female and a male professional story teller. <br> - Auditory stimuli data (hrtf): Auditory stimuli after head-related transfer function (HRTF) filtering (simulating sound coming from ± 90 deg).<br> - Preprocessing code<br> - AAD baseline (CNN model)</p>
Syllable level speech sequencing
Open the record for dataset details and reuse information.
German Political Speeches Corpus
<p>This text archive focuses on German political speeches held by top officials mostly from 1990 onwards, selected according to their political relevance. The currently included speeches come from the following sources:</p> <ul> <li>Official pages of the German <a href="http://www.bundespraesident.de/EN/Home/home_node.html">Presidency</a>, <a href="https://www.bundeskanzlerin.de/Webs/BKin/EN/Chancellery/federal_chancellery_node.html">Chancellery</a>, <a href="http://www.bundestag.de/en/parliament/presidium/function_neu">Bundestag</a>, <a href="https://www.auswaertiges-amt.de/en/">Ministry of Foreign Affairs</a></li> <li>Personal pages of the <a href="https://www.helmut-kohl.de/dokumente_reden.html">Helmut Kohl archive</a>, <a href="http://www.thierse.de/reden-und-texte/reden/">Wolfgang Thierse</a> and <a href="http://www.norbert-lammert.de/01-lammert/texte.php">Norbert Lammert</a></li> </ul> <p>This resource is available online:</p> <ul> <li><a href="https://www.dwds.de/r?corpus=politische_reden">Online queries on the DWDS website</a> and <a href="https://www.dwds.de/d/korpussuche">usage instructions</a> (the text base may be newer than the downloadable archives)</li> <li><a href="http://purl.org/corpus/german-speeches">http://purl.org/corpus/german-speeches</a></li> </ul> <p>The files below consist of texts with metadata encoded in <a href="http://xml.silmaril.ie/whatisxml.html">XML format</a>. For appropriate tooling see:</p> <ul> <li>Python tutorial using the speeches: <a href="https://www.timmer-net.de/2019/03/24/nlp_basics/">Natural Language Processing — Einsteigen und Loslegen!</a></li> <li><a href="http://www.corpusexplorer.de">CorpusExplorer</a>, corpus linguistics and text mining software featuring the speeches</li> <li><a href="https://github.com/adbar/german-nlp">List of off-the-shelf NLP tools for German</a></li> </ul> <p>This is work in progress, updated and extended versions will follow.</p>
First Speech Separation Challenge
<p>The first international Speech Separation Challenge took place in 2006, with results disseminated at Interspeech in Pittsburgh and later in a special issue of Computer Speech and Language (volume 24, 2010). The main focus of the challenge was to compare algorithms and human listeners on the task of identifying words in sentences from one talker when mixed with similar utterances from another talker, using a single channel (i.e., operating monaurally). Development and test data was also provided for a stationary noise masking condition. For technical details and a summary of the outcome of the Challenge, see the article by Cooke, Hershey and Rennie (pp 1--15) of the special issue (available as cooke_csl2010.pdf in this dataset). </p> <p>The current dataset consisted of the following zip files:</p> <p>twotalker_dev.zip two-talker development set<br> twotalker_test.zip two-talker test set<br> 1.zip, 2.zip etc training data for each of 34 talkers (500 sentences each)<br> ssn_dev.zip speech-shaped noise (SSN) development set<br> ssn_test.zip SSN test set<br> <br> The Challenge was funded by the EU Network of Excellence PASCAL (Pattern Analysis, Statistical Modeling and Computational Learning) </p>
Listening experiment and Stimuli for: Adjustable Deterministic Pseudonymization of Speech
<p>Web pages for listening experiment, with stimuli included, as reported in: Adjustable Deterministic Pseudonymization of Speech</p> <p>The listening experiments can be run locally offline. After unpacking the files, the listening experiment can be run locally or in a web site by pointing a web browser to the index.html file.</p> <p>A report discussing the pseudonymization results can be found at doi: 10.5281/zenodo.3773931</p> <p>A dataset created with this expriment can be found at doi: 10.5281/zenodo.3773936</p> <p>The <em>akouste</em> listening experiment software can be found at doi: 10.5281/zenodo.3712142 on Github</p> <p>The <em>Pseudonymize</em> <em>Speech</em> <em>Praat</em> script can be found at doi: 10.5281/zenodo.3712140</p> <p>The <em>Praat</em> speech software can be found at <em>www.praat.org</em></p>
BreathBase: Intra-Speech Breathing Dataset
<p>BreathBase contains 5070 breath instances detected on the recordings of 20 participants reading pre-prepared random pseudo texts in 5 different postures with 4 different microphones, simultaneously.</p> <p>It is recorded in a studio with a maximum background noise of 40 dB SPL and with professional recording equipment. It also provides tagging for 5 different postures and 4 different channels as different recording conditions for data variety.</p> <p>More than 90% of the recordings is shorter than 600 milliseconds. The minimum number of breath instances per participant is 89, the maximum number of instances is 710 and the average for all participants is 253.5 breath instances.</p>
MigParl. A Corpus of Speeches on Migration and Integration in Germany's Regional Parliaments
<p>MigParl is an indexed and linguistically annotated corpus of speeches on migration and integration affairs in Germany’s regional parliaments (“Landtage”). The corpus has been prepared in the MigTex Project (principal investigators: Andreas Blätte / University of Duisburg-Essen, Ruud Koopmans / Berlin Social Science Center), using the resources and the infrastructure of the <a href="http://polmine.github.io">PolMine Project</a>.</p> <p>MigTex was part of a larger joint project to establish the research community of the <em>German Centre for Migration and Integration Affairs</em> (<em>Deutsches Zentrum für Migration and Integrationsforschung</em> / DeZIM). Funding awarded by Germany’s <em>Federal Ministry for Family Affairs, Senior Citizens, Women and Youth</em> (<em>Bundesministerium für Familie, Senioren, Frauen und Jugend</em> / BMFSFJ) is gratefully acknowleged.</p>
TweetBLM: A Hate Speech Dataset and Analysis of BlackLivesMatter-related Microblogs on Twitter
<p>Collection of BLM related tweets and their corresponding labels of hate speech.</p>
SPEECH-COCO
<p><strong>SpeechCoco</strong></p> <p><em>Introduction</em></p> <p>Our corpus is an extension of the MS COCO image recognition and captioning dataset. MS COCO comprises images paired with a set of five captions. Yet, it does not include any speech. Therefore, we used <a href="https://www.voxygen.fr/">Voxygen's text-to-speech system</a> to synthesise the available captions.</p> <p>The addition of speech as a new modality enables MSCOCO to be used for researches in the field of language acquisition, unsupervised term discovery, keyword spotting, or semantic embedding using speech and vision.</p> <p>Our corpus is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 License</a>.</p> <p><em>Data Set</em></p> <ul> <li> <p>This corpus contains <strong>616,767</strong> spoken captions from MSCOCO's val2014 and train2014 subsets (respectively 414,113 for train2014 and 202,654 for val2014).</p> </li> <li> <p>We used 8 different voices. 4 of them have a British accent (Paul, Bronwen, Judith, and Elizabeth) and the 4 others have an American accent (Phil, Bruce, Amanda, Jenny).</p> </li> <li> <p>In order to make the captions sound more natural, we used SOX <em>tempo</em> command, enabling us to change the speed without changing the pitch. 1/3 of the captions are 10% slower than the original pace, 1/3 are 10% faster. The last third of the captions was kept untouched.</p> </li> <li> <p>We also modified approximately 30% of the original captions and added <strong>disfluencies</strong> such as "um", "uh", "er" so that the captions would sound more natural.</p> </li> <li> <p>Each WAV file is paired with a JSON file containing various information: timecode of each word in the caption, name of the speaker, name of the WAV file, etc. The JSON files have the following data structure:</p> </li> </ul> <pre><code class="language-json">{ "duration": float, "speaker": string, "synthesisedCaption": string, "timecode": list, "speed": float, "wavFilename": string, "captionID": int, "imgID": int, "disfluency": list }</code></pre> <ul> <li> <p>On average, each caption comprises 10.79 tokens, disfluencies included. The WAV files are on average 3.52 seconds long.</p> </li> </ul> <p><em>Repository</em></p> <p>The repository is organized as follows:</p> <ul> <li> <p>CORPUS-MSCOCO (~75GB once decompressed)</p> <blockquote> <ul> <li> <p><strong>train2014/</strong> : folder contains 413,915 captions</p> <ul> <li> <p>json/</p> </li> <li> <p>wav/</p> </li> <li> <p>translations/</p> <ul> <li> <p>train_en_ja.txt</p> </li> <li> <p>train_translate.sqlite3</p> </li> </ul> </li> <li> <p>train_2014.sqlite3</p> </li> </ul> </li> <li> <p><strong>val2014/</strong> : folder contains 202,520 captions</p> <ul> <li> <p>json/</p> </li> <li> <p>wav/</p> </li> <li> <p>translations/</p> <ul> <li> <p>train_en_ja.txt</p> </li> <li> <p>train_translate.sqlite3</p> </li> </ul> </li> <li> <p>val_2014.sqlite3</p> </li> </ul> </li> <li> <p><strong>speechcoco_API/</strong></p> <ul> <li> <p>speechcoco/</p> <ul> <li> <p>__init__.py</p> </li> <li> <p>speechcoco.py</p> </li> </ul> </li> <li> <p>setup.py</p> </li> </ul> </li> </ul> </blockquote> </li> </ul> <p><em>Filenames</em></p> <p><strong>.wav</strong> files contain the spoken version of a caption</p> <p><strong>.json</strong> files contain all the metadata of a given WAV file</p> <p><strong>.sqlite3</strong> files are SQLite databases containing all the information contained in the JSON files</p> <p>We adopted the following naming convention for both the WAV and JSON files:</p> <p><em>imageID_captionID_Speaker_DisfluencyPosition_Speed[.wav/.json]</em></p> <p><em>Script</em></p> <p>We created a script called <strong>speechcoco.py</strong> in order to handle the metadata and allow the user to easily find captions according to specific filters. The script uses the *.db files.</p> <p>Features:</p> <ul> <li> <p><strong>Aggregate all the information in the JSON files into a single SQLite database</strong></p> </li> <li> <p><strong>Find captions according to specific filters (name, gender and nationality of the speaker, disfluency position, speed, duration, and words in the caption).</strong> <em>The script automatically builds the SQLite query. The user can also provide his own SQLite query.</em></p> </li> </ul> <p><em>The following Python code returns all the captions spoken by a male with an American accent for which the speed was slowed down by 10% and that contain "keys" at any position</em></p> <pre><code class="language-python"># create SpeechCoco object db = SpeechCoco(train_2014.sqlite3, train_translate.sqlite3, verbose=True) # filter captions (returns Caption Objects) captions = db.filterCaptions(gender="Male", nationality="US", speed=0.9, text='%keys%') for caption in captions: print('\n{}\t{}\t{}\t{}\t{}\t{}\t\t{}'.format(caption.imageID, caption.captionID, caption.speaker.name, caption.speaker.nationality, caption.speed, caption.filename, caption.text))</code></pre> <pre><code>... 298817 26763 Phil 0.9 298817_26763_Phil_None_0-9.wav A group of turkeys with bushes in the background. 108505 147972 Phil 0.9 108505_147972_Phil_Middle_0-9.wav Person using a, um, slider cell phone with blue backlit keys. 258289 154380 Bruce 0.9 258289_154380_Bruce_None_0-9.wav Some donkeys and sheep are in their green pens . 545312 201303 Phil 0.9 545312_201303_Phil_None_0-9.wav A man walking next to a couple of donkeys. ...</code></pre> <ul> <li> <p><strong>Find all the captions belonging to a specific image</strong></p> </li> </ul> <pre><code class="language-python">captions = db.getImgCaptions(298817) for caption in captions: print('\n{}'.format(caption.text))</code></pre> <pre><code>Birds wondering through grassy ground next to bushes. A flock of turkeys are making their way up a hill. Um, ah. Two wild turkeys in a field walking around. Four wild turkeys and some bushes trees and weeds. A group of turkeys with bushes in the background.</code></pre> <ul> <li> <p><strong>Parse the timecodes and have them structured</strong></p> </li> </ul> <p><strong>input</strong>:</p> <pre><code>... [1926.3068, "SYL", ""], [1926.3068, "SEPR", " "], [1926.3068, "WORD", "white"], [1926.3068, "PHO", "w"], [2050.7955, "PHO", "ai"], [2144.6591, "PHO", "t"], [2179.3182, "SYL", ""], [2179.3182, "SEPR", " "] ...</code></pre> <p><strong>output</strong>:</p> <pre><code class="language-python">print(caption.timecode.parse())</code></pre> <pre><code>... { 'begin': 1926.3068, 'end': 2179.3182, 'syllable': [{'begin': 1926.3068, 'end': 2179.3182, 'phoneme': [{'begin': 1926.3068, 'end': 2050.7955, 'value': 'w'}, {'begin': 2050.7955, 'end': 2144.6591, 'value': 'ai'}, {'begin': 2144.6591, 'end': 2179.3182, 'value': 't'}], 'value': 'wait'}], 'value': 'white' }, ...</code></pre> <ul> <li> <p><strong>Convert the timecodes to Praat TextGrid files</strong></p> </li> </ul> <pre><code class="language-python">caption.timecode.toTextgrid(outputDir, level=3)</code></pre> <ul> <li> <p><strong>Get the words, syllables and phonemes between</strong> <em>n</em> <strong>seconds/milliseconds</strong></p> </li> </ul> <p><em>The following Python code returns all the words between 0.2 and 0.6 seconds for which at least 50% of the word's total length is within the specified interval</em></p> <pre><code class="language-python">pprint(caption.getWords(0.20, 0.60, seconds=True, level=1, olapthr=50))</code></pre> <pre><code>... 404537 827239 Bruce US 0.9 404537_827239_Bruce_None_0-9.wav Eyeglasses, a cellphone, some keys and other pocket items are all laid out on the cloth. . [ { 'begin': 0.0, 'end': 0.7202778, 'overlapPercentage': 55.53412863758955, 'word': 'eyeglasses' } ] ...</code></pre> <ul> <li> <p><strong>Get the translations of the selected captions</strong></p> </li> </ul> <p><em>As for now, only japanese translations are available. We also used</em> <a href="http://www.phontron.com/kytea/">Kytea</a> <em>to tokenize and tag the captions translated with Google Translate</em></p> <pre><code class="language-python">captions = db.getImgCaptions(298817) for caption in captions: print('\n{}'.format(caption.text)) # Get translations and POS print('\tja_google: {}'.format(db.getTranslation(caption.captionID, "ja_google"))) print('\t\tja_google_tokens: {}'.format(db.getTokens(caption.captionID, "ja_google"))) print('\t\tja_google_pos: {}'.format(db.getPOS(caption.captionID, "ja_google"))) print('\tja_excite: {}'.format(db.getTranslation(caption.captionID, "ja_excite")))</code></pre> <pre><code> Birds wondering through grassy ground next to bushes. ja_google: 鳥は茂みの下に茂った地面を抱えています。 ja_google_tokens: 鳥 は 茂み の 下 に 茂 っ た 地面 を 抱え て い ま す 。 ja_google_pos: 鳥/名詞/とり は/助詞/は 茂み/名詞/しげみ の/助詞/の 下/名詞/した に/助詞/に 茂/動詞/しげ っ/語尾/っ た/助動詞/た 地面/名詞/じめん を/助詞/を 抱え/動詞/かかえ て/助詞/て い/動詞/い ま/助動詞/ま す/語尾/す 。/補助記号/。 ja_excite: 低木と隣接した草深いグラウンドを通って疑う鳥。 A flock of turkeys are making their way up a hill. ja_google: 七面鳥の群れが丘を上っています。 ja_google_tokens: 七 面 鳥 の 群れ が 丘 を 上 っ て い ま す 。 ja_google_pos: 七/名詞/なな 面/名詞/めん 鳥/名詞/とり の/助詞/の 群れ/名詞/むれ が/助詞/が 丘/名詞/おか を/助詞/を 上/動詞/のぼ っ/語尾/っ て/助詞/て い/動詞/い ま/助動詞/ま す/語尾/す 。/補助記号/。 ja_excite: 七面鳥の群れは丘の上で進んでいる。 Um, ah. Two wild turkeys in a field walking around. ja_google: 野生のシチメンチョウ、野生の七面鳥 ja_google_tokens: 野生 の シチメンチョウ 、 野生 の 七 面 鳥 ja_google_pos: 野生/名詞/やせい の/助詞/の シチメンチョウ/名詞/しちめんちょう 、/補助記号/、 野生/名詞/やせい の/助詞/の 七/名詞/なな 面/名詞/めん 鳥/名詞/ちょう ja_excite: まわりで移動しているフィールドの2羽の野生の七面鳥 Four wild turkeys and some bushes trees and weeds. ja_google: 4本の野生のシチメンチョウといくつかの茂みの木と雑草 ja_google_tokens: 4 本 の 野生 の シチメンチョウ と いく つ か の 茂み の 木 と 雑草 ja_google_pos: 4/名詞/4 本/接尾辞/ほん の/助詞/の 野生/名詞/やせい の/助詞/の シチメンチョウ/名詞/しちめんちょう と/助詞/と いく/名詞/いく つ/接尾辞/つ か/助詞/か の/助詞/の 茂み/名詞/しげみ の/助詞/の 木/名詞/き と/助詞/と 雑草/名詞/ざっそう ja_excite: 4羽の野生の七面鳥およびいくつかの低木木と雑草 A group of turkeys with bushes in the background. ja_google: 背景に茂みを持つ七面鳥の群 ja_google_tokens: 背景 に 茂み を 持 つ 七 面 鳥 の 群 ja_google_pos: 背景/名詞/はいけい に/助詞/に 茂み/名詞/しげみ を/助詞/を 持/動詞/も つ/語尾/つ 七/名詞/なな 面/名詞/めん 鳥/名詞/ちょう の/助詞/の 群/名詞/むれ ja_excite: 背景の低木を持つ七面鳥のグループ</code></pre> <p> </p>
French Emotional Speech Database - Oréau
<p>This document presents the French emotional speech database - Oréau, recorded in a quiet environment. The database is designed for general study of emotional speech and analysis of emotion characteristics for speech synthesis purposes. It contains 79 utterances which could be used in everyday life in the classroom. Between 10 and 13 utterances were written for each of the 7 emotions in French language by 32 non-professional speakers.</p> <p>2 versions are available, the first one contains 502 sentences. A perception test was performed to evaluate the recognition of emotions and their naturalness. 90% of utterances (434 utterances) were correctly identified and retained after the test and various analyses, which constitutes the second version of database.</p>
Five graphs on speech acts in late-Georgian satires
<p>Data, R scripts, and image files for 'Five graphs on speech acts in late-Georgian satires', cradledincaricature.com, 21 April 2015 If you have any questions or queries, please email me at james.baker@bl.uk. All responses will be logged for the benefit of future researchers. Share nicely. James Baker (james.baker@bl.uk)</p> <p><br /> This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.</p>
Speech circuit (Ferdinand de Saussure)
<p>Speech circuit as described by Ferdinand de Saussure in <em>Course in Gneral Linguistics</em> (1916). The image does not appear in the book, but is based on it.</p>
Hansard Speeches and Sentiment V1.0.1
<p>A public dataset of speeches in the Hansard, the record of the speeches, votes and legislation in the UK Parliament. The dataset provides information on each speech of ten words or longer, made in the House of Commons between 1980 and 2016, with information on the speaking MP, their party, gender and age at the time of the speech. The dataset also includes all speeches of ten words made from 1936 to 1979, without identifying information on the speaker.</p> <p>The speeches have been classified for sentiment using a total of five libraries from the R packages `sentimentr`, `syuzhet` and `lexicon`.</p> <p>The integrity of the public Hansard record is questionable at times, and while I have improved it, the data is presented 'as is'. More details on the dataset are available at: http://evanodell.com/datasets/hansard-data/</p>
A part-of-speech (POS) tagged corpus of Classical Tibetan
<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>
A rule based Tibetan part-of-speech (POS) tagger for the creation of gold standard training data
<p>This rule based Tibetan part-of-speech (POS) tagger was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. For a description of the tagger itself see Garrett et al. 2014. Note that the tagger must be used together with a lexicon (for example Hill & Garrett 2017a). One must use one's own script to tag all words with all tags in the lexicon and then apply the tagger to remove incorrect tags.</p> <p>On the associated corpus of 318,230 words (Hill & Garrett 2017b) the lexical tagger (i.e. simply applying all available tags to all words) tags 141,911 words with the correct unique tag, achieves as accuracy of 1.000 (by definition getting the right tag among others for each word) with an ambiguity of 2.73111. In contrast, the Rule Tagger tags 241,256 words with the correct unique tag, achieves an accuracy of 0.99893 and an ambiguity of 1.38577.</p> <p>Because this tagger does not achieve ambiguity 1.000 it is not suitable for tagging large scale corpora, but instead is useful for the creation of gold standard training data.</p> <p>N.B. In some rare cases the tagger removes all POS-tags.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.