Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “speech corpus”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)

<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from &#39;Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais&#39; de J.-F. Blad&eacute;, &#39;Coundes biarn&eacute;s, cou&eacute;ilhuts a&uuml;s pars&agrave;as mi&eacute;ytad&egrave;s dou p&eacute;ys d&eacute; Biarn&#39; de J.-V. Lalanne, &#39;Contes populaires du Languedoc&#39; de L. Lambert and &#39;Contes populaires recueillis dans la Grande-Lande&#39; de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

Oral cancer speech corpus for paper "Detecting and analysing spontaneous oral cancer speech in the wild"

<p>This is the oral cancer speech corpus used in the paper <em>&quot;Detecting and analysing spontaneous oral cancer speech in the wild&quot;.</em></p> <p><strong>Description</strong></p> <p>This dataset contains approximately 3 hours of oral cancer speech data collected from YouTube, including a file with additional metadata. We use this dataset to perform an oral cancer speech detection task in our paper.</p> <p><strong>Funding</strong></p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under Marie Sklodowska-Curie grant agreement No 766287. The Department of Head and Neck Oncology and surgery of the Netherlands Cancer Institute receives a research grant from Atos Medical (Horby, Sweden),<br> which contributes to the existing infrastructure for quality of life research.</p> <p><strong>Citation:</strong></p> <p>If you use this dataset please cite:</p> <pre><code>@misc{halpern2020detecting, title={Detecting and analysing spontaneous oral cancer speech in the wild}, author={Bence Mark Halpern and Rob van Son and Michiel van den Brekel and Odette Scharenborg}, year={2020}, eprint={2007.14205}, archivePrefix={arXiv}, primaryClass={eess.AS} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Speech Corpus of Interpreted Premier Press Conferences (SCIPPC)

<p>SCIPPC v1.0 is a parallel corpus of consecutive interpreting between Mandarin Chinese and English and vice versa in two Chinese premiers&rsquo; press conferences in March 2003&ndash;2007 and 2013&ndash;2017. The conferences were held after sessions of the National People&rsquo;s Congress and the Chinese People&rsquo;s Political Consultative Conference. They were moderated by spokespersons of the Congress and the Chinese Ministry of Foreign Affairs and attended by journalists, who asked the premiers questions.</p> <p>SCIPPC v1.0 includes source speeches by approximately 170 speakers and interpretations by six different staff interpreters of the Chinese Ministry of Foreign Affairs, who worked into their B language. It contains 192,209 tokens (source: 108,296, target: 83,913; Chinese: 112,528, English: 79,681) and 19 h 43 min 5 s of video recordings. It is fully transcribed and aligned at the recording&ndash;transcript and source&ndash;target transcript levels.</p>

opencc-by-sa-4.0Sep 2024View details →
zenodo44/100

AVbook, a high-frame-rate corpus of narrative audiovisual speech for investigating multimodal speech perception

<p><strong>Please cite</strong><br> Varano E, Guilleminot P, Reichenbach T. <em>AVbook, a&nbsp;high-frame-rate corpus of narrative audiovisual speech for investigating multimodal speech perception</em>. J Acoust Soc Am. 2023 May 1;153(5):3130. doi: 10.1121/10.0019460. PMID: 37249407.<br> <br> Seeing a speaker&#39;s face can help substantially in understanding them, in particular in challenging listening conditions. Research into the neurobiological mechanisms behind the audiovisual integration has recently begun to employ continuous natural speech. However, these efforts are impeded by a lack of high-quality audiovisual recordings of a speaker narrating a longer text. Here we seek to close this gap by developing AVbook, an audiovisual speech corpus designed for cognitive neuroscience studies and audiovisual speech recognition. The corpus consists of 3.6 hours of audiovisual recordings of two speakers, one male and one female, reading 59 passages from a narrative English text. The recordings were acquired at a high frame rate of 119.88 frames per second. The corpus includes a sets of multiple-choice questions to test attention to the different passages. We verified the efficacy of these questions in a pilot study. A short written summary is also provided for each recording. To enable audiovisual synchronization when presenting the stimuli, four videos of an electronic clapperboard were recorded with the corpus. The corpus is&nbsp;available for download to support research into the neurobiology of audiovisual speech processing as well as the development of computer algorithms for audiovisual speech recognition.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

German Political Speeches Corpus

<p>This text archive focuses on German political speeches held by top officials mostly from 1990 onwards, selected according to their political relevance. The currently included speeches come from the following sources:</p> <ul> <li>Official pages of the German <a href="http://www.bundespraesident.de/EN/Home/home_node.html">Presidency</a>, <a href="https://www.bundeskanzlerin.de/Webs/BKin/EN/Chancellery/federal_chancellery_node.html">Chancellery</a>, <a href="http://www.bundestag.de/en/parliament/presidium/function_neu">Bundestag</a>, <a href="https://www.auswaertiges-amt.de/en/">Ministry of Foreign Affairs</a></li> <li>Personal pages of the <a href="https://www.helmut-kohl.de/dokumente_reden.html">Helmut Kohl archive</a>, <a href="http://www.thierse.de/reden-und-texte/reden/">Wolfgang Thierse</a> and <a href="http://www.norbert-lammert.de/01-lammert/texte.php">Norbert Lammert</a></li> </ul> <p>This resource is available online:</p> <ul> <li><a href="https://www.dwds.de/r?corpus=politische_reden">Online queries on the DWDS website</a> and <a href="https://www.dwds.de/d/korpussuche">usage instructions</a> (the text base may be newer than the downloadable archives)</li> <li><a href="http://purl.org/corpus/german-speeches">http://purl.org/corpus/german-speeches</a></li> </ul> <p>The files below consist of texts with metadata encoded in <a href="http://xml.silmaril.ie/whatisxml.html">XML format</a>. For appropriate tooling see:</p> <ul> <li>Python tutorial using the speeches: <a href="https://www.timmer-net.de/2019/03/24/nlp_basics/">Natural Language Processing &mdash; Einsteigen und Loslegen!</a></li> <li><a href="http://www.corpusexplorer.de">CorpusExplorer</a>, corpus linguistics and text mining software featuring the speeches</li> <li><a href="https://github.com/adbar/german-nlp">List of off-the-shelf NLP tools for German</a></li> </ul> <p>This is work in progress, updated and extended versions will follow.</p>

opencc-by-sa-4.0Jun 2019View details →
zenodo40/100

MigParl. A Corpus of Speeches on Migration and Integration in Germany's Regional Parliaments

<p>MigParl is an indexed and linguistically annotated corpus of speeches on migration and integration affairs in Germany&rsquo;s regional parliaments (&ldquo;Landtage&rdquo;). The corpus has been prepared in the MigTex Project (principal investigators: Andreas Bl&auml;tte / University of Duisburg-Essen, Ruud Koopmans / Berlin Social Science Center), using the resources and the infrastructure of the <a href="http://polmine.github.io">PolMine Project</a>.</p> <p>MigTex was part of a larger joint project to establish the research community of the <em>German Centre for Migration and Integration Affairs</em> (<em>Deutsches Zentrum f&uuml;r Migration and Integrationsforschung</em> / DeZIM). Funding awarded by Germany&rsquo;s <em>Federal Ministry for Family Affairs, Senior Citizens, Women and Youth</em> (<em>Bundesministerium f&uuml;r Familie, Senioren, Frauen und Jugend</em> / BMFSFJ) is gratefully acknowleged.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

A part-of-speech (POS) tagged corpus of Classical Tibetan

<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"

<p>Dataset accompanying the paper &quot;<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>&quot;</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

<p>We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the population in Zambia. To assess its usefulness for training and testing ASR systems for Bemba, we explored different approaches; supervised pre-training (training from scratch), cross-lingual transfer learning from a monolingual English pre-trained model using DeepSpeech on the portion of the dataset and fine-tuning large scale self-supervised Wav2Vec2.0 based multilingual pre-trained models on the complete BembaSpeech corpus. From our experiments, the 1 billion XLS-R parameter model gives the best results. The model achieves a word error rate (WER) of 32.91%, results demonstrating that model capacity significantly improves performance and that multilingual pre-trained models transfers cross-lingual acoustic representation better than monolingual pre-trained English model on the BembaSpeech for the Bemba ASR. Lastly, results also show that the corpus can be used for building ASR systems for Bemba language</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Emozionalmente: a crowdsourced Italian speech emotional corpus

<p>This repository contains Emozionalmente: an extensive simulated speech emotional corpus in Italian. The dataset comprises 6,902 labeled samples acted out by 431 amateur actors, each verbalizing 18 different sentences to express the Big Six emotions (anger, disgust, fear, joy, sadness, surprise) plus neutrality. The labels represent the emotional communicative intention of the actors (i.e., the seven emotional states).</p> <p>Key details about the dataset:<br>- **Recording specifications**: The recordings were generally obtained with non-professional equipment. They are .wav files, mono-channel, with a sample size of 16 bits and a sample rate of 16,000 Hz. Each audio recording lasts 3.81 seconds on average (SD = 0.99 seconds).<br>- **Validation**: To validate the emotional content of the clips, 829 humans evaluated each audio recording, providing five evaluations per audio. The general Unweighted Average Recall (UAR) achieved by the evaluators was 66%, which is comparable to previous literature in the field.</p> <p>The repository includes the following additional resources:<br>1. **Demographic information**: Three .csv files describing the demographics of the actors and evaluators, as well as the emotions they expressed and recognized for each audio sample.<br>2. **Data splits**: A speaker-independent train-dev-test split, stratified by emotion, gender, and age.</p> <p>If you use this dataset, please cite the following paper:</p> <blockquote> <p>F. Catania, J. W. Wilke and F. Garzotto,<br><em>"Emozionalmente: A Crowdsourced Corpus of Simulated Emotional Speech in Italian,"</em><br>IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1142&ndash;1155, 2025.<br>doi: <a target="_new" rel="noopener">10.1109/TASLPRO.2025.3540662</a></p> </blockquote>

opencc-by-4.0May 2022View details →
zenodo40/100

ParlamentParla - Speech corpus of Catalan Parliamentary sessions

<p>This is the <a href="http://github.com/CollectivaT-dev/ParlamentParla">ParlamentParla</a> speech corpus for Catalan prepared by <a href="https://collectivat.cat/">Col&middot;lectivaT</a>. The audio segments were extracted from recordings the Catalan Parliament (<a href="https://www.parlament.cat/">Parlament de Catalunya</a>) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their <a href="https://www.parlament.cat/pcat/serveis-parlament/avis-legal/">terms of use</a>.</p> <p>Preparation of this corpus was partly supported by the <a href="http://cultura.gencat.cat/">Department of Culture</a> of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project <a href="http://aina.gencat.cat/">AINA</a> of the <a href="https://politiquesdigitals.gencat.cat/">Departament de Pol&iacute;tiques Digitals</a>.</p> <p>As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file.</p> <p>For more information, go to <a href="https://github.com/CollectivaT-dev/ParlamentParla">https://github.com/CollectivaT-dev/ParlamentParla</a> or mail info@collectivat.cat.</p> <p><strong>Revision log:</strong></p> <ul> <li> <p><em>2.0:</em> Major changes in the file structure; speaker ids with respective<br> genders added. The speakers of train, test and dev corpora do not overlap.<br> A major increase in size with a total time of 611 hours 43 minutes.</p> </li> <li> <p><em>1.0:</em> Much better quality due to improved segmentation, corpus separated<br> into clean and other.</p> </li> <li> <p><em>0.2:</em> First public release of approx. 320 hours.</p> </li> </ul>

opencc-by-4.0Oct 2021View details →
zenodo40/100

VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining

<p>The <em>VivesDebate-Speech</em> corpus contains the acoustic information of 29 different&nbsp;argumentative debates and the annotations of the segmentation (i.e., BIO tags) of the Argumentative Discourse Units identified in the spoken natural language discourse.&nbsp;</p>

opencc-by-nc-sa-4.0Sep 2022View details →
zenodo36/100

The Grid Audio-Visual Lombard Speech Corpus

<p>Lombard Grid is a bi-view audiovisual Lombard speech corpus which can be used to support joint computational/behavioral studies in speech perception. The corpus includes 54 talkers, with 100 utterances per talker (50 Lombard and 50 plain utterances). This dataset follows the same sentence format as the audiovisual&nbsp;<a href="https://asa.scitation.org/doi/10.1121/1.5042758">Grid corpus</a>, and can thus be considered as an extension of that corpus. The sentence sets used in the Lombard Grid corpus are unique, however, and have not been utilized by the Grid corpus.</p> <p>It offers two synchronised views of the talkers (front and side) to facilitate analysis of speech from different angles. A bespoke head-mounted camera system was used to collect both front and profile views of the talkers.</p> <p><strong>Statistics</strong>: 54 talkers: 30 female talkers and 24 male talkers; 5,400 (audio, front video and side video) utterances (16,200 files in total): 50% Lombard utterances, 50% plain reference utterances.</p> <p>The dataset is described in detail in the paper,</p> <p>Najwa Alghamdi, Steve Maddock, Ricard Marxer, Jon Barker and Guy J. Brown,, &quot;A corpus of audio-visual Lombard speech with frontal and profile views&quot;, The Journal of the Acoustical Society of America 143, El523 (2018)&nbsp;</p> <p>The paper is available online at <a href="http://eprints.whiterose.ac.uk/131924/">White Rose Online Research</a>.</p> <p>------------------------------------------------------------------------------------</p> <p><strong>Notes on Filenaming</strong></p> <p><strong>Filename format</strong></p> <p>SPKR_COND_UTTERANCE.wav|.mov - e.g., s8_p_sbbi9p.wav</p> <p>*SPKR = s1 to s55</p> <p>*COND = l or p, where l=&gt; Lombard, p=&gt; plain (i.e. non-Lombard)</p> <p>*UTTERANCE = 6-character Grid utterance code, e.g. &#39;pgag6a&#39; which means &#39;place green at g 6 again&#39;</p> <p><strong>Metadata format</strong></p> <p>*SPKR = s1 to s55</p> <p>*SESSION = 1 or 2</p> <p>*INDEX = 1 to 10 for ordering of the recording blocks</p> <p>*SUBINDEX = 1 to 10 for ordering of utterance in a 10-utterance block.</p> <p>*COND = l or r, where l=&gt; Lombard, p=&gt; plain (i.e. non-Lombard)</p> <p>*UTTERANCE = 6-character Grid utterance code, e.g. &#39;pgag6a&#39; which means &#39;place green at g 6 again&#39;</p> <p>If a sentence is spoken incorrectly then the filename will be</p> <p>_WRONG.wav e.g. s8_2_38_8_r_lrwizp_WRONG_lrbizp.wav</p> <p>*TRANS = the Grid utterance code for what was actually said.</p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

voiceHome-2 corpus - automatic speech recognition baseline - acoustic model

<p>This entry contains the acoustic model used for evaluation of distant-microphone speech recognition performance in:</p> <p>Nancy Bertin, Ewen Camberlein, Romain Lebarbenchon, Emmanuel Vincent, Sunit Sivasankaran, Irina Illina, Fr&eacute;d&eacute;ric Bimbot<br> <a href="https://hal.inria.fr/hal-01923108">VoiceHome-2, an extended corpus for multichannel speech processing in real homes</a><br> <em>Speech Communication</em>, 2019, 106, pp.68-78.&nbsp;<a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">&lang;10.1016/j.specom.2018.11.002&rang;</a></p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Corona speeches mini corpus

<div> <div> <div> <p>This data collection (mini corpus) collects fifteen speeches of three speakers, Emmanuel Macron, Pedro S&aacute;nchez, and Angela Merkel, with five speeches per speaker. Four of their speeches are from between March and June, 2020, and one speech per speaker is from October or November, 2020. The speeches share important parallels in content and the speakers have similar intentions.</p> </div> </div> </div>

opencc-by-nc-nd-4.0Dec 2023View details →
zenodo36/100

Corpus of Distorted Speech

<p>This dataset contains audio stimuli (wav format, 16 kHz) used to test the perception of speech that has undergone artificial distortion. The corpus consists of 240 sentences for each of eight types of distortion. The original (unmodified) sentences come from the public-domain Sharvard Corpus, sentence numbers 241-480.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

CUCO Database: A voice and speech corpus of patients who underwent upper airway surgery in pre- and post-operative states

<p>The data set comprises 3,800 speech audio files of 3 types of upper respiratory tract surgeries and 1 control set. The dataset has an average of <strong>35.51 +- 5.91</strong> audio recordings per patient. It provides valuable resources to the scientific community to systematically investigate the objective effects of upper respiratory tract surgery on voice and speech. &nbsp;</p> <p>This data set is a complete corpus comprising data from 107 Spanish Castilian speakers. This corpus encompasses voice and speech recordings from both control speakers and patients who underwent upper airway surgical procedures in pre- and post-operative stages. The surgeries in focus include <strong>Tonsillectomy</strong>, <strong>Functional Endoscopic Sinus Surgery</strong>, and <strong>Septoplasty</strong>, all consistently performed by a single surgeon.</p> <p>This corpus has been the basis for different previous studies to evaluate changes in voice and its quality due to surgery. The results do not suggest significant changes in the most relevant acoustic parameters studied for the voice, which is consistent with the initial hypothesis. However, the analysis of speech recordings remains open, with a special focus on the nasalised segments, which are expected to change due to surgical intervention.</p> <p>This data set also opens the way to study the effect of upper airway surgery on the performance of speaker recognition and identification methods, as well as to be used to test anti-spoofing methodologies to make them more robust.&nbsp;</p> <p>Please, if you use this database, cite this open-access paper where the data acquisition is explained and detailed:</p> <p>Hern&aacute;ndez-Garc&iacute;a, E., Guerrero-L&oacute;pez, A., Arias-Londo&ntilde;o, J. D., &amp; Godino-Llorente, J. I. (2024). A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states.&nbsp;<em>Scientific Data</em>,&nbsp;<em>11</em>(1), 746.</p>

opencc-by-nc-nd-4.0May 2024View details →
zenodo36/100

VoiceHome-2 corpus : A corpus dedicated to distant-microphone speech processing in domestic environments

<p><strong>Purpose: </strong></p> <p>This corpus includes reverberated, noisy speech signals spoken by 12 native French talkers in 4 houses (3 rooms per house) and recorded by an 8-microphone device at various angles and distances and in various noise conditions.</p> <p>This corpus stands apart from other corpora in the field by the number of rooms and homes considered by the diversity of acoustic conditions recorded and by the facts that it is publicly available at no cost.</p> <p><strong>Other materials:</strong></p> <ul> <li>Article : N. Bertin, E. Camberlein, R. Lebarbenchon, E. Vincent, S. Sivasankaran, I. Illina and F. Bimbot: <a href="https://hal.inria.fr/hal-01923108"><strong>VoiceHome-2, an extended corpus for multichannel speech processing in real homes</strong></a>, <em>Speech Communication</em>, Elsevier : North-Holland, 2019, 106, pp.68-78. <a href="https://dx.doi.org/10.1016/j.specom.2018.11.002">&lang;10.1016/j.specom.2018.11.002&rang;</a>.</li> <li>Code baseline to reproduce article&#39;s results: <ul> <li><a href="https://hal.inria.fr/hal-02963528">Localization and speech enhancement</a></li> <li><a href="https://doi.org/10.5281/zenodo.4079314">Acoustic models</a> and <a href="https://hal.inria.fr/hal-02963802">recognition scripts</a> for automatic speech recognition</li> </ul> </li> <li>Related software to execute the baseline: <ul> <li><a href="https://gitlab.inria.fr/bass-db/mbss_locate">MBSS Locate (v2.0)</a></li> <li><a href="https://gitlab.inria.fr/bass-db/fasst">FASST</a></li> </ul> </li> </ul> <p><strong>Documentation:</strong></p> <p>The corpus documentation is both available into the archive and hereafter by clicking on voiceHome-2_corpus_v1.0_documentation.pdf .</p> <p><strong>Terms of use</strong></p> <p>You may exploit the corpus for a non-commercial scientific purpose provided you mention it in any written work or software you derive from its use. Within a published article, paper or report, the corpus must appear in the bibliographical references.</p> <p><strong>Speaker records diffusion consent</strong></p> <p>All participants have given an informed and signed consent about public diffusion of recorded sentences.</p> <p><strong>Contact:</strong></p> <p>nancy [dot] bertin [at] irisa [dot] fr</p>

opencc-by-nc-sa-4.0May 2018View details →
zenodo36/100

voiceHome corpus: A corpus dedicated to distant-microphone speech processing in domestic environments

<p><strong>Purpose: </strong></p> <p>This corpus includes reverberated, noisy speech signals spoken by native French talkers in a lounge and recorded by an 8-microphone device at various angles and distances and in various noise conditions.</p> <p>Room impulse responses and noise-only signals recorded in various real rooms and homes and baseline speaker localization and enhancement software are also provided.</p> <p>This corpus stands apart from other corpora in the field by the number of rooms and homes considered and by the fact that it is publicly available at no cost.</p> <p>&nbsp;</p> <p><strong>Other materials:</strong></p> <ul> <li>Article: N. Bertin, E. Camberlein, E. Vincent, R. Lebarbenchon, S. Peillon, E. Lamand&eacute;, S. Sivasankaran, F. Bimbot, I. Illina, A. Tom, S. Fleury and E. Jamet: <a href="https://hal.inria.fr/hal-01343060"><strong>A French corpus for distant-microphone speech processing in real homes</strong></a>, Interspeech2016, Sep 2016, San Francisco, United States, 2016.</li> <li>Related software to reproduce article&#39;s results: <ul> <li><a href="https://gitlab.inria.fr/bass-db/mbss_locate">Multi-Channel BSS Locate (v1.3)</a></li> <li><a href="https://gitlab.inria.fr/bass-db/fasst">FASST (v2.2.1)</a></li> </ul> </li> </ul> <p><strong>Documentation (in french):</strong></p> <p>The corpus documentation is both available into the archive and hereafter by clicking on voiceHome_corpus_french_documentation_v1.2.pdf .</p> <p><strong>Terms of use</strong></p> <p>You may exploit the corpus for a non-commercial scientific purpose provided you mention it in any written work or software you derive from its use. Within a published article, paper or report, the corpus must appear in the bibliographical references.</p> <p><strong>Speaker records diffusion consent</strong></p> <p>All participants have given an informed and signed consent about public diffusion of recorded sentences.</p> <p><strong>New corpus version available : voiceHome-2 corpus</strong></p> <p>A new version of the corpus is available : <a href="https://doi.org/10.5281/zenodo.1252143"><strong>voiceHome-2 corpus web page</strong></a></p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Jul 2018View details →
zenodo36/100

A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus

<p>We produce Sejong-style morphological analysis and part-of-speech tagging results which have been the de facto&nbsp;standard for Korean language processing by using UDPipe (http://ufal.mff.cuni.cz/udpipe)&nbsp;</p> <p>&nbsp;</p> <p>udpipe --tokenize --tag sjmorph.model input &gt; output</p> <p>see&nbsp;https://github.com/jungyeul/sjmorph</p>

opencc-by-4.0May 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record