Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “speech synthesis”
FestCat speech synthesis dataset in Catalan, raw at 48kHz, 16bits/sample
<p>This dataset contains audio recordings from the 10 speakers of the FestCat project and an additional speaker "uri".</p> <p>The recordings are available in the "-raw" compressed archives and are provided in the following format:</p> <pre><code class="language-bash">SAMPLERATE=48000 # Hz BITSPERSAMPLE=16 ENCODING="SIGNED_INTEGER" AUDIOHEADER="RAW" # (no header) # For the prompts and utts: TEXT_ENCODING="ISO-8859-15"</code></pre> <p>These sampling conditions were downsampled from the original recordings captured at 96kHz and using 24 bits/sample for all the FestCat speakers. The additional speaker "uri" had recordings 16kHz and were here upsampled to 48kHz.</p> <p>The recordings have been automatically segmented. The segmentation results are available at the "-utts" archives.</p> <p>The text prompts are available in the "-prompts" archive.</p> <p>The text information is encoded using ISO-8859-15.</p> <p> </p>
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis
<p>This is the public release of the Samsung Open Mean Opinion Scores (SOMOS) dataset for the evaluation of neural text-to-speech (TTS) synthesis, which consists of audio files generated with a public domain voice from trained TTS models based on bibliography, and numbers assigned to each audio as quality (naturalness) evaluations by several crowdsourced listeners.<br><br><strong>Description</strong><br><br>The SOMOS dataset contains 20,000 synthetic utterances (wavs), 100 natural utterances and 374,955 naturalness evaluations (human-assigned scores in the range 1-5). The synthetic utterances are single-speaker, generated by training several Tacotron-like acoustic models and an LPCNet vocoder on the LJ Speech voice public dataset. 2,000 text sentences were synthesized, selected from Blizzard Challenge texts of years 2007-2016, the LJ Speech corpus as well as Wikipedia and general domain data from the Internet.<br>Naturalness evaluations were collected via crowdsourcing a listening test on Amazon Mechanical Turk in the US, GB and CA locales. The records of listening test participants (workers) are fully anonymized. Statistics on the reliability of the scores assigned by the workers are also included, generated through processing the scores and validation controls per submission page.</p> <p>To listen to audio samples of the dataset, please see <a href="https://innoetics.github.io/publications/somos-dataset/index.html">our Github page</a>.</p> <p>The dataset release comes with a carefully designed train-validation-test split (70%-15%-15%) with unseen systems, listeners and texts, which can be used for experimentation on MOS prediction.</p> <p><em>This version also contains the necessary resources to obtain the transcripts corresponding to all dataset audios.</em></p> <p><strong>Terms of use</strong></p> <ul> <li>The dataset may be used for <strong>research</strong> purposes only, for <strong>non-commercial</strong> purposes only, and may be distributed with the same terms.</li> <li>Every time you produce research that has used this dataset, please <strong>cite </strong>the dataset appropriately.</li> </ul> <p>Cite as:</p> <pre><code>@inproceedings{maniati22_interspeech, author={Georgia Maniati and Alexandra Vioni and Nikolaos Ellinas and Karolos Nikitaras and Konstantinos Klapsas and June Sig Sung and Gunu Jho and Aimilios Chalamandaris and Pirros Tsiakoulis}, title={{SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis}}, year=2022, booktitle={Proc. Interspeech 2022}, pages={2388--2392}, doi={10.21437/Interspeech.2022-10922} } </code></pre> <p><br><strong>References of resources & models used</strong></p> <p>Voice & synthesized texts:<br>K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017.</p> <p>Vocoder:<br>J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP, 2019.<br>R. Vipperla, S. Park, K. Choo, S. Ishtiaq, K. Min, S. Bhattacharya, A. Mehrotra, A. G. C. P. Ramos, and N. D. Lane, “Bunched lpcnet: Vocoder for low-cost neural text-to-speech systems,” in Proc. Interspeech, 2020.</p> <p>Acoustic models:<br>N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High quality streaming speech synthesis with low, sentence-length-independent latency,” in Proc. Interspeech, 2020.<br>Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al., “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech, 2017.<br>J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al., “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in Proc. ICASSP, 2018.<br>J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, “Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling,” arXiv preprint arXiv:2010.04301, 2020.<br>M. Honnibal and M. Johnson, “An Improved Non-monotonic Transition System for Dependency Parsing,” in Proc. EMNLP, 2015.<br>M. Dominguez, P. L. Rohrer, and J. Soler-Company, “PyToBI: A Toolkit for ToBI Labeling Under Python,” in Proc. Interspeech, 2019.<br>Y. Zou, S. Liu, X. Yin, H. Lin, C. Wang, H. Zhang, and Z. Ma, “Fine-grained prosody modeling in neural speech synthesis using ToBI representation,” in Proc. Interspeech, 2021.<br>K. Klapsas, N. Ellinas, J. S. Sung, H. Park, and S. Raptis, “WordLevel Style Control for Expressive, Non-attentive Speech Synthesis,” in Proc. SPECOM, 2021.<br>T. Raitio, R. Rasipuram, and D. Castellani, “Controllable neural text-to-speech synthesis using intuitive prosodic features,” in Proc. Interspeech, 2020.</p> <p>Synthesized texts from the Blizzard Challenges 2007, 2008, 2009, 2010, 2011, 2012, 2013, 2016:<br>M. Fraser and S. King, "The Blizzard Challenge 2007," in Proc. SSW6, 2007.<br>V. Karaiskos, S. King, R. A. Clark, and C. Mayo, "The Blizzard Challenge 2008," in Proc. Blizzard Challenge Workshop, 2008.<br>A. W. Black, S. King, and K. Tokuda, "The Blizzard Challenge 2009," in Proc. Blizzard Challenge, 2009.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2010," 2010.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2011," 2011.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2012," 2012.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2013," 2013.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2016," 2016.</p> <p><strong>Contact</strong></p> <p>Alexandra Vioni - a.vioni@samsung.com</p> <ul> <li>If you have any questions or comments about the dataset, please feel free to write to us.</li> <li>We are interested in knowing if you find our dataset useful! If you use our dataset, please email us and tell us about your research.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.