Nos_Celtia-GL: Galician TTS corpus
<p><em><strong>This corpus is publicly accessible upon accepting T&Cs and requesting access.</strong></em></p> <p><br>Galician TTS single speaker corpus of approximately 25 hours of speech.</p> <p>Nos_Celtia-GL is a phonetically and morphosyntactically rich corpus of 20,000 phrases (approximately 200,000 words) comprising two subcorpora: a previously compiled corpus created by the Grupo de Tecnoloxías Multimedia (GTM), together with the Centro Ramón Piñeiro para a Investigación en Humanidades (CRPIH), and a corpus compiled by the Nós Project from multi-domain texts.</p> <p>The text corpus statistics are detailed in the table below:</p> <table> <tbody> <tr> <td><strong>Subcorpus</strong></td> <td> <strong>Sentence no</strong>.</td> <td><strong>Word no. </strong></td> <td><strong>Sentence length (words)</strong></td> <td><strong>Sentence domain / type</strong></td> </tr> <tr> <td>GTM</td> <td>10,000</td> <td>121,726</td> <td>1-44</td> <td> <ul> <li>Journalistic (written) text</li> <li>Manually designed sentences (interrogative, exclamative, imperative, lists of numbers…)</li> </ul> </td> </tr> <tr> <td>Nós</td> <td>10,000</td> <td>99,622</td> <td> 1-36</td> <td> <ul> <li>21,8% transcripts of oral discourse</li> <li>17,5% dictionary definitions</li> <li>12.7% transcripts of parliamentary speeches</li> <li>20% transcripts of news broadcasts</li> <li>28% short (<4 words), interrogative, exclamative, imperative, and elliptical sentences</li> </ul> </td> </tr> </tbody> </table> <p> </p> <p>While the Nós subcorpus has undergone a thorough linguistic review, we have decided not to adapt the GTM corpus to the current grammatical norms of the Galician language with a view to obtaining a parallel corpus to the previously recorded <a href="../record/8027725">CRPIH_UVigo-GL-Voices</a>.</p> <p>Nos_Celtia-GL was recorded in a controlled environment (recording studio) by a professional female voice talent selected among four speakers through a perceptual listening test in which more than 50 participants assessed the speakers' clarity, prosody, likeability, and language proficiency.</p> <p>The file naming scheme of the audio files consists of a series of lowercase elements indicating the type of audio (raw), the creators of the corpus (nos), the name of the voice (celtia), and the ISO code for the Galician language (gl), followed by a 5-digit number identifying the utterance. All components are separated by underscores (e. g., raw_nos_celtia_gl_00001.wav).</p> <p>Metadata is provided in "metadata.csv". This file consists of one record per line, delimited by the vertical bar character (0x7c). The fields are:</p> <p> 1. Audio file: name of the corresponding .wav file</p> <p> 2. Transcription: non-normalized text read by speaker (UTF-8)</p> <p>The audio files are available in the format in which they were originally recorded, 48 kHz, 16-bit WAV format, and amount to approximately 25 hours.</p> <p>Version 1.0.0 contains the raw sound files with no editing nor normalization, together with the corresponding text.</p> <p>For more information, please go to <a href="https://nos.gal/">https://nos.gal/</a> or contact the Nós project at <a href="mailto:proxecto.nos@usc.gal">proxecto.nos@usc.gal</a>.</p> <p><strong>Terms and conditions</strong></p> <p>The property to the speech data contained in this dataset has been transferred to the University of Santiago de Compostela (USC) for the duration of 15 years. Starting 30/11/2037, this data will be removed. After this date, the USC is not liable for any use by third parties who might have downloaded the dataset. </p> <p><strong>Citing</strong></p> <p>Please refer to our paper for more details: <a title="https://www.isca-archive.org/iberspeech_2024/garciadiaz24_iberspeech.html" href="https://www.isca-archive.org/iberspeech_2024/garciadiaz24_iberspeech.html" target="_blank" rel="noreferrer noopener">Nos_Celtia-GL: an Open High-Quality Speech Synthesis Resource for Galician</a></p> <p>If you use this data in your work, please cite: García Díaz, N., Vázquez Abuín, M., Magariños, C., Vladu, A.I., Moscoso Sánchez, A., Fernández Rei, E. (2024) Nos_Celtia-GL: an Open High-Quality Speech Synthesis Resource for Galician. Proc. IberSPEECH 2024, 91-95, doi: 10.21437/IberSPEECH.2024-19</p> <p><strong>Funding and acknowledgements</strong></p> <p>"The Nós project: Galician in the society and economy of Artificial Intelligence" is possible thanks to the funding resulting from the agreement 2021-CP080 between the Xunta de Galicia and the University of Santiago de Compostela, and thanks to the Investigo program, within the National Recovery, Transformation and Resilience Plan, within the framework of the European Recovery Fund (NextGenerationEU).</p> <p>We would like to thank the speaker, Consuelo Díaz Isorna, for kindly providing her voice to this project.</p> <p>We would also like to thank the following entities for their kind collaboration in providing the data for the text corpus: <a href="http://gtm.uvigo.es/en/">Grupo de Tecnoloxías Multimedia (GTM)</a>, <a href="https://www.cirp.es/">Centro Ramón Piñeiro para a Investigación en Humanidades (CRPIH)</a>, <a href="https://academia.gal/inicio">Real Academia Galega</a>, <a href="https://www.crtvg.es/">Corporación Radio Televisión de Galicia S.A.</a>, <a href="https://www.parlamentodegalicia.gal/">Parlamento de Galicia</a>, and the <a href="http://ilg.usc.es/ago/">Arquivo do Galego Oral (ILG) project</a>.</p> <p>Our gratitude also to Xoán Carlos Goris García, Elia Lago Pereira and Alicia López Besteiro for reviewing part of the audio corpus.</p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 4