USPDATRO: Underrepresented Speech Dataset from Romanian language Open Data
<p> USPDATRO<br> ==========</p> <p>Underrepresented Speech Dataset from Open Data: Case Study on the Romanian Language (USPDATRO) is a manually created Romanian language speech corpus.<br> It was created specifically using speech types that are underrepresented in other speech datasets.<br> Sources for this dataset are represented by open data available on multimedia platforms under a Creative Commons license.<br> The data was manually transcribed and aligned at segment level.<br> In addition to the text and audio files, we offer text annotations (lemmatization, part of speech tags, dependency parsing) in CoNLL-U Plus format.</p> <p>Each datasource is mentioned by URL in the metadata.csv file with associated license (a Creative Commons variant).</p> <p>Dataset structure:<br> - audio: Folder with audio segments in WAV format<br> - text: Folder with corresponding transcriptions<br> - conllup: Folder with corresponding token-based annotations<br> - metadata.csv: Contains information about each segment</p> <p>LICENSING</p> <p>This work (transcriptions, alignment, metadata, annotations) is provided under the license CC BY-NC-SA 4.0 (Attribution-NonCommercial-ShareAlike 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-sa/4.0/<br> and the full text here: https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode .<br> The original works considered for audio sources are available under their respective licenses (Creative Commons variants) as described in the metadata.csv file.</p> <p><br> CONTACT</p> <p>Research Institute for Artificial Intelligence "Mihai Drăgănescu", Romanian Academy<br> Web: http://www.racai.ro<br> Contact emails: vasile@racai.ro</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0