Skip to main content
zenodoopen

Voice of America: Ukrainian ASR Dataset of Broadcast Speech

<p>The dataset is based on public recordings of Voice of America (<a href="https://ukrainian.voanews.com">https://ukrainian.voanews.com</a>) extracted from their&nbsp;videos.</p> <p>The dataset contains&nbsp;398 hours of speech.</p> <p>The dataset is created by the ASR Corpus Creator (<a href="https://zenodo.org/record/7396705">https://zenodo.org/record/7396705</a>).</p> <p>The format of files: WAV with 16 kHz.</p> <p>&nbsp;</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
4
Engagement
4

Topics