Skip to main content
zenodoopen

The Consonant Challenge Corpus

<p>The Consonant Challenge Corpus provides a dataset to support&nbsp;human-machine comparisons of consonant recognition in quiet and noise.&nbsp;Twelve female and 12 male native English talkers contributed to the corpus. All speakers produced each of the 24 English consonants / b, d, g, p, t, k, s, ʃ, f, v, &eth;, &theta;, ʧ, z, ʒ, h, ʤ, m, n, ŋ, w, r, j, l / in nine vowel contexts consisting of all possible combinations of the three vowels / iː / (as in &ldquo;beat&rdquo;), / uː / (as in &ldquo;boot&rdquo;), and / &aelig; / (as in &ldquo;bat&rdquo;). Each VCV was produced using both front and end stress (e.g. / &lsquo;&aelig; b &aelig; / vs / &aelig; b &lsquo;&aelig; /) giving a total of 24 (speakers) * 24 (consonants) * 2 (stress types) * 9 (vowel contexts) = 10368 tokens. Tokens are distributed into training, development and test sets for the purposes of automatic speech recognition experiments.</p> <p>The Consonant Challenge is described in this article:&nbsp;Cooke, M., Scharenborg, O. (2008), &ldquo;The Interspeech 2008 Consonant Challenge&rdquo;, Proceedings of Interspeech, Brisbane, Australia, September 2008.</p> <p>The&nbsp;distribution consists of the following elements:&nbsp;</p> <p>Technical description:</p> <ul> <li><em>readme</em></li> </ul> <p>Speech/noise waveforms</p> <ul> <li><em>train.zip</em> contains noise-free training data</li> <li><em>test.zip</em> contains the 7 test sets as well as practice items for perceptual tests, and MATLAB format files containing offsets identifying the time location of the speech token within the mixture</li> <li><em>test_binaural.zip</em> contains 2-channel wavs with the speech and noise on separate channels (left=noise, right=speech), for test sets 2-7 (test set 1 is noise-free)</li> <li><em>dev.zip</em> development set</li> <li><em>dev_binaural.zip</em>&nbsp;is the 2-channel version of the development set</li> </ul> <p>Phoneme segmentation data</p> <ul> <li><em>handsegm.91.mlf.txt:</em> 91 hand-segmented VCVs in HTK format.&nbsp;This set consists of at least three items per consonant in a context in which the first and the second vowel were identical, added to that were 19 randomly selected VCVs.</li> <li><em>segmentation_training.mlf.txt:</em> automatically generated phoneme segmentation of the clean training material&nbsp;in HTK format</li> <li><em>segmentation_testsets.zip</em>: zip file containing automatically generated phoneme segmentations of each test set&nbsp;in HTK format&nbsp;</li> </ul> <p>Automatic speech recognition</p> <ul> <li><em>asr.zip</em> contais&nbsp;scripts and models</li> </ul>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4

Topics