AnglistikVoices: L2 English speech dataset
<h1><strong>AnglistikVoices: an L2 English speech dataset</strong> </h1> <p>This repository contains an L2 (second language) English speech corpus consisting of <strong>74 minutes</strong> of recorded audio from <strong>15 non-native English speaking participants</strong>. The dataset was created as part of a university course, with all participants being students who are also the authors of this dataset.</p> <h2>Dataset Specifications</h2> <ul> <li><strong>Total participants</strong>: 15 non-native English speakers</li> <li><strong>Total audio duration</strong>: 74 minutes</li> <li><strong>Recordings per participant</strong>: 60 audio samples each</li> <li><strong>Sentence alignment</strong>: Available for 8 out of 15 participant</li> <li><strong>Recording equipment</strong>: Audio-Technica ATM75 microphone</li> <li><strong>Stimuli: </strong>All sentences are from the Artie Bias Corpus (https://github.com/artie-inc/artie-bias-corpus)</li> <li><strong>Recording environment</strong>: Recording booth</li> </ul> <p>The dataset contains individual recordings of non-native English speakers <strong>organized by participant ID</strong>. For 8 participants, sentence-level alignments are provided. All recordings were captured in a controlled acoustic environment using Audio-Technica ATM75 microphone to ensure high audio quality. </p> <p>The recordings consist of spoken English utterances from each participant. <strong>Detailed linguistic profiles for each participant are available in the metadata.xlsx file</strong>, which is indexed by participant ID and contains information on native language, proficiency level, language learning history, and other relevant linguistic background data.</p> <p>The audio files are organized by participant ID, matching the identifiers used in the metadata file for easy cross-referencing between the audio recordings and participant linguistic profiles.</p> <h2>Authors and Contributors</h2> <p>This dataset was created by the student participants themselves as part of their coursework.</p> <p><strong>Course Instructor</strong>: Akhilesh Kakolu Ramarao</p> <p><strong>Teaching Assistant</strong>: Anna Sophia Stein</p> <p>If you have any questions, you can contact: kakolura@hhu.de</p> <p>If you use this dataset in your research, please cite:</p> <pre><code>@dataset{kakolu_ramarao_2024_anglistikvoices, author = {Kakolu Ramarao, A. and Stein, A. S. and Tahiri, A. and Rodrigues, D. C. and Antonia Weismann, C. and Schäfer, O. S. and Kaczor, J. and Tran, N. H. and Elena Telaar, C. and Bauer, L. and Jütten, M. and Mafuta, C. and Agelopoulou, V. V. and Grabowski, Q. A. G.}, title = {AnglistikVoices: L2 English speech dataset}, publisher = {Zenodo}, version = {v1.0.0}, year = {2024}, month = jun, doi = {10.5281/zenodo.12525952}, url = {https://doi.org/10.5281/zenodo.12525952}, note = {LabPhon 19, Hanyang Institute for Phonetics and Cognitive Sciences of Language (HIPCS), Hanyang University in Seoul, Korea} }</code></pre>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4