THCHS-30 - Aligned IPA transcriptions
<p>This upload contains aligned IPA transcriptions for the <a href="https://www.openslr.org/18/">THCHS-30 dataset from OpenSLR</a>. Thereby, punctuation is added, silence marked and duration markers for each phoneme are assigned. Furthermore, the silence on the beginning and ending of each file is marked.</p> <p>The words were transcribed using <a href="https://pypi.org/project/pypinyin/">pypinyin</a> (v0.47.1) via <a href="https://pypi.org/project/dict-from-pypinyin/">dict-from-pypinyin</a> (v0.0.1) and mapped to IPA using the <code>pinyin-ipa-map-TONE3-all.json</code> mapping <a href="https://zenodo.org/record/7525638">from here</a>. The alignment was done using <a href="https://zenodo.org/record/6796264">Montreal Forced Aligner</a> (v2.0.5) and the acoustic model <a href="https://mfa-models.readthedocs.io/en/latest/acoustic/Mandarin/Mandarin%20MFA%20acoustic%20model%20v2_0_0a.html">Mandarin MFA</a> (v2.0.0a).</p> <p>Phoneme duration markers:</p> <ul> <li><code>˘</code> -> [0, 20) percentile (speaker-wise), e.g., <code>a˥˩˘</code></li> <li>(none) -> [20, 80) percentile (speaker-wise), e.g., <code>a˥˩</code></li> <li><code>ˑ</code> -> [80, 90) percentile (speaker-wise), e.g., <code>a˥˩ˑ</code></li> <li><code>ː</code> -> [90, inf) percentile (speaker-wise), e.g., <code>a˥˩ː</code></li> </ul> <p>Thereby each phoneme (including tones) was considered on its own, i.e., phonemes with different tones were not considered together for the percentile calculation.</p> <p>Silence markers:</p> <ul> <li><code>SILX</code> -> silence at start/end of a recording (aligned on all tiers)</li> <li><code>SIL0</code> -> no silence</li> <li><code>SIL1</code> -> [0, 33.33333333) percentile of all silences (speaker-wise)</li> <li><code>SIL2</code> -> [33.33333333, 66.66666666) percentile of all silences (speaker-wise)</li> <li><code>SIL3</code> -> [66.66666666, inf) percentile of all silences (speaker-wise)</li> </ul> <p>Files:</p> <ul> <li><code>grids.zip</code> <ul> <li>contains TextGrids for all audio files containing three tiers <code>words</code>, <code>phonemes</code> and <code>transcription</code> <ul> <li><code>words</code> contains the aligned Chinese words</li> <li><code>phonemes</code> contains the IPA pronunciations including silence markers at start and end (<code>SILX</code>)</li> <li><code>transcription</code> contains unaligned phonemes including punctuation and word boundary labels (<code>SIL0</code>)</li> </ul> </li> <li>the folder structure is equal to one from the dataset</li> </ul> </li> <li><code>grids-sdp.zip</code> <ul> <li>same as <code>grids.zip</code> except the folder structure is the one from <a href="https://pypi.org/project/speech-dataset-parser/">speech-dataset-parser</a> (v0.0.4)</li> </ul> </li> <li><code>preview-start/middle/end.png</code> <ul> <li>preview of the first TextGrid from speaker <code>A2</code> opened in Praat in different positions</li> </ul> </li> <li><code>words-vocabulary.txt</code> <ul> <li>contains all Chinese words from tier <code>words</code></li> </ul> </li> <li><code>phonemes-vocabulary.txt</code> <ul> <li>contains all phonemes from tier <code>phonemes</code></li> </ul> </li> <li><code>transcription-vocabulary.txt</code> <ul> <li>contains all phonemes/punctuation from tier <code>transcription</code></li> </ul> </li> <li><code>phonemes-durations.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code></li> </ul> </li> <li><code>phonemes-durations-simple.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers are ignored</li> </ul> </li> <li><code>phonemes-durations-simple-toneless.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers and tones are ignored</li> </ul> </li> <li><code>pronunciations-broad.dict</code> <ul> <li>contains the broad pronunciations for each word including punctuation but not duration markers</li> <li>e.g., <code>一下。 i˥ ɕ j a˥˩ 。</code></li> </ul> </li> <li><code>pronunciations-narrow.dict</code> <ul> <li>contains the narrow pronunciations for each word including punctuation, duration markers and weights (= occurrence) over all speakers</li> <li>e.g., <code>一下。 3 i˥ ɕ j a˥˩ː 。</code></li> </ul> </li> <li><code>pronunciations-narrow-speakers.zip</code> <ul> <li>contains the narrow pronunciations separated for each speaker</li> </ul> </li> <li><code>script.sh</code> <ul> <li>contains the script to reproduce all results</li> <li>error in line 32: replace with <code>speech-dataset-parser==0.0.4</code></li> </ul> </li> </ul>
ShareScore
48/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4