Skip to main content
zenodoopen

THCHS-30 - Aligned IPA transcriptions

<p>This upload contains aligned IPA transcriptions for the&nbsp;<a href="https://www.openslr.org/18/">THCHS-30 dataset from OpenSLR</a>. Thereby, punctuation is added, silence marked and duration markers for each phoneme are assigned. Furthermore, the silence on the beginning and ending of each file is marked.</p> <p>The words were transcribed using&nbsp;<a href="https://pypi.org/project/pypinyin/">pypinyin</a>&nbsp;(v0.47.1) via&nbsp;<a href="https://pypi.org/project/dict-from-pypinyin/">dict-from-pypinyin</a>&nbsp;(v0.0.1) and mapped to IPA using the&nbsp;<code>pinyin-ipa-map-TONE3-all.json</code>&nbsp;mapping&nbsp;<a href="https://zenodo.org/record/7525638">from here</a>. The alignment was done using&nbsp;<a href="https://zenodo.org/record/6796264">Montreal Forced Aligner</a>&nbsp;(v2.0.5) and the acoustic model&nbsp;<a href="https://mfa-models.readthedocs.io/en/latest/acoustic/Mandarin/Mandarin%20MFA%20acoustic%20model%20v2_0_0a.html">Mandarin MFA</a>&nbsp;(v2.0.0a).</p> <p>Phoneme duration markers:</p> <ul> <li><code>˘</code>&nbsp;-&gt; [0, 20) percentile (speaker-wise), e.g.,&nbsp;<code>a˥˩˘</code></li> <li>(none) -&gt; [20, 80) percentile (speaker-wise), e.g.,&nbsp;<code>a˥˩</code></li> <li><code>ˑ</code>&nbsp;-&gt; [80, 90) percentile (speaker-wise), e.g.,&nbsp;<code>a˥˩ˑ</code></li> <li><code>ː</code>&nbsp;-&gt; [90, inf) percentile (speaker-wise), e.g.,&nbsp;<code>a˥˩ː</code></li> </ul> <p>Thereby each phoneme (including tones) was considered on its own, i.e., phonemes with different tones were not considered together for the percentile calculation.</p> <p>Silence markers:</p> <ul> <li><code>SILX</code>&nbsp;-&gt; silence at start/end of a recording (aligned on all tiers)</li> <li><code>SIL0</code>&nbsp;-&gt; no silence</li> <li><code>SIL1</code>&nbsp;-&gt; [0, 33.33333333) percentile of all silences (speaker-wise)</li> <li><code>SIL2</code>&nbsp;-&gt; [33.33333333, 66.66666666) percentile of all silences (speaker-wise)</li> <li><code>SIL3</code>&nbsp;-&gt; [66.66666666, inf) percentile of all silences (speaker-wise)</li> </ul> <p>Files:</p> <ul> <li><code>grids.zip</code> <ul> <li>contains TextGrids for all audio files containing three tiers&nbsp;<code>words</code>,&nbsp;<code>phonemes</code>&nbsp;and&nbsp;<code>transcription</code> <ul> <li><code>words</code>&nbsp;contains the aligned Chinese words</li> <li><code>phonemes</code>&nbsp;contains the IPA pronunciations including silence markers at start and end (<code>SILX</code>)</li> <li><code>transcription</code>&nbsp;contains unaligned phonemes including punctuation and word boundary labels (<code>SIL0</code>)</li> </ul> </li> <li>the folder structure is equal to one from the dataset</li> </ul> </li> <li><code>grids-sdp.zip</code> <ul> <li>same as&nbsp;<code>grids.zip</code>&nbsp;except the folder structure is the one from&nbsp;<a href="https://pypi.org/project/speech-dataset-parser/">speech-dataset-parser</a>&nbsp;(v0.0.4)</li> </ul> </li> <li><code>preview-start/middle/end.png</code> <ul> <li>preview of the first TextGrid from speaker&nbsp;<code>A2</code>&nbsp;opened in Praat in different positions</li> </ul> </li> <li><code>words-vocabulary.txt</code> <ul> <li>contains all Chinese words from tier&nbsp;<code>words</code></li> </ul> </li> <li><code>phonemes-vocabulary.txt</code> <ul> <li>contains all phonemes from tier&nbsp;<code>phonemes</code></li> </ul> </li> <li><code>transcription-vocabulary.txt</code> <ul> <li>contains all phonemes/punctuation from tier&nbsp;<code>transcription</code></li> </ul> </li> <li><code>phonemes-durations.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier&nbsp;<code>phonemes</code></li> </ul> </li> <li><code>phonemes-durations-simple.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier&nbsp;<code>phonemes</code>&nbsp;if all duration markers are ignored</li> </ul> </li> <li><code>phonemes-durations-simple-toneless.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier&nbsp;<code>phonemes</code>&nbsp;if all duration markers and tones are ignored</li> </ul> </li> <li><code>pronunciations-broad.dict</code> <ul> <li>contains the broad pronunciations for each word including punctuation but not duration markers</li> <li>e.g.,&nbsp;<code>一下。 i˥ ɕ j a˥˩ 。</code></li> </ul> </li> <li><code>pronunciations-narrow.dict</code> <ul> <li>contains the narrow pronunciations for each word including punctuation, duration markers and weights (= occurrence) over all speakers</li> <li>e.g.,&nbsp;<code>一下。 3 i˥ ɕ j a˥˩ː 。</code></li> </ul> </li> <li><code>pronunciations-narrow-speakers.zip</code> <ul> <li>contains the narrow pronunciations separated for each speaker</li> </ul> </li> <li><code>script.sh</code> <ul> <li>contains the script to reproduce all results</li> <li>error in line 32: replace with&nbsp;<code>speech-dataset-parser==0.0.4</code></li> </ul> </li> </ul>

ShareScore

48/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
20
Reuse readiness
8
Engagement
4

Topics