Skip to main content
zenodoopen

Toward Visual Pronunciation Learning: A Speech-to-Articulatory Animation Pipeline Leveraging wav2vec 2.0 and rtMRI Landmarks

<p>Each Video Includes 5 sections:</p> <ul> <li>Top left-most: Phoneme Transcription from Original dataset of USC-TIMIT Dataset.</li> <li>Top middle-left: Generated Articulatory Animation from speech input from this paper.</li> <li>Top middle-right: Ground Truth from refined contour dataset from Refined rtMRI Landmark-Based Vocal Tract Contour Labels.</li> <li>Top right-most: rtMRI data from USC-TIMIT Dataset.</li> <li>Below: Sentences and words</li> </ul>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
8
Access
16
Reuse readiness
4
Engagement
4