The Grid Audio-Visual Speech Corpus
<p>The Grid Corpus is a large multitalker audiovisual sentence corpus designed to support joint computational-behavioral studies in speech perception. In brief, the corpus consists of high-quality audio and video (facial) recordings of 1000 sentences spoken by each of 34 talkers (18 male, 16 female), for a total of 34000 sentences. Sentences are of the form "put red at G9 now".</p> <p>audio_25k.zip contains the wav format utterances at a 25 kHz sampling rate in a separate directory per talker<br> alignments.zip provides word-level time alignments, again separated by talker<br> s1.zip, s2.zip etc contain .jpg videos for each talker [note that due to an oversight, no video for talker t21 is available]</p> <p>The Grid Corpus is described in detail in the paper jasagrid.pdf included in the dataset.</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4