Skip to main content
zenodoopen

The Grid Audio-Visual Speech Corpus

<p>The Grid Corpus is a large multitalker audiovisual sentence corpus designed to support joint computational-behavioral studies in speech perception. In brief, the corpus consists of high-quality audio and video (facial) recordings of 1000 sentences spoken by each of 34 talkers (18 male, 16 female), for a total of 34000 sentences. Sentences are of the form &quot;put red at G9 now&quot;.</p> <p>audio_25k.zip &nbsp;contains the wav format utterances at a 25 kHz sampling rate in a separate directory per talker<br> alignments.zip provides word-level time alignments, again separated by talker<br> s1.zip, s2.zip etc contain .jpg videos for each talker [note that due to an oversight, no video for talker t21 is available]</p> <p>The Grid Corpus is described in detail in the paper jasagrid.pdf included in the dataset.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics