Skip to main content
zenodoopen

ADEPT: A Dataset for Evaluating Prosody Transfer

<p>The ADEPT dataset consists of prosodically-varied natural speech samples for&nbsp;evaluating prosody transfer in english text-to-speech&nbsp;models.&nbsp;The samples include global variations reflecting emotion and interpersonal attitude, and local variations reflecting topical emphasis, propositional attitude, syntactic phrasing and marked tonicity.</p> <p>Txt and wav files are organised according to the folder structure {speech_class}/{subcategory_or_interpretation}/{filename}, where filename&nbsp;follows the naming convention {speaker}_{utterance_id}. Speakers comprise &#39;ad00&#39; (female voice) and &#39;ad01&#39; (male voice). For classes with multiple&nbsp;interpretations, we provide the interpretations used in&nbsp;the disambiguation tasks in&nbsp;&#39;adept_prompts.json&#39;.</p> <p>The corpus only includes prosodic variations that listeners are able to distinguish with reasonable accuracy, and we report these figures as a benchmark against which text-to-speech prosody transfer can be compared. More details can be found in our pre-print about the dataset (https://arxiv.org/abs/2106.08321).</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics