ADEPT: A Dataset for Evaluating Prosody Transfer
<p>The ADEPT dataset consists of prosodically-varied natural speech samples for evaluating prosody transfer in english text-to-speech models. The samples include global variations reflecting emotion and interpersonal attitude, and local variations reflecting topical emphasis, propositional attitude, syntactic phrasing and marked tonicity.</p> <p>Txt and wav files are organised according to the folder structure {speech_class}/{subcategory_or_interpretation}/{filename}, where filename follows the naming convention {speaker}_{utterance_id}. Speakers comprise 'ad00' (female voice) and 'ad01' (male voice). For classes with multiple interpretations, we provide the interpretations used in the disambiguation tasks in 'adept_prompts.json'.</p> <p>The corpus only includes prosodic variations that listeners are able to distinguish with reasonable accuracy, and we report these figures as a benchmark against which text-to-speech prosody transfer can be compared. More details can be found in our pre-print about the dataset (https://arxiv.org/abs/2106.08321).</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4