Skip to main content
zenodoopen

Clotho Analysis Set

<p>This dataset is derived from the evaluation subset of <a href="https://zenodo.org/record/4783391">Clotho dataset</a>. It is designed to analyze the behavior of the captioning system under certain perturbation in order to try and identify some open challenges in automated audio captioning. The original audio clips are transformed with <a href="https://github.com/emilio-molina/audio_degrader">audio_degrader</a>. The transformations applied are the following:</p> <ul> <li> <p>Microphone response simulation</p> </li> <li> <p>Mixup with another clip from the dataset (ratio -6dB, -3dB and 0dB)</p> </li> <li> <p>Additive noise from <a href="https://zenodo.org/record/6026841">DESED</a> (ratio -12dB, -6dB, 0dB)</p> </li> </ul>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics