Clotho Analysis Set
<p>This dataset is derived from the evaluation subset of <a href="https://zenodo.org/record/4783391">Clotho dataset</a>. It is designed to analyze the behavior of the captioning system under certain perturbation in order to try and identify some open challenges in automated audio captioning. The original audio clips are transformed with <a href="https://github.com/emilio-molina/audio_degrader">audio_degrader</a>. The transformations applied are the following:</p> <ul> <li> <p>Microphone response simulation</p> </li> <li> <p>Mixup with another clip from the dataset (ratio -6dB, -3dB and 0dB)</p> </li> <li> <p>Additive noise from <a href="https://zenodo.org/record/6026841">DESED</a> (ratio -12dB, -6dB, 0dB)</p> </li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0