Skip to main content
zenodoopen

Clotho-AQA dataset

<p>Clotho-AQA is an audio question-answering dataset consisting of 1991 audio samples taken from Clotho dataset [1]. Each audio sample has 6 associated questions collected through crowdsourcing. For each question, the answers are provided by three different annotators making a total of 35,838 question-answer pairs. For each audio sample, 4 questions are designed to be answered with &#39;yes&#39; or &#39;no&#39;, while the remaining two questions are designed to be answered in a single word. More details about the data collection process and data splitting process can be found in our following paper.</p> <p><em><strong>S. Lipping, P. Sudarsanam, K. Drossos, T. Virtanen &lsquo;Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering.&rsquo; </strong></em>The paper is available online at&nbsp;<a href="https://arxiv.org/pdf/2204.09634.pdf">2204.09634.pdf (arxiv.org)</a></p> <p>If you use the Clotho-AQA dataset, please cite the paper mentioned above. A sample baseline model to use the Clotho-AQA dataset can be found at <a href="https://github.com/partha2409/AquaNet">partha2409/AquaNet (github.com)</a></p> <p>To use the dataset,</p> <p>&bull; Download and extract <strong>&lsquo;audio_files.zip&rsquo;</strong>. This contains all the 1991 audio samples in the dataset.</p> <p>&bull; Download <strong>&lsquo;clotho_aqa_train.csv&rsquo;, &lsquo;clotho_aqa_val.csv&rsquo;, </strong>and <strong>&lsquo;clotho_aqa_test.csv&rsquo;</strong>. These files contain the train, validation, and test splits, respectively. They contain the audio file name, questions, answers, and confidence scores provided by the annotators.</p> <p><strong>License:</strong></p> <p>The audio files in the archive &lsquo;audio_files.zip&rsquo; are under the corresponding licenses (mostly CreativeCommons with attribution) of Freesound [2] platform, mentioned explicitly in the CSV file <strong>&rsquo;clotho_aqa_metadata.csv&rsquo;</strong> for each of the audio files. That is, each audio file in the archive is listed in the CSV file with meta-data. The meta-data for each file are:</p> <p>&bull; File name</p> <p>&bull; Keywords</p> <p>&bull; URL for the original audio file</p> <p>&bull; Start and ending samples for the excerpt that is used in the Clotho dataset</p> <p>&bull; Uploader/user in the Freesound platform (manufacturer)</p> <p>&bull; Link to the license of the file.</p> <p>The questions and answers in the files:</p> <p>&bull; clotho_aqa_train.csv</p> <p>&bull; clotho_aqa_val.csv</p> <p>&bull; clotho_aqa_test.csv</p> <p>are under the MIT license, described in the LICENSE file.</p> <p><strong>References:</strong></p> <p>[1] K. Drossos, S. Lipping and T. Virtanen, &quot;Clotho: An Audio Captioning Dataset,&quot; IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 736- 740, doi: 10.1109/ICASSP40776.2020.9052990.</p> <p>[2] Frederic Font, Gerard Roma, and Xavier Serra. 2013. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia (MM &#39;13). ACM, New York, NY, USA, 411-412. DOI: https://doi.org/10.1145/2502081.2502245</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
12
Reuse readiness
8
Engagement
0

Topics