Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2 results for “Clotho”

Learn how ShareScore rates datasets ↗
zenodo36/100

Clotho Analysis Set

<p>This dataset is derived from the evaluation subset of <a href="https://zenodo.org/record/4783391">Clotho dataset</a>. It is designed to analyze the behavior of the captioning system under certain perturbation in order to try and identify some open challenges in automated audio captioning. The original audio clips are transformed with <a href="https://github.com/emilio-molina/audio_degrader">audio_degrader</a>. The transformations applied are the following:</p> <ul> <li> <p>Microphone response simulation</p> </li> <li> <p>Mixup with another clip from the dataset (ratio -6dB, -3dB and 0dB)</p> </li> <li> <p>Additive noise from <a href="https://zenodo.org/record/6026841">DESED</a> (ratio -12dB, -6dB, 0dB)</p> </li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Clotho-AQA dataset

<p>Clotho-AQA is an audio question-answering dataset consisting of 1991 audio samples taken from Clotho dataset [1]. Each audio sample has 6 associated questions collected through crowdsourcing. For each question, the answers are provided by three different annotators making a total of 35,838 question-answer pairs. For each audio sample, 4 questions are designed to be answered with &#39;yes&#39; or &#39;no&#39;, while the remaining two questions are designed to be answered in a single word. More details about the data collection process and data splitting process can be found in our following paper.</p> <p><em><strong>S. Lipping, P. Sudarsanam, K. Drossos, T. Virtanen &lsquo;Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering.&rsquo; </strong></em>The paper is available online at&nbsp;<a href="https://arxiv.org/pdf/2204.09634.pdf">2204.09634.pdf (arxiv.org)</a></p> <p>If you use the Clotho-AQA dataset, please cite the paper mentioned above. A sample baseline model to use the Clotho-AQA dataset can be found at <a href="https://github.com/partha2409/AquaNet">partha2409/AquaNet (github.com)</a></p> <p>To use the dataset,</p> <p>&bull; Download and extract <strong>&lsquo;audio_files.zip&rsquo;</strong>. This contains all the 1991 audio samples in the dataset.</p> <p>&bull; Download <strong>&lsquo;clotho_aqa_train.csv&rsquo;, &lsquo;clotho_aqa_val.csv&rsquo;, </strong>and <strong>&lsquo;clotho_aqa_test.csv&rsquo;</strong>. These files contain the train, validation, and test splits, respectively. They contain the audio file name, questions, answers, and confidence scores provided by the annotators.</p> <p><strong>License:</strong></p> <p>The audio files in the archive &lsquo;audio_files.zip&rsquo; are under the corresponding licenses (mostly CreativeCommons with attribution) of Freesound [2] platform, mentioned explicitly in the CSV file <strong>&rsquo;clotho_aqa_metadata.csv&rsquo;</strong> for each of the audio files. That is, each audio file in the archive is listed in the CSV file with meta-data. The meta-data for each file are:</p> <p>&bull; File name</p> <p>&bull; Keywords</p> <p>&bull; URL for the original audio file</p> <p>&bull; Start and ending samples for the excerpt that is used in the Clotho dataset</p> <p>&bull; Uploader/user in the Freesound platform (manufacturer)</p> <p>&bull; Link to the license of the file.</p> <p>The questions and answers in the files:</p> <p>&bull; clotho_aqa_train.csv</p> <p>&bull; clotho_aqa_val.csv</p> <p>&bull; clotho_aqa_test.csv</p> <p>are under the MIT license, described in the LICENSE file.</p> <p><strong>References:</strong></p> <p>[1] K. Drossos, S. Lipping and T. Virtanen, &quot;Clotho: An Audio Captioning Dataset,&quot; IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 736- 740, doi: 10.1109/ICASSP40776.2020.9052990.</p> <p>[2] Frederic Font, Gerard Roma, and Xavier Serra. 2013. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia (MM &#39;13). ACM, New York, NY, USA, 411-412. DOI: https://doi.org/10.1145/2502081.2502245</p>

openother-atApr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record