Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Clotho”
Clotho Analysis Set
<p>This dataset is derived from the evaluation subset of <a href="https://zenodo.org/record/4783391">Clotho dataset</a>. It is designed to analyze the behavior of the captioning system under certain perturbation in order to try and identify some open challenges in automated audio captioning. The original audio clips are transformed with <a href="https://github.com/emilio-molina/audio_degrader">audio_degrader</a>. The transformations applied are the following:</p> <ul> <li> <p>Microphone response simulation</p> </li> <li> <p>Mixup with another clip from the dataset (ratio -6dB, -3dB and 0dB)</p> </li> <li> <p>Additive noise from <a href="https://zenodo.org/record/6026841">DESED</a> (ratio -12dB, -6dB, 0dB)</p> </li> </ul>
Clotho-AQA dataset
<p>Clotho-AQA is an audio question-answering dataset consisting of 1991 audio samples taken from Clotho dataset [1]. Each audio sample has 6 associated questions collected through crowdsourcing. For each question, the answers are provided by three different annotators making a total of 35,838 question-answer pairs. For each audio sample, 4 questions are designed to be answered with 'yes' or 'no', while the remaining two questions are designed to be answered in a single word. More details about the data collection process and data splitting process can be found in our following paper.</p> <p><em><strong>S. Lipping, P. Sudarsanam, K. Drossos, T. Virtanen ‘Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering.’ </strong></em>The paper is available online at <a href="https://arxiv.org/pdf/2204.09634.pdf">2204.09634.pdf (arxiv.org)</a></p> <p>If you use the Clotho-AQA dataset, please cite the paper mentioned above. A sample baseline model to use the Clotho-AQA dataset can be found at <a href="https://github.com/partha2409/AquaNet">partha2409/AquaNet (github.com)</a></p> <p>To use the dataset,</p> <p>• Download and extract <strong>‘audio_files.zip’</strong>. This contains all the 1991 audio samples in the dataset.</p> <p>• Download <strong>‘clotho_aqa_train.csv’, ‘clotho_aqa_val.csv’, </strong>and <strong>‘clotho_aqa_test.csv’</strong>. These files contain the train, validation, and test splits, respectively. They contain the audio file name, questions, answers, and confidence scores provided by the annotators.</p> <p><strong>License:</strong></p> <p>The audio files in the archive ‘audio_files.zip’ are under the corresponding licenses (mostly CreativeCommons with attribution) of Freesound [2] platform, mentioned explicitly in the CSV file <strong>’clotho_aqa_metadata.csv’</strong> for each of the audio files. That is, each audio file in the archive is listed in the CSV file with meta-data. The meta-data for each file are:</p> <p>• File name</p> <p>• Keywords</p> <p>• URL for the original audio file</p> <p>• Start and ending samples for the excerpt that is used in the Clotho dataset</p> <p>• Uploader/user in the Freesound platform (manufacturer)</p> <p>• Link to the license of the file.</p> <p>The questions and answers in the files:</p> <p>• clotho_aqa_train.csv</p> <p>• clotho_aqa_val.csv</p> <p>• clotho_aqa_test.csv</p> <p>are under the MIT license, described in the LICENSE file.</p> <p><strong>References:</strong></p> <p>[1] K. Drossos, S. Lipping and T. Virtanen, "Clotho: An Audio Captioning Dataset," IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 736- 740, doi: 10.1109/ICASSP40776.2020.9052990.</p> <p>[2] Frederic Font, Gerard Roma, and Xavier Serra. 2013. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia (MM '13). ACM, New York, NY, USA, 411-412. DOI: https://doi.org/10.1145/2502081.2502245</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.