Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Audio captioning”
Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024
<h1><strong>Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024</strong></h1> <h3><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2"><em>Navigating the World of Wikidata for Research, Science and Cultural Heritage</em></a></h3> <h3><em><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht" target="_blank" rel="noopener">Wikidata's Twelfth Birthday: Workshop in Maastricht </a></em></h3> <p>Wikidata is a free, collaborative, multilingual database, collecting structured open data for anyone in the world to use. It also plays a crucial role in supporting Wikimedia projects, such as Wikipedia and Wikimedia Commons. Over the last 12 years it has strongly increased in popularity among the scientific and cultural heritage communities.</p> <p>In this 2,5 hours workshop you will learn the basics of working with Wikidata, both in theory and practice. You will learn</p> <ol> <li>The basics of Wikidata: A first look at what Wikidata is and how it works, both technically and socially (Wikidata community)</li> <li>How Wikidata can be relevant for research, science and cultural heritage (GLAM), and</li> <li>First steps in contributing to Wikidata yourself, with a focus on the topic of UM professors from past and present.</li> </ol> <p>As part of the <a title="Wikidata:Twelfth Birthday" href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday">Wikidata 12th Birthday celebrations</a> this workshop is open to academics, researchers, students, and professionals interested in working with Wikidata in the intersection of open data, research, and science. Whether you are new to Wikidata or looking to deepen your understanding, this session will provide valuable insights for improving your work.</p> <h2><strong>Workshop outline</strong></h2> <h3><strong>Part 1: Theory, Wikidata basics (45-60 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><em>1) Wikidata basics</em></p> <ul> <li>What is Wikidata?</li> <li>What are the principles of Wikidata?</li> <li>How are things described in Wikidata?</li> <li>Who builds Wikidata? - The Wikidata community</li> </ul> <p><em>2) Wikidata for research, science and cultural heritage</em></p> <ul> <li>To what extent is Wikidata used throughout science, research and GLAM?</li> <li>Six anecd<em>a</em>tic cases <ol> <li>Scientometrics - Scholia</li> <li>Life and biomedical sciences</li> <li>Astronomy</li> <li>Language technology / AI / LLMs</li> <li>GLAM – KB collection highlights</li> <li>Representation of (female) scientists</li> </ol> </li> </ul> <h3><strong>Break (15 minutes)</strong></h3> <h3><strong>Part 2: Practice, contributing to Wikidata (75-90 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm?download=1" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> <li><strong>Handout for participants in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.docx?download=1" rel="nofollow">Word</a> or <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.pdf?download=1" rel="nofollow">PDF</a></strong> <strong>are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><strong> </strong>The goals of this hands-on part are:</p> <ul> <li>Get familiar with basic data editing via the Wikidata interface</li> <li>Understand WD data models and structures related to professors (of Maastricht University)</li> <li>Extend existing <a title="Wikidata:Wiki-wetenschappers/Universiteit Maastricht/hoogleraren" href="https://www.wikidata.org/wiki/Wikidata:Wiki-wetenschappers/Universiteit_Maastricht/hoogleraren">Wikidata items about UM professors</a>, based on information in public sources.</li> <li>If time allows: Create new Wikidata items about UM professors</li> </ul> <p>The visual slides and the textual handout explain the same content, blocks and exercises, albeit in a slightly different order.<strong> </strong></p> <h2><strong>Required preparation</strong></h2> <p>To make optimal use of our time, participants must create a Wikidata account in the weeks before the workshop. See <a href="https://www.wikidata.org/w/index.php?title=Special:CreateAccount" target="_blank" rel="noopener">https://www.wikidata.org/w/index.php?title=Special:CreateAccount</a>.</p> <p>This is important because very fresh accounts may have limited editing rights. Furthermore only 6 Wikidata accounts can be created per day from UM IP addresses, so creating a lot of new accounts during the workshop might overstretch this limit.</p> <h2><strong>Workshop leader</strong></h2> <p>This workshop was given by <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Olaf Janssen</a>, the Wikimedia coordinator of the <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Koninklijke Bibliotheek</a>, the national library of the Netherlands.</p> <p>In this role he stimulates and facilitates collaboration between the collections, knowledge, open data and staff of the KB on the one hand, and the projects of the Wikimedia movement, such as Wikipedia, Wikimedia Commons and Wikidata on the other. He is also active as a volunteer within the community. Feel free to contact Olaf via olaf.janssen(at)<a href="http://kb.nl">kb.nl</a></p> <h2><strong>Materials on Wikimedia Commons</strong></h2> <p>Photos , videos and presentations related to this event can be found on Wikimedia Commons: <a title="c:Category:Wikidata Workshop at Maastricht University, 15 October 2024" href="https://commons.wikimedia.org/wiki/Category:Wikidata_Workshop_at_Maastricht_University,_15_October_2024">Category:Wikidata Workshop at Maastricht University, 15 October 2024</a></p> <h2>Relevant URLs </h2> <ul> <li><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht">https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht </a></li> <li><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2" target="_blank" rel="noopener">https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2</a> + <a href="https://web.archive.org/web/20240926154021/https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2/">archived version</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/">https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/</a></li> </ul> <p>Earlier LinkedIn posts (April-May 2024, before rescheduling the worlshop to October)</p> <ul> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/</a></li> </ul> <h3> </h3>
Audio captioning DCASE 2020 evaluation (testing) split
<p>This is the <strong>evaluation split for Task 6, Automated Audio Captioning, in DCASE 2020 Challenge</strong>. </p> <p>This evaluation split is the Clotho testing split, which is thoroughly described in the corresponding paper: </p> <p><em>K. Drossos, S. Lipping and T. Virtanen, "Clotho: an Audio Captioning Dataset," IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 736-740, doi: 10.1109/ICASSP40776.2020.9052990.</em></p> <p>available online at: <a href="https://arxiv.org/abs/1910.09387">https://arxiv.org/abs/1910.09387</a> and at: <a href="https://ieeexplore.ieee.org/document/9052990 ">https://ieeexplore.ieee.org/document/9052990 </a></p> <p>This evaluation split is meant to be used for the purposes of the Task 6 at the scientific challenge DCASE 2020. This split it is not meant to be used for developing audio captioning methods. For developing audio captioning methods, you should use the development and evaluation splits of Clotho. </p> <p>If you want the development and evaluation splits of Clotho dataset, you can find them also in Zenodo, at: <a href="https://zenodo.org/record/3490684">https://zenodo.org/record/3490684</a></p> <p>--------------------------------------------------------------------------------------------------------</p> <p><strong>== License ==</strong></p> <p>The audio files in the archives:</p> <ul> <li>clotho_audio_test.7z </li> </ul> <p>and the associated meta-data in the CSV file:</p> <ul> <li>clotho_metadata_test.csv</li> </ul> <p>are under the corresponding licences (mostly CreativeCommons with attribution) of Freesound [1] platform, mentioned explicitly in the CSV file for each of the audio files. That is, each audio file in the 7z archive is listed in the CSV file with the meta-data. The meta-data for each file are: </p> <ul> <li>File name</li> <li>Start and ending samples for the excerpt that is used in the Clotho dataset</li> <li>Uploader/user in the Freesound platform (manufacturer)</li> <li>Link to the licence of the file</li> </ul> <p>--------------------------------------------------------------------------------------------------------</p> <p><strong>== References ==</strong><br> [1] Frederic Font, Gerard Roma, and Xavier Serra. 2013. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia (MM '13). ACM, New York, NY, USA, 411-412. DOI: https://doi.org/10.1145/2502081.2502245</p>
Places Audio Captions (Japanese) 100k
<p>The Places Audio Caption (Japanese) 100K Corpus contains approximately 100,000 Japanese spoken captions for natural images drawn from the Places 205 image dataset.</p> <p>This speech corpus was collected to investigate the learning of spoken language (words, sub-word units, higher-level semantics, etc.) from visually-grounded speech. For a description of the corpus, see:</p> <pre><code>@INPROCEEDINGS{Ohishi2020trilingual, author={Ohishi, Yasunori and Kimura, Akisato and Kawanishi, Takahito and Kashino, Kunio and Harwath, David and Glass, James}, booktitle={ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, title={Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms}, year={2020}, pages={4352-4356}, }</code></pre> <p>The corpus only includes audio recordings, and not the associated images. You will need to separately download the Places image dataset <a href="http://places.csail.mit.edu/">here</a>.</p> <p>The data is distributed under the Creative Commons Attribution-ShareAlike (CC BY-SA) license <a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">(link)</a>.</p> <p>If you use this data in your own publications, please cite the paper above.</p>
Audio Caption Dataset (Hospital & Car)
<p>This dataset consists of the Hospital scene of our Audio Caption dataset. Details can be seen in our paper <a href="https://arxiv.org/abs/1902.09254">Audio Caption: Listen and Tell</a> published at ICASSP2019. </p> <p>Car scene, detailed in <a href="https://arxiv.org/abs/1905.13448">Audio Caption in a Car Setting with a Sentence-Level Loss</a> published at ISCSLP 2021.</p> <p>Original captions in Mandarin Chinese, with English translations provided. </p>
SiVi-CAFE dataset - Sighted and Visually-impaired Captions for Audio in Finnish and English
<p>This is a dataset containing <strong>audio captions</strong> for audio files of the <a href="../records/2589280">TAU Urban Acoustic Scenes 2019 development</a> dataset (airport, public square, and park) for 10 cities. </p> <p>The files were annotated using a web-based tool as presented in:</p> <p>Martin Morato, I., & Mesaros, A. (2021). <a href="https://doi.org/10.5281/zenodo.5770113" target="_blank" rel="noopener"><em>Diversity and bias in audio captioning datasets</em></a>. In F. Font, A. Mesaros, D. P.W. Ellis, E. Fonseca, M. Fuentes, & B. Elizalde (Eds.), Proceedings of the 6th Workshop on Detection and Classication of Acoustic Scenes and Events (DCASE 2021) (pp. 90-94)</p> <p>Each file is annotated by multiple annotators that provided a one-sentence description of the audio content.</p> <p><br>Data is provided in csv files:</p> <ul> <li>sighted-EN-bias-original</li> <li>sighted-FI-bias-translated</li> <li>sighted-EN-no_bias-original</li> <li>sighted-FI-no_bias-translated</li> <li>visually_impaired-FI-original</li> <li>visually_impaired-EN-translated</li> <li>sighted-FI-original</li> <li>sighted-EN-translated</li> </ul> <p>original = original descriptions, non-translated<br>translated = Translated descriptions using <a href="https://www.deepl.com/pro-api">automatic deep learning tool</a></p> <p>900 annotated audio files, Finnish audio descriptions provided by visual-impaired and sighted people.<br>2050 annotated audio files, English audio descriptions provided by international students (not-necessarily English native-speakers).<br>3930 annotated audio files, English audio descriptions provided by international students (not-necessarily English native-speakers) biased by the provided audio tags.</p> <p> </p> <p>The audio files can be downloaded from https://zenodo.org/record/2589280 and are covered by their own license.</p>
EmotionCaps: A Synthetic Emotion-Enriched Audio Captioning Dataset
<p>Version 1.0, October 2024</p> <h2>Created by</h2> <p>Mithun Manivannan (1), Vignesh Nethrapalli (1), Mark Cartwright (1)</p> <ol> <li>Sound Interaction and Computer Lab, New Jersey Institute of Technology</li> </ol> <h2>Publication</h2> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Manivannan, M., Nethrapalli, V., Cartwright, M. EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation. arXiv preprint arXiv:2410.12028, 2024.</p> <h2>Description</h2> <p>EmotionCaps is a ChatGPT-assisted, weakly-labeled audio captioning dataset developed to bridge the gap between soundscape emotion recognition (SER) and automated audio captioning (AAC). Created through a three-stage pipeline, the dataset leverages ground-truth annotations from AudioSet SL, which are enhanced by ChatGPT using tailored prompts and emotions assigned via a soundscape emotion recognition model trained on Emo-Soundscapes Dataset. It comprises four subsets of captions for 120,071 audio clips, each reflecting a different prompt variation: WavCaps-like, Scene-Focused, Emotion Addon, and Emotion Rewrite. The average word counts for these subsets are: WavCaps-like (12.61), Scene-Focused (14.04), Emotion Addon (18.35), and Emotion Rewrite (18.65). The increase in word count for the emotion prompts illustrates the difference in sentence length when integrating emotion information into the captions.</p> <h2>Audio Data</h2> <p>The audio data is from AudioSet SL, the strongly-labled subset of 120,071 audio clips from the larger AudioSet dataset.</p> <h2>Synthetic Captions</h2> <p>The synthetic captions were generated using a three-stage pipeline, beginning with training a soundscape emotion recognition model. This model assesses the valence and arousal of each audio clip, mapping the resulting vector to an emotion identifier. Next, we leveraged the ground-truth annotations from AudioSet SL, and extracted the list of sound events. Using these sound events, we employed ChatGPT to create different variations of captions by applying distinct prompts.</p> <p>We first used the WavCaps prompt for AudioSet SL as a base, the output of which we call WavCaps-like. Building on this, we created three new prompt variations (1) <strong>scene-focused</strong> which is a modified WavCaps prompt that describes the scene, (2) <strong>emotion addon</strong> which is an extension of the scene-Focused prompt, where an emotion is appended to the list of sound events to guide the caption generation, and (3) <strong>emotion rewrite</strong> which consists of two-step prompt where ChatGPT first generates the scene-focused caption, then is instructed to rewrite it with a specific emotion in mind.</p> <p>Using these four prompt styles — WavCaps, Scene-Focused, Emotion Addon, and Emotion Rewrite — along with the AudioSet SL sound events and predicted emotions, we employed ChatGPT-3.5 Turbo to generate four corresponding caption variations for the dataset.</p> <p>Each caption variation has been organized into separate CSV files for clarity and accessibility. All files correspond to the same set of audio clips from AudioSet SL, with the key distinction being the caption variation associated with each clip. The different subsets are designed to be used independently, as they each fulfill specific roles in understanding the impact of emotion in audio captions.</p> <ul> <li> <p><strong>wavcaps-like.csv</strong>: Contains captions generated using the WavCaps prompt, serving as the baseline before emotion is introduced.</p> </li> <li> <p><strong>scene-focused.csv</strong>: Provides captions focused on describing the scene or environment of the audio clip, without emotion integration.</p> </li> <li> <p><strong>emotion-addon.csv</strong>: Captions where emotion data is appended to the scene-focused base caption.</p> </li> <li> <p><strong>emotion-rewrite.csv</strong>: Captions that are completely rewritten based on the scene-focused base caption and the assigned emotion.</p> </li> </ul> <p>This structure allows users to explore how emotional content influences captioning models by comparing the variations both with and without emotional enrichment.</p> <h2>Columns in CSV files</h2> <p><em><strong>segment_id</strong></em> : The ID of the audio recording in AudioSet SL. These are in the form <em><YouTube ID>_<start time in ms>_<end time in ms></em></p> <p><em><strong>caption</strong></em> : The caption generated for each audio clip, corresponding to the specific subset (e.g., WavCaps, Scene-Focused, Emotion Addon, or Emotion Rewrite) as indicated by the file name.</p> <h2>Conditions of use</h2> <p>Dataset created by Mithun Manivannan, Vignesh Nethrapalli, Mark Cartwright</p> <p>The EmotionCaps dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:</p> <p><a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New Jersey Institute of Technology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the EmotionCaps dataset or any part of it.</p> <h2>Feedback</h2> <p>Please help us improve EmotionCaps by sending your feedback to:</p> <ul> <li>Mithun Manivannan: <a href="mailto:mithun.mani01@gmail.com">mithun.mani01@gmail.com</a></li> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <h2>Acknowledgments</h2> <p>This work was partially supported by the New Jersey Institute of Technology Honors Summer Research Institute (HSRI).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.