Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “Captioning”
CUB-200 Speech captions (Part-II)
<p>Speech captions of CUB-200 database. These speech captions are synthesized by Tacotron-v2 according to the original textual descriptions.</p>
EmotionCaps: A Synthetic Emotion-Enriched Audio Captioning Dataset
<p>Version 1.0, October 2024</p> <h2>Created by</h2> <p>Mithun Manivannan (1), Vignesh Nethrapalli (1), Mark Cartwright (1)</p> <ol> <li>Sound Interaction and Computer Lab, New Jersey Institute of Technology</li> </ol> <h2>Publication</h2> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Manivannan, M., Nethrapalli, V., Cartwright, M. EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation. arXiv preprint arXiv:2410.12028, 2024.</p> <h2>Description</h2> <p>EmotionCaps is a ChatGPT-assisted, weakly-labeled audio captioning dataset developed to bridge the gap between soundscape emotion recognition (SER) and automated audio captioning (AAC). Created through a three-stage pipeline, the dataset leverages ground-truth annotations from AudioSet SL, which are enhanced by ChatGPT using tailored prompts and emotions assigned via a soundscape emotion recognition model trained on Emo-Soundscapes Dataset. It comprises four subsets of captions for 120,071 audio clips, each reflecting a different prompt variation: WavCaps-like, Scene-Focused, Emotion Addon, and Emotion Rewrite. The average word counts for these subsets are: WavCaps-like (12.61), Scene-Focused (14.04), Emotion Addon (18.35), and Emotion Rewrite (18.65). The increase in word count for the emotion prompts illustrates the difference in sentence length when integrating emotion information into the captions.</p> <h2>Audio Data</h2> <p>The audio data is from AudioSet SL, the strongly-labled subset of 120,071 audio clips from the larger AudioSet dataset.</p> <h2>Synthetic Captions</h2> <p>The synthetic captions were generated using a three-stage pipeline, beginning with training a soundscape emotion recognition model. This model assesses the valence and arousal of each audio clip, mapping the resulting vector to an emotion identifier. Next, we leveraged the ground-truth annotations from AudioSet SL, and extracted the list of sound events. Using these sound events, we employed ChatGPT to create different variations of captions by applying distinct prompts.</p> <p>We first used the WavCaps prompt for AudioSet SL as a base, the output of which we call WavCaps-like. Building on this, we created three new prompt variations (1) <strong>scene-focused</strong> which is a modified WavCaps prompt that describes the scene, (2) <strong>emotion addon</strong> which is an extension of the scene-Focused prompt, where an emotion is appended to the list of sound events to guide the caption generation, and (3) <strong>emotion rewrite</strong> which consists of two-step prompt where ChatGPT first generates the scene-focused caption, then is instructed to rewrite it with a specific emotion in mind.</p> <p>Using these four prompt styles — WavCaps, Scene-Focused, Emotion Addon, and Emotion Rewrite — along with the AudioSet SL sound events and predicted emotions, we employed ChatGPT-3.5 Turbo to generate four corresponding caption variations for the dataset.</p> <p>Each caption variation has been organized into separate CSV files for clarity and accessibility. All files correspond to the same set of audio clips from AudioSet SL, with the key distinction being the caption variation associated with each clip. The different subsets are designed to be used independently, as they each fulfill specific roles in understanding the impact of emotion in audio captions.</p> <ul> <li> <p><strong>wavcaps-like.csv</strong>: Contains captions generated using the WavCaps prompt, serving as the baseline before emotion is introduced.</p> </li> <li> <p><strong>scene-focused.csv</strong>: Provides captions focused on describing the scene or environment of the audio clip, without emotion integration.</p> </li> <li> <p><strong>emotion-addon.csv</strong>: Captions where emotion data is appended to the scene-focused base caption.</p> </li> <li> <p><strong>emotion-rewrite.csv</strong>: Captions that are completely rewritten based on the scene-focused base caption and the assigned emotion.</p> </li> </ul> <p>This structure allows users to explore how emotional content influences captioning models by comparing the variations both with and without emotional enrichment.</p> <h2>Columns in CSV files</h2> <p><em><strong>segment_id</strong></em> : The ID of the audio recording in AudioSet SL. These are in the form <em><YouTube ID>_<start time in ms>_<end time in ms></em></p> <p><em><strong>caption</strong></em> : The caption generated for each audio clip, corresponding to the specific subset (e.g., WavCaps, Scene-Focused, Emotion Addon, or Emotion Rewrite) as indicated by the file name.</p> <h2>Conditions of use</h2> <p>Dataset created by Mithun Manivannan, Vignesh Nethrapalli, Mark Cartwright</p> <p>The EmotionCaps dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:</p> <p><a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New Jersey Institute of Technology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the EmotionCaps dataset or any part of it.</p> <h2>Feedback</h2> <p>Please help us improve EmotionCaps by sending your feedback to:</p> <ul> <li>Mithun Manivannan: <a href="mailto:mithun.mani01@gmail.com">mithun.mani01@gmail.com</a></li> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <h2>Acknowledgments</h2> <p>This work was partially supported by the New Jersey Institute of Technology Honors Summer Research Institute (HSRI).</p>
FaceAttDB: A Multilingual Dataset for Facial Attribute Captioning
<p>The FaceCaption dataset is a curated collection specifically created for the purpose of research in the field of facial attribute captioning. It consists of 2,000 portrait images sourced from the CelebA dataset, showcasing a diverse range of facial characteristics such as age, gender, expression, and hair color. The dataset includes five captions per image, providing both English and Google-translated Bangla versions.</p> <p>The dataset is designed to facilitate the exploration of multilingual caption generation on portrait images. Each image in the dataset is accompanied by descriptive and informative captions that accurately describe the visual characteristics present in the image. The captions were generated based on the attribute annotations available in the CelebA dataset, ensuring a close alignment between the captions and the visual attributes.</p> <p>The images in the BanglaFaceCaption dataset are conveniently stored in a single folder, making them easily accessible for training and evaluation purposes. Additionally, an accompanying Excel sheet is provided, linking each image file with its corresponding English and Bangla captions.</p> <p>While the current version of the dataset comprises 2,000 images with five captions each, future work aims to expand the dataset size to enhance the diversity and robustness of models trained on it. The BanglaFaceCaption dataset serves as a valuable resource for researchers and practitioners interested in advancing the field of facial attribute captioning and exploring multilingual caption generation capabilities.</p>
CAPTION AI to Minimize Risk of COVID Exposure
ClinicalTrials.gov study NCT04336774. IPD Sharing: NO. Countries: 1. Publications: 0.
Oxford-102 Speech Captions
<p>Speech captions (synthesized by Tacotron2) of Oxford-102 using their original textual captions.</p>
Fig. 1 Braula coeca female reproductive tract, historic illustrations with original captions. a in Setting the records straight II: "single spermatheca" of Braula coeca (Diptera: Braulidae) is really the ventral receptacle
Fig. 1 Braula coeca female reproductive tract, historic illustrations with original captions. a Skaife (1922). b Alfonsus and Brown (1931). agl accessory gland, cod common ovarian duct, ot ovarian tubes, ov ovipositor, ovd oviduct, ovt ovarian tube, sp spermatheca, v vagina. Skaife, S. H. (1922). On Braula caeca, Nitzsch, a dipterous parasite of the honey bee. Transactions of the Royal Society of South Africa, 10(1), 41–48. Alfonsus, E. C., & Braun, E. (1931). Preliminary studies of the internal structures of Braula coeca Nitzsch. Annals of the Entomological Society of America, 24(3), 561–582
Figure captions
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.