Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

135

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

135 results for “Multimodal dataset”

Learn how ShareScore rates datasets ↗
Figshare52/100

MAMEM Phase I Dataset - A dataset for multimodal human-computer interaction using biosignals and eye tracking information

<p>This dataset combines multimodal biosignals and eye tracking information gathered under a human-computer interaction framework. The dataset was developed in the vein of the MAMEM project that aims to endow people with motor disabilities with the ability to edit and author multimedia content through mental commands and gaze activity. The dataset includes EEG, eye-tracking, and physiological (GSR and Heart rate) signals along with demographic, clinical and behavioral data collected from 36 individuals (18 able-bodied and 18 motor-impaired). Data were collected during the interaction with specifically designed interface for web browsing and multimedia content manipulation and during imaginary movement tasks. Alongside these data we also include evaluation reports both from the subjects and the experimenters as far as the experimental procedure and collected dataset are concerned. We believe that the presented dataset will contribute towards the development and evaluation of modern human-computer interaction systems that would foster the integration of people with severe motor impairments back into society.</p>

opencc-by-4.0Dec 2016View details →
zenodo52/100

PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures

<h1>PAN-AR</h1> <p>This is <strong>PAN-AR</strong> (Panoramas, Ambient Noise &amp; Ambisonics RIRs), a dataset described in the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <blockquote> <p>Filippo Denti, Davide Fantini, Federico Avanzini and Giorgio Presti. PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures. In <em>Proceedings of the 19th International Audio Mostly Conference</em>, Milan, Italy, September 2024.</p> </blockquote> <p>The dataset includes Spatial Room Impulse Responses (SRIRs) in second-order Ambisonics format, ambient noise recordings, and spherical photos. These data have been captured in four environments with different configurations of the source and listener positions:</p> <ol> <li>Printer room</li> <li>Meeting room</li> <li>Classroom</li> <li>Underground parking area</li> </ol> <p>Panoramas and planimetries are provided in a temporary version. The final version with post-processed panoramas and complete planimetries will be available soon. An example of the final panoramas is provided for position A of the printer room, while an example of complete planimetry is provided for the printer and the meeting rooms.</p> <h2>SOFA</h2> <p>The SRIRs are also provided in SOFA format&nbsp;<a href="https://sofacoustics.org/data/database/pan-ar/" target="_blank" rel="noopener">here</a>.</p> <h2>How to cite</h2> <p>If you use the PAN-AR dataset, please cite the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <pre><code>@inproceedings{denti2024panar,</code><br><code> title = {{PAN-AR}: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures},</code><br><code> author = {Denti, Filippo and Fantini, Davide and Avanzini, Federico and Presti, Giorgio},</code><br><code> year = {2024},</code><br><code> month = {September},</code><br><code> booktitle = {Proceedings of the 19th International Audio Mostly Conference (AM '24)},</code><br><code> location = {Milan, Italy},</code><br><code> publisher = {ACM},</code><br><code> isbn = {979-8-4007-0968-5/24/09},</code><br><code> doi = {10.1145/3678299.3678332}</code><br><code>}</code></pre>

opencc-by-sa-4.0Dec 2024View details →
zenodo52/100

PE-HRI-temporal: A Multimodal Temporal Dataset in a robot mediated Collaborative Educational Setting

<p><em><strong>Please note that this dataset corresponds to the training data used in "Social robots as skilled ignorant peers for supporting learning "[7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint).&nbsp;</strong></em></p> <p>&nbsp;</p> <p>This data set consists of&nbsp;<strong>multi-modal temporal team behaviors as well as learning outcomes </strong>collected in the context of a robot mediated collaborative and constructivist learning activity called JUSThink [1,2]. The data set can be useful for those looking to explore evolution of log actions, speech behavior, affective states, and gaze patterns for students to model constructs such as engagement, motivation, collaboration, etc. in educational settings.&nbsp;</p> <p>In this data set, team level data is collected from 34 teams of two (68 children) where the children are&nbsp;aged between 9 and 12. There are two files:&nbsp;&nbsp;</p> <p><strong>PE-HRI_learning_and_performance.csv:</strong> This file consists of the <strong>team level&nbsp;performance and learning metrics</strong> which are defined below:&nbsp;</p> <ul> <li> <p><em>last_error:</em> This is the error of the last submitted solution. Note that if a team has found an optimal solution (error = 0) the game stops, therefore making last error = 0. This is a metric for performance in the task.&nbsp;</p> </li> <li> <p><em>T_LG_absolute:</em>&nbsp;It is a&nbsp;team-level&nbsp;learning outcome that&nbsp;we calculate by taking&nbsp;the average of the two individual absolute&nbsp;learning gains of the team members. The individual absolute&nbsp;gain is the difference between a participant&rsquo;s post-test and pre-test score, divided by the maximum score that can be achieved (10), which grasps how much the participant learned of all the knowledge available.</p> </li> <li> <p><em>T_LG_relative:</em>&nbsp;It is a&nbsp;team-level&nbsp;learning outcome that&nbsp;we calculate by taking&nbsp;the average of the two individual relative learning gains of the team members. The individual relative gain is the difference between a participant&rsquo;s post-test and pre-test score, divided by the difference between the maximum score that can be achieved and the pre-test score. This grasps how much the participant learned of the knowledge that he/she didn&rsquo;t possess before the activity.&nbsp;</p> </li> <li> <p><em>T_LG_joint_abs:&nbsp;</em>It is a team-level learning outcome defined as the difference between the&nbsp;number of questions that both of the team members answer correctly in the post-test and in the pre-test, which grasps the amount of knowledge acquired together by the team members during the activity</p> </li> </ul> <p><strong>PE-HRI_behavioral_timeseries_w_labels.csv:</strong> In this file, for each team, the interaction of around 20-25&nbsp;minutes&nbsp;is organized in windows of 10 seconds; hence, we have a total of 5048 windows of 10 seconds each. We report team level log actions, speech behavior, affective states, and gaze patterns for each window.&nbsp;More specifically, within each window, 26 features are generated in two ways:&nbsp;</p> <ol> <li>non-incremental</li> <li>incremental</li> </ol> <p>A non-incremental type would mean the value of a feature <em>in</em> that particular time window while an incremental type would mean the value of a feature <em>until</em> that particular time window. The incremental type is indicated by an "_inc" at the end of the feature name. Hence, in the end, within each window, we have 52 values:&nbsp;</p> <ul> <li> <p><em>T_add/(_inc):&nbsp;</em>The number of times a team added an edge on the map in that window/(until that window).</p> </li> <li> <p><em>T_remove/(_inc):&nbsp;</em>The number of times a team removed an edge from the map in that window/(until that window).</p> </li> <li> <p><em>T_ratio_add_rem/(_inc):&nbsp;</em>The ratio of addition of edges over deletion of edges by a team in that window/(until that window).</p> </li> <li> <p><em>T_action/(_inc):</em>&nbsp;The total number of actions taken by a team (add, delete, submit, presses on the screen)&nbsp;in that window/(until that window).</p> </li> <li> <p><em>T_hist/(_inc):&nbsp;</em>The number of times a team opened the sub-window with history of their previous solutions&nbsp;in that window/(until that window).</p> </li> <li> <p><em>T_help/(_inc):&nbsp;</em>The number of times a team opened the instructions manual in that window/(until that window). Please note that the robot initially gives all the instructions before the game-play while a video is played for demonstration of the functionality of the game.&nbsp;</p> </li> <li> <p><em>T1_T1_rem/(_inc):&nbsp;</em>The number of times either&nbsp;of the two members in the team followed the pattern consecutively: I add an edge, I then delete it&nbsp;in that window/(until that window).</p> </li> <li> <p><em>T1_T1_add/(_inc):&nbsp;</em>The number of times either&nbsp;of the two members in the team followed the pattern consecutively: I delete an edge, I add it back&nbsp;in that window/(until that window).</p> </li> <li> <p><em>T1_T2_rem/(_inc):&nbsp;</em>The number of times the members of the team&nbsp;followed the pattern consecutively: I add an edge, you then delete it&nbsp;in that window/(until that window).</p> </li> <li> <p><em>T1_T2_add/(_inc):&nbsp;</em>The number of times the members of the team&nbsp;followed the pattern consecutively: I delete an edge, you add it back&nbsp;in that window/(until that window).</p> </li> <li> <p><em>redundant_exist/(_inc):&nbsp;</em>The number of times the team had redundant edges in their map&nbsp;in that window/(until that window).</p> </li> <li> <p><em>positive_valence/(_inc):&nbsp;</em>The average value of positive valence for the team&nbsp;in that window/(until that window).</p> </li> <li> <p><em>negative_valence/(_inc):&nbsp;</em>The average value of negative valence for the team&nbsp;in that window/(until that window).</p> </li> <li> <p><em>difference_in_valence/(_inc):&nbsp;</em>The difference of the average value of positive and negative valence for the team&nbsp;in that window/(until that window).</p> </li> <li> <p><em>arousal/(_inc):&nbsp;</em>The average value of arousal for the team&nbsp;in that window/(until that window).</p> </li> <li> <p><em>gaze_at_partner/(_inc):&nbsp;</em>The average of the the two team member's gaze when looking at their partner&nbsp;in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>gaze_at_robot/(_inc):&nbsp;</em>The average of the the two team member's gaze when&nbsp;looking at the robot&nbsp;in that window/(until that window).&nbsp;Each individual member's gaze is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>gaze_other/(_inc):&nbsp;</em>The average of the the two team member's gaze when&nbsp;looking in the direction opposite to the robot&nbsp;in that window/(until that window).&nbsp;Each individual member's gaze is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>gaze_at_screen_left/(_inc):&nbsp;</em>The average of the the two team member's gaze when&nbsp;looking at the left side of the screen&nbsp;in that window/(until that window).&nbsp;Each individual member's gaze is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>gaze_at_screen_right/(_inc):</em>&nbsp;The average of the the two team member's gaze when looking at the right side of the screen&nbsp;in that window/(until that window).&nbsp;Each individual member's gaze is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>T_speech_activity/(_inc):&nbsp;</em>The average of the two team member's speech activity in that window/(until that window). Each individual member's speech activity is calculated as a percentage of time that they are speaking in that window/(until that window).&nbsp;</p> </li> <li> <p><em>T_silence/(_inc):&nbsp;</em>The average of the two team member's silence in that window/(until that window). Each individual member's silence is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>T_short_pauses/(_inc):&nbsp;</em>The average of the two team member's short pauses over their speech activity&nbsp;in that window/(until that window). Each individual member's short pause&nbsp;refers to a brief pause of 0.15 seconds and is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>T_long_pauses/(_inc):&nbsp;</em>The average of the two team members long pauses over their speech activity&nbsp;in that window/(until that window). Each individual member's long&nbsp;pause&nbsp;refers to a pause of 1.5&nbsp;seconds and is calculated as a percentage of time in that window/(until that window).&nbsp;</p> </li> <li> <p><em>T_overlap/(_inc):&nbsp;</em>The average percentage of time the speech of the team members overlaps in that window/(until that window).</p> </li> <li> <p><em>T_overlap_to_speech_ratio/(_inc):&nbsp;</em>The ratio of the speech overlap over the speech activity of the team&nbsp;in that window/(until that window).</p> </li> </ul> <p>Apart from these 52&nbsp;values, within each window, we also indicate:&nbsp;</p> <ul> <li><em>team: </em>The team to which the window belongs to.</li> <li><em>time_in_secs:</em> Time in seconds until that window.</li> <li><em>window: </em>The window number.</li> <li><em>normalized_time: </em>The time when this window occurred with respect to the total duration of the task for a particular team.&nbsp;</li> <li>cluster_labels: The cluster number associated with each time window in reference to the productive and non-productive clusters found in [3]</li> <li>PE_score: The Productive Engagement score in each window</li> </ul> <p>Lastly, we briefly elaborate on how the features&nbsp;are operationalised. We extract log behaviors from the recorded rosbags while the behaviors related to both gaze and affective states are computed through the open source library OpenFace [6] that returns both facial actions units (AUs) as well as gaze angles.&nbsp;For voice activity detection (VAD), that classifies if a piece of audio is voiced or unvoiced, we made use of the python wrapper for the open source Google WebRTC VAD. The literature that inspired our&nbsp;log, audio and video features as well as the tools used to extract them are&nbsp;described in more detail in [3,4]. However, in those papers, we make use of only the aggregate version of this&nbsp;data [5].</p> <p><em><strong>Please note that this dataset corresponds to the training data used in [7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint).&nbsp;</strong></em></p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)

<p>Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine. However, most publicly available datasets on human body movements cannot be used to study both problems in an out-of-the-lab movement acquisition setting. The objective of the VIDIMU dataset is to pave the way towards affordable patient tracking solutions for remote daily life activities recognition and kinematic analysis.&nbsp;</p> <p>The VIDIMU dataset includes 54 healthy young adults that were recorded on video and 16 of them were simultaneously recorded using custom IMUs.&nbsp; For each subject, 13 activities were registered using a low-resolution video camera and five Inertial Measurement Units (IMUs). Inertial sensors were placed in the lower or the upper limbs of the subject, respectively for activities that involve movement with the lower or the upper body. Video recordings were postprocessed using the state-of-the-art pose estimator <em>BodyTrack</em> (similar to OpenPose, and&nbsp;included in NVIDIA Maxine-AR-SDK) to provide a sequence of 3D joint positions for each movement. Raw IMU recordings were post-processed to compute joint angles by inverse kinematics with <em>OpenSim</em>. For recordings including simultaneous acquisition of video and IMU data types, these signals were used for data file synchronization. Collected data can be further used in applications related to human activity recognition and biomechanics related experiments in simulated home-like settings.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Synthetic Multimodal Dataset for Daily Life Activities

<p><strong>Outline</strong></p> <ul> <li>This dataset is originally created for the&nbsp;<a href="https://challenge.knowledge-graph.jp/2022/">Knowledge Graph Reasoning Challenge for Social Issue</a>s (KGRC4SI)</li> <li>Video data that simulates daily life actions in a virtual space from Scenario Data.</li> <li>Knowledge graphs, and transcriptions of the Video Data content (&quot;who&quot; did what &quot;action&quot; with what &quot;object,&quot; when and where, and the resulting &quot;state&quot; or &quot;position&quot; of the object).</li> <li>Knowledge Graph Embedding Data are created for reasoning based on machine learning&nbsp;</li> <li>This data&nbsp;is open to the public as open data</li> </ul> <p><strong>Details</strong></p> <ul> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Movie">Videos</a></p> <ul> <li>mp4 format</li> <li>203&nbsp;action scenarios</li> <li>For each scenario, there is a character rear view (file name ending in 0), an indoor camera switching view (file name ending in 1), and a fixed camera view placed in each corner of the room (file name ending in 2-5). Also, for each action scenario, data was generated for a minimum of 1 to a maximum of 7 patterns with different room layouts (scenes). A total of 1,218&nbsp;videos</li> <li>Videos with slowly moving characters simulate the movements of elderly people.</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/RDF">Knowledge Graphs</a></p> <ul> <li>RDF format</li> <li>203&nbsp;knowledge graphs corresponding to the videos</li> <li>Includes schema and location supplement information</li> <li>The schema is described below</li> <li><a href="http://kgrc4si.ml:7200/sparql">SPARQL endpoints</a>&nbsp;and&nbsp;<a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/tree/kgrc4si#%E3%83%8A%E3%83%AC%E3%83%83%E3%82%B8%E3%82%B0%E3%83%A9%E3%83%95%E3%81%AE%E4%BD%BF%E7%94%A8%E6%96%B9%E6%B3%95">query examples</a>&nbsp;are available</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Program">Script Data</a></p> <ul> <li>txt format</li> <li>Data provided to VirtualHome2KG to generate videos and knowledge graphs</li> <li>Includes the action title and a brief description in text format.</li> </ul> </li> <li>Embedding <ul> <li>Embedding Vectors in TransE, ComplEx, and RotatE. Created with DGL-KE (<a href="https://dglke.dgl.ai/doc/">https://dglke.dgl.ai/doc/</a>)</li> <li>Embedding Vectors created with jRDF2vec (<a href="https://github.com/dwslab/jRDF2Vec">https://github.com/dwslab/jRDF2Vec</a>).</li> </ul> </li> </ul> <p><strong>Specification of Ontology</strong></p> <ul> <li>Please refer to the&nbsp;specification for descriptions of all classes, instances, and properties:&nbsp;<a href="https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.html">https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.htm</a></li> </ul> <p><strong>Related Resources</strong></p> <ul> <li><a href="https://www.youtube.com/watch?v=Ajbn8hNXiZ8&amp;list=PLHaRK-B0LUwjvrPgmIBTrf3DsPhmdnFTW">KGRC4SI Final Presentations with automatic English subtitles (YouTube)</a></li> <li><a href="https://github.com/aistairc/VirtualHome2KG">VirtualHome2KG (Software)</a></li> <li><a href="https://github.com/aistairc/virtualhome_unity_aist">VirtualHome-AIST (Unity</a>)</li> <li><a href="https://github.com/aistairc/virtualhome_aist">VirtualHome-AIST (Python API</a>)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_visualization">Visualization Tool</a>&nbsp;(Software)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_generation">Script Editor</a>&nbsp;(Software)</li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Raw dataset for: High-throughput multimodal wide-field Fourier-transform Raman microscope

<p>This is the Raw spectral dataset of the data published in&nbsp;10.1364/OPTICA.488860</p> <p>Data are arranged as follows:</p> <p>wavenumber [Nx1]</p> <p>Hyperspectrum_cube [Nx2, A, B]: hyperspectral datacube, where:&nbsp;Hyperspectrum_cube (1:N, :, :) is the real part;&nbsp;Hyperspectrum_cube (N+1:2N, :, :) is the imaginary part</p> <p>maximum [1x1]</p> <p>minimum [1x1].</p> <p>N: number of spectral bands</p> <p>A and&nbsp;B:&nbsp;size of the spatial coordinates</p> <p>Spectral amplitudes are obtained by:&nbsp;Hyperspectrum_cube=double(Hyperspectrum_cube)./(2.^16-1).*maximum+minimum</p>

opencc-by-4.0May 2023View details →
zenodo44/100

JDC2015 - A multimodal dataset from facilitating multi-tabletop lessons in an open-doors day

<p><strong>IMPORTANT NOTE: Two of the files in this dataset are incorrect, see this dataset's errata at https://zenodo.org/record/204063 and https://zenodo.org/record/204819</strong></p> <p>This dataset contains eye-tracking, EEG, accelerometer, indoor location and video coding data from a single subject (a researcher with limited teacher experience), facilitating four maths lessons in a simulated multi-tabletop classroom,  with four cohorts of 10-12 year old students, using tangible paper tabletops and a projector. These sessions were recorded in the frame of the MIOCTI project (http://chili.epfl.ch/miocti).</p> <p>This dataset has been used in several scientific works, such a submitted journal paper "Orchestration Load Indicators and Patterns: In-the-wild Studies Using Mobile Eye-tracking", by Luis P. Prieto, Kshitij Sharma, Lukasz Kidzinski &amp; Pierre Dillenbourg (the analysis and usage of this dataset is available publicly at https://github.com/chili-epfl/paper-IEEETLT-orchestrationload)</p>

opencc-by-sa-4.0Dec 2016View details →
zenodo44/100

Dataset of the scientific paper " Multimodal robotic system for upper-limb rehabilitation in physical environment" (Advances in Mechanical Engineering)

<p>There are eight files with the following information:<br>     - pos_stateXX.bin, binary file with information of the end effector position of the robot device in meters along the three axis (X, Y, Z) during state XX of the experiment<br>     - target_stateXX.bin, binary file with information of the target position for the robot device in meters along the three axis (X, Y, Z) during state XX of the experiment<br>     - emg_channelXX.bin, binary file with information of channel 1 of the EMG sensor in mV during during the whole time of the experiment<br>     - color_stateXX.bin, binary file with information of color filter information during state XX of the experiment. This information is the percentage of pixels with the correct color (yellow, cyan or magenta) inside the region of interest</p> <p> </p>

opencc-by-4.0Aug 2016View details →
zenodo44/100

MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications

<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>

openNov 2023View details →
zenodo44/100

N20EMv2 dataset for automatic music transcription from multimodal singing

<p>N20EMv2 dataset for multimodal automatic music transcription from multimodal singing, presented in our TOMM 2024 paper, Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing. This dataset contains recordings of two modalities: audio and video.&nbsp;</p> <p>Our paper is available at: https://dl.acm.org/doi/10.1145/3651310.</p> <p>Code is available at: https://github.com/guxm2021/SVT_SpeechBrain</p> <p>Please cite our work as:</p> <pre>@article{gu2024automatic, title={Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing}, author={Gu, Xiangming and Ou, Longshen and Zeng, Wei and Zhang, Jianan and Wong, Nicholas and Wang, Ye}, journal={ACM Transactions on Multimedia Computing, Communications and Applications}, publisher={ACM New York, NY}, year={2024} }</pre>

opencc-by-sa-4.0Mar 2024View details →
zenodo44/100

ROCOv2: Radiology Objects in COntext Version 2, An Updated Multimodal Image Dataset

<p>Recent advances in deep learning techniques have enabled the development of systems for automatic analysis of medical images. These systems often require large amounts of training data with high quality labels, which is difficult and time consuming to generate.</p> <p>Here, we introduce Radiology Object in COntext Version 2 (ROCOv2), a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PubMed Open Access subset. Concepts for clinical modality, anatomy (X-ray), and directionality (X-ray) were manually curated and additionally evaluated by a radiologist. Unlike MIMIC-CXR, ROCOv2 includes seven different clinical modalities.</p> <p>It is an updated version of the ROCO dataset published in 2018, and includes 35,705 new images added to PubMed since 2018, as well as manually curated medical concepts for modality, body region (X-ray) and directionality (X-ray). The dataset consists of 79,789 images and has been used, with minor modifications, in the concept detection and caption prediction tasks of ImageCLEFmedical 2023. The participants had access to the training and validation sets after signing a user agreement.</p> <p>The dataset is suitable for training image annotation models based on image-caption pairs, or for multi-label image classification using the UMLS concepts provided with each image, e.g., to build systems to support structured medical reporting.</p> <p>Additional possible use cases for the ROCOv2 dataset include the pre-training of models for the medical domain, and the evaluation evaluation of deep learning models for multi-task learning.</p>

opencc-by-nc-4.0Nov 2023View details →
zenodo44/100

scGeneAI pbmc multimodal dataset

<p>The input pbmc_multimodal_2023&nbsp;dataset used in the full-size examples in scGenAI is uploaded here</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

CREATTIVE3D multimodal dataset of user behavior in virtual reality

<p>In the context of the <a href="https://project.inria.fr/creattive3d/">ANR CREATTIVE3D</a> project, we join the expertise of computer science, neuroscience, and clinical practitioners, with the aim to analyze the impact that a simulated low-vision condition has on user navigation behavior in complex road crossing scenes: a common daily situation where the difficulty to access and process visual information (e.g., traffic lights, approaching cars) in a timely fashion can lead to serious consequences on a person's safety and well-being. As a secondary objective, we also aim to investigate the potential role virtual reality could play in rehabilitation and training protocols for low-vision patients.</p> <p>This dataset contains the data as part of the study described in <a href="https://hal.science/hal-04102737">An Integrated Framework for Understanding Multimodal Embodied Experiences in Interactive Virtual Reality</a>.</p> <p>The dataset is metadata for the pre-print <a href="https://inria.hal.science/hal-04429351">Exploring, walking, and interacting in virtual reality with simulated low vision: a living contextual dataset</a></p> <p>To use this dataset, please cite:</p> <blockquote> <pre>@unpublished{wu:hal-04429351, TITLE = {{Exploring, walking, and interacting in virtual reality with simulated low vision: a living contextual dataset}}, AUTHOR = {Wu, Hui-Yin and Robert, Florent Alain Sauveur and Gallo, Franz Franco and <br> Pirkovets, Kateryna and Quere, Cl{\'e}ment and Delachambre, Johanna and <br> Ramano{\"e}l, Stephen and Gros, Auriane and Winckler, Marco and Sassatelli, Lucile and <br> Hayotte, Meggy and Menin, Aline and Kornprobst, Pierre}, URL = {https://inria.hal.science/hal-04429351}, NOTE = {working paper or preprint}, YEAR = {2023}, MONTH = Dec, KEYWORDS = {Virtual reality ; Dataset ; Context ; Low vision ; 3D environments ; User study}, PDF = {https://inria.hal.science/hal-04429351/file/2023_CREATTIVE3D_dataset_arxiv_.pdf}, HAL_ID = {hal-04429351}, HAL_VERSION = {v1}, }<br><br>@inproceedings{robert2023integrated, title={An integrated framework for understanding multimodal embodied experiences in interactive virtual reality}, author={Robert, Florent and Wu, Hui-Yin and Sassatelli, Lucile and Ramanoel, Stephen and <br> Gros, Auriane and Winckler, Marco}, booktitle={Proceedings of the 2023 ACM International Conference on Interactive Media Experiences}, pages={14--26}, year={2023} }</pre> </blockquote> <h3>&nbsp;</h3> <h3>Versions</h3> <p>2024-12-18: Updated readme with description of labels, columns, and suggestions on how to start exploring the dataset. We also provide the questionnaire responses and observation notes in English (questionnaire_translation_EN.csv).</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Synthetic Multimodal Drone Delivery Dataset

<h1><strong>README: Synthetic Logistics Dataset Structure and Components</strong></h1> <p>This dataset provides a structured representation of logistics data designed to evaluate and optimize hybrid truck-and-drone delivery networks. It captures a comprehensive set of parameters essential for modeling real-world logistics scenarios, including spatial coordinates, environmental conditions, and operational constraints. The data is meticulously organized into distinct keys, each representing a critical aspect of the delivery network, enabling researchers and practitioners to conduct flexible and in-depth analyses. &nbsp;</p> <p>The dataset is a curated subset derived from the research presented in the paper "Synthetic Dataset Generation for Optimizing Multimodal Drone Delivery Systems" by Altinsel et al. (2024), published in Drones. It serves as a practical resource for studying the interplay between ground-based and aerial delivery systems, with a focus on efficiency, environmental impact, and operational feasibility. &nbsp;</p> <p><strong>Altinses, D., Torres, D. O. S., Gobachew, A. M., Lier, S., &amp; Schwung, A. (2024). Synthetic Dataset Generation for Optimizing Multimodal Drone Delivery Systems.&nbsp;<em>Drones (2504-446X)</em>,&nbsp;<em>8</em>(12).</strong></p> <p>Each data file contains information on ten customer locations, specified by their x and y coordinates, which facilitate the modeling of delivery routes and service areas. Additionally, the dataset includes communication data represented as a two-dimensional grid, which can be used to assess signal strength, connectivity, or other network-related factors that influence drone operations. &nbsp;</p> <p>A key feature of this dataset is the inclusion of wind data, structured as a two-dimensional grid with four distinct features per grid point. These features likely represent wind velocity components (such as horizontal and vertical directions) along with auxiliary parameters like turbulence intensity or wind shear, which are crucial for drone path planning and energy consumption estimation. The wind data enables researchers to simulate realistic environmental conditions and evaluate their impact on drone performance, stability, and battery life. &nbsp;</p> <p>By integrating geospatial, environmental, and operational data, this dataset supports a wide range of applications, from route optimization and energy efficiency studies to risk assessment and resilience planning in multimodal delivery systems. Its synthetic nature ensures reproducibility while maintaining relevance to real-world logistics challenges, making it a valuable tool for advancing research in drone-assisted delivery networks. &nbsp;</p> <p>&nbsp;</p> <h3><strong>The 4 wind channels represent:</strong></h3> <ol> <li> <p><strong><code>X</code>&nbsp;and&nbsp;<code>Y</code>&nbsp;(Grid Positions)</strong></p> <ul> <li> <p>These define&nbsp;<strong>where the arrows start</strong>&nbsp;(usually a meshgrid).</p> </li> </ul> </li> <li> <p><strong><code>U</code>&nbsp;and&nbsp;<code>V</code>&nbsp;(Arrow Directions)</strong></p> <ul> <li> <p><code>U</code>&nbsp;= Horizontal component (e.g., gradient in&nbsp;<code>x</code>).</p> </li> <li> <p><code>V</code>&nbsp;= Vertical component (e.g., gradient in&nbsp;<code>y</code>).</p> </li> </ul> </li> </ol> <p>&nbsp;</p> <h3><strong>How to load the files using Python:</strong></h3> <p>data = np.loadtxt('data.txt')</p> <p>#### Just for Wind data:</p> <p>data = data.reshape((4,16,16))&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

SILKNOW Multimodal Cultural Heritage Dataset

<p>SILKNOW Multimodal Cultural Heritage Dataset. Includes text descriptions, images, labels, and predictions made by individual modality classifiers.</p> <p>The data resulted from an export of the SILKNOW&nbsp;Knowledge Graph. See:&nbsp;<a href="https://zenodo.org/record/5743090">https://zenodo.org/record/5743090</a></p> <p>Repository with code using this dataset&nbsp;available at:&nbsp;<a href="https://github.com/silknow/multimodal_cultural_heritage">https://github.com/silknow/multimodal_cultural_heritage</a></p>

opencc-by-4.0May 2022View details →
zenodo44/100

Multimodal Dataset of 3D point clouds and CT-volumes

<p>The multimodal dataset for evaluating algorithms for aligning CT volumes and point clouds which is presented in &#39;Multimodal registration across 3D point clouds and CT-volumes&#39;. (Saiti, E., and T. Theoharis. &quot;Multimodal registration across 3D point clouds and CT-volumes.&quot;&nbsp;<em>Computers &amp; Graphics</em>&nbsp;106 (2022): 259-266.) The&nbsp;multimodal dataset consistsof real micro-CT scans and their synthetically generated 3D models (point clouds) .</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

MSMD - Multimodal Sheet Music Dataset

<p>MSMD is a synthetic dataset of 497 pieces of (classical) music that contains both audio and score representations of the pieces aligned at a fine-grained level (344,742 pairs of noteheads aligned to their audio/MIDI counterpart). It can be used for training and evaluating multimodal models that enable crossing from one modality to the other, such as retrieving sheet music using recordings or following a performance in the score image.</p> <p>Please find further information and a corresponding Python package on this Github page: <a href="https://github.com/CPJKU/msmd">https://github.com/CPJKU/msmd</a></p> <p>If you use this dataset, please cite:<br> [1] Matthias Dorfer, Jan Hajič jr., Andreas Arzt, Harald Frostel, Gerhard Widmer.<br> <a href="https://transactions.ismir.net/articles/10.5334/tismir.12/">Learning Audio-Sheet Music Correspondences for Cross-Modal Retrieval and Piece Identification</a> (<a href="https://transactions.ismir.net/articles/10.5334/tismir.12/galley/8/download/">PDF</a>).<br> Transactions of the International Society for Music Information Retrieval, issue 1, 2018.</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Dataset of the study: Explaining recovery from coma with multimodal neuroimaging

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo44/100

OpenMapCD: A Multimodal Benchmark Dataset for Change Detection Between Optical Remote Sensing and Map Data

<p><strong>Overview:&nbsp;</strong></p> <ol> <li>OpenMapCD, the&nbsp;<strong>first large-scale multimodal dataset</strong>&nbsp;for change detection on optical remote sensing imagery and map (OpenStreetMap) data,&nbsp;<strong>supporing basic binary change detection and further semantic change detection</strong></li> <li>OpenMapCD is highly geographically diverse, with&nbsp;<strong>1288</strong>&nbsp;benchmark samples with 1024x1024 pixels from&nbsp;<strong>40&nbsp;</strong>regions across six continents and out-of-distribution data in two areas in Japan</li> <li>Advancing land-cover mapping, binary change detection and semantic change detection tasks, and GIS system updating<br><br></li> </ol> <p><strong>Research Paper:&nbsp;<br></strong></p> <ul> <li>Arxiv paper:&nbsp;<a href="https://arxiv.org/abs/2310.02674v3">https://arxiv.org/html/2310.02674v3</a></li> <li>TGRS paper:&nbsp;<a href="https://ieeexplore.ieee.org/document/10551264">https://ieeexplore.ieee.org/document/10551264</a></li> </ul> <p><strong><br>Project Page:</strong><br>The benchmark code is available at: <a href="https://github.com/ChenHongruixuan/ObjFormer">https://github.com/ChenHongruixuan/ObjFormer</a><br><br><strong>Reference:</strong></p> <pre><code>@ARTICLE{Chen2024ObjFormer, author={Chen, Hongruixuan and Lan, Cuiling and Song, Jian and Broni-Bediako, Clifford and Xia, Junshi and Yokoya, Naoto}, journal={IEEE Transactions on Geoscience and Remote Sensing}, title={ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer}, year={2024}, volume={62}, number={}, pages={1-22}, doi={10.1109/TGRS.2024.3410389} }</code></pre>

opencc-by-4.0Jun 2024View details →
zenodo44/100

MultiSubs: A Large-scale Multimodal and Multilingual Dataset

<p>MultiSubs&nbsp;is a dataset of multilingual subtitles gathered from&nbsp;<a href="https://opus.nlpl.eu/OpenSubtitles.php">the OPUS OpenSubtitles dataset</a>,&nbsp;which in turn was sourced from <a href="http://www.opensubtitles.org/">opensubtitles.org</a>. We have supplemented some text fragments (visually salient nouns in this release) within the subtitles with web images, where the word sense of the fragment has been disambiguated using a cross-lingual approach.&nbsp;</p> <p>Please refer to our&nbsp;paper for a more detailed description of the dataset:</p> <p>Josiah Wang, Pranava Madhyastha, Josiel Figueiredo, Chiraag Lala, Lucia Specia (2021). <a href="https://arxiv.org/abs/2103.01910">MultiSubs: A Large-scale Multimodal and Multilingual Dataset</a>. CoRR, abs/2103.01910. Available at: <a href="https://arxiv.org/abs/2103.01910">https://arxiv.org/abs/2103.01910</a></p>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record