Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,389
datasets available to search
ShareScore release 0.9.0
Dataset results
1,389 results for “Multimodal”
Disentangling the origins of confidence in speeded perceptual judgments through multimodal imaging
Open the record for dataset details and reuse information.
MAMEM Phase I Dataset - A dataset for multimodal human-computer interaction using biosignals and eye tracking information
<p>This dataset combines multimodal biosignals and eye tracking information gathered under a human-computer interaction framework. The dataset was developed in the vein of the MAMEM project that aims to endow people with motor disabilities with the ability to edit and author multimedia content through mental commands and gaze activity. The dataset includes EEG, eye-tracking, and physiological (GSR and Heart rate) signals along with demographic, clinical and behavioral data collected from 36 individuals (18 able-bodied and 18 motor-impaired). Data were collected during the interaction with specifically designed interface for web browsing and multimedia content manipulation and during imaginary movement tasks. Alongside these data we also include evaluation reports both from the subjects and the experimenters as far as the experimental procedure and collected dataset are concerned. We believe that the presented dataset will contribute towards the development and evaluation of modern human-computer interaction systems that would foster the integration of people with severe motor impairments back into society.</p>
PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures
<h1>PAN-AR</h1> <p>This is <strong>PAN-AR</strong> (Panoramas, Ambient Noise & Ambisonics RIRs), a dataset described in the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <blockquote> <p>Filippo Denti, Davide Fantini, Federico Avanzini and Giorgio Presti. PAN-AR: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures. In <em>Proceedings of the 19th International Audio Mostly Conference</em>, Milan, Italy, September 2024.</p> </blockquote> <p>The dataset includes Spatial Room Impulse Responses (SRIRs) in second-order Ambisonics format, ambient noise recordings, and spherical photos. These data have been captured in four environments with different configurations of the source and listener positions:</p> <ol> <li>Printer room</li> <li>Meeting room</li> <li>Classroom</li> <li>Underground parking area</li> </ol> <p>Panoramas and planimetries are provided in a temporary version. The final version with post-processed panoramas and complete planimetries will be available soon. An example of the final panoramas is provided for position A of the printer room, while an example of complete planimetry is provided for the printer and the meeting rooms.</p> <h2>SOFA</h2> <p>The SRIRs are also provided in SOFA format <a href="https://sofacoustics.org/data/database/pan-ar/" target="_blank" rel="noopener">here</a>.</p> <h2>How to cite</h2> <p>If you use the PAN-AR dataset, please cite the following <a href="https://doi.org/10.1145/3678299.3678332" target="_blank" rel="noopener">paper</a>:</p> <pre><code>@inproceedings{denti2024panar,</code><br><code> title = {{PAN-AR}: A Multimodal Dataset of Higher-Order Ambisonics Room Impulse Responses, Ambient Noise and Spherical Pictures},</code><br><code> author = {Denti, Filippo and Fantini, Davide and Avanzini, Federico and Presti, Giorgio},</code><br><code> year = {2024},</code><br><code> month = {September},</code><br><code> booktitle = {Proceedings of the 19th International Audio Mostly Conference (AM '24)},</code><br><code> location = {Milan, Italy},</code><br><code> publisher = {ACM},</code><br><code> isbn = {979-8-4007-0968-5/24/09},</code><br><code> doi = {10.1145/3678299.3678332}</code><br><code>}</code></pre>
PE-HRI-temporal: A Multimodal Temporal Dataset in a robot mediated Collaborative Educational Setting
<p><em><strong>Please note that this dataset corresponds to the training data used in "Social robots as skilled ignorant peers for supporting learning "[7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint). </strong></em></p> <p> </p> <p>This data set consists of <strong>multi-modal temporal team behaviors as well as learning outcomes </strong>collected in the context of a robot mediated collaborative and constructivist learning activity called JUSThink [1,2]. The data set can be useful for those looking to explore evolution of log actions, speech behavior, affective states, and gaze patterns for students to model constructs such as engagement, motivation, collaboration, etc. in educational settings. </p> <p>In this data set, team level data is collected from 34 teams of two (68 children) where the children are aged between 9 and 12. There are two files: </p> <p><strong>PE-HRI_learning_and_performance.csv:</strong> This file consists of the <strong>team level performance and learning metrics</strong> which are defined below: </p> <ul> <li> <p><em>last_error:</em> This is the error of the last submitted solution. Note that if a team has found an optimal solution (error = 0) the game stops, therefore making last error = 0. This is a metric for performance in the task. </p> </li> <li> <p><em>T_LG_absolute:</em> It is a team-level learning outcome that we calculate by taking the average of the two individual absolute learning gains of the team members. The individual absolute gain is the difference between a participant’s post-test and pre-test score, divided by the maximum score that can be achieved (10), which grasps how much the participant learned of all the knowledge available.</p> </li> <li> <p><em>T_LG_relative:</em> It is a team-level learning outcome that we calculate by taking the average of the two individual relative learning gains of the team members. The individual relative gain is the difference between a participant’s post-test and pre-test score, divided by the difference between the maximum score that can be achieved and the pre-test score. This grasps how much the participant learned of the knowledge that he/she didn’t possess before the activity. </p> </li> <li> <p><em>T_LG_joint_abs: </em>It is a team-level learning outcome defined as the difference between the number of questions that both of the team members answer correctly in the post-test and in the pre-test, which grasps the amount of knowledge acquired together by the team members during the activity</p> </li> </ul> <p><strong>PE-HRI_behavioral_timeseries_w_labels.csv:</strong> In this file, for each team, the interaction of around 20-25 minutes is organized in windows of 10 seconds; hence, we have a total of 5048 windows of 10 seconds each. We report team level log actions, speech behavior, affective states, and gaze patterns for each window. More specifically, within each window, 26 features are generated in two ways: </p> <ol> <li>non-incremental</li> <li>incremental</li> </ol> <p>A non-incremental type would mean the value of a feature <em>in</em> that particular time window while an incremental type would mean the value of a feature <em>until</em> that particular time window. The incremental type is indicated by an "_inc" at the end of the feature name. Hence, in the end, within each window, we have 52 values: </p> <ul> <li> <p><em>T_add/(_inc): </em>The number of times a team added an edge on the map in that window/(until that window).</p> </li> <li> <p><em>T_remove/(_inc): </em>The number of times a team removed an edge from the map in that window/(until that window).</p> </li> <li> <p><em>T_ratio_add_rem/(_inc): </em>The ratio of addition of edges over deletion of edges by a team in that window/(until that window).</p> </li> <li> <p><em>T_action/(_inc):</em> The total number of actions taken by a team (add, delete, submit, presses on the screen) in that window/(until that window).</p> </li> <li> <p><em>T_hist/(_inc): </em>The number of times a team opened the sub-window with history of their previous solutions in that window/(until that window).</p> </li> <li> <p><em>T_help/(_inc): </em>The number of times a team opened the instructions manual in that window/(until that window). Please note that the robot initially gives all the instructions before the game-play while a video is played for demonstration of the functionality of the game. </p> </li> <li> <p><em>T1_T1_rem/(_inc): </em>The number of times either of the two members in the team followed the pattern consecutively: I add an edge, I then delete it in that window/(until that window).</p> </li> <li> <p><em>T1_T1_add/(_inc): </em>The number of times either of the two members in the team followed the pattern consecutively: I delete an edge, I add it back in that window/(until that window).</p> </li> <li> <p><em>T1_T2_rem/(_inc): </em>The number of times the members of the team followed the pattern consecutively: I add an edge, you then delete it in that window/(until that window).</p> </li> <li> <p><em>T1_T2_add/(_inc): </em>The number of times the members of the team followed the pattern consecutively: I delete an edge, you add it back in that window/(until that window).</p> </li> <li> <p><em>redundant_exist/(_inc): </em>The number of times the team had redundant edges in their map in that window/(until that window).</p> </li> <li> <p><em>positive_valence/(_inc): </em>The average value of positive valence for the team in that window/(until that window).</p> </li> <li> <p><em>negative_valence/(_inc): </em>The average value of negative valence for the team in that window/(until that window).</p> </li> <li> <p><em>difference_in_valence/(_inc): </em>The difference of the average value of positive and negative valence for the team in that window/(until that window).</p> </li> <li> <p><em>arousal/(_inc): </em>The average value of arousal for the team in that window/(until that window).</p> </li> <li> <p><em>gaze_at_partner/(_inc): </em>The average of the the two team member's gaze when looking at their partner in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_robot/(_inc): </em>The average of the the two team member's gaze when looking at the robot in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_other/(_inc): </em>The average of the the two team member's gaze when looking in the direction opposite to the robot in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_screen_left/(_inc): </em>The average of the the two team member's gaze when looking at the left side of the screen in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_screen_right/(_inc):</em> The average of the the two team member's gaze when looking at the right side of the screen in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_speech_activity/(_inc): </em>The average of the two team member's speech activity in that window/(until that window). Each individual member's speech activity is calculated as a percentage of time that they are speaking in that window/(until that window). </p> </li> <li> <p><em>T_silence/(_inc): </em>The average of the two team member's silence in that window/(until that window). Each individual member's silence is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_short_pauses/(_inc): </em>The average of the two team member's short pauses over their speech activity in that window/(until that window). Each individual member's short pause refers to a brief pause of 0.15 seconds and is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_long_pauses/(_inc): </em>The average of the two team members long pauses over their speech activity in that window/(until that window). Each individual member's long pause refers to a pause of 1.5 seconds and is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_overlap/(_inc): </em>The average percentage of time the speech of the team members overlaps in that window/(until that window).</p> </li> <li> <p><em>T_overlap_to_speech_ratio/(_inc): </em>The ratio of the speech overlap over the speech activity of the team in that window/(until that window).</p> </li> </ul> <p>Apart from these 52 values, within each window, we also indicate: </p> <ul> <li><em>team: </em>The team to which the window belongs to.</li> <li><em>time_in_secs:</em> Time in seconds until that window.</li> <li><em>window: </em>The window number.</li> <li><em>normalized_time: </em>The time when this window occurred with respect to the total duration of the task for a particular team. </li> <li>cluster_labels: The cluster number associated with each time window in reference to the productive and non-productive clusters found in [3]</li> <li>PE_score: The Productive Engagement score in each window</li> </ul> <p>Lastly, we briefly elaborate on how the features are operationalised. We extract log behaviors from the recorded rosbags while the behaviors related to both gaze and affective states are computed through the open source library OpenFace [6] that returns both facial actions units (AUs) as well as gaze angles. For voice activity detection (VAD), that classifies if a piece of audio is voiced or unvoiced, we made use of the python wrapper for the open source Google WebRTC VAD. The literature that inspired our log, audio and video features as well as the tools used to extract them are described in more detail in [3,4]. However, in those papers, we make use of only the aggregate version of this data [5].</p> <p><em><strong>Please note that this dataset corresponds to the training data used in [7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint). </strong></em></p>
Neural correlates of the LSD experience revealed by multimodal neuroimaging
Open the record for dataset details and reuse information.
Robust joint registration of multiple stains and MRI for multimodal 3D histology reconstruction: Application to the Allen human brain atlas
Open the record for dataset details and reuse information.
Written and spoken digits database for multimodal learning
<p><strong>Database description:</strong></p> <p>The written and spoken digits database is not a new database but a constructed database from existing ones, in order to provide a ready-to-use database for multimodal fusion [1].</p> <p>The written digits database is the original MNIST handwritten digits database [2] with no additional processing. It consists of 70000 images (60000 for training and 10000 for test) of 28 x 28 = 784 dimensions.</p> <p>The spoken digits database was extracted from Google Speech Commands [3], an audio dataset of spoken words that was proposed to train and evaluate keyword spotting systems. It consists of 105829 utterances of 35 words, amongst which 38908 utterances of the ten digits (34801 for training and 4107 for test). A pre-processing was done via the extraction of the Mel Frequency Cepstral Coefficients (MFCC) with a framing window size of 50 ms and frame shift size of 25 ms. Since the speech samples are approximately 1 s long, we end up with 39 time slots. For each one, we extract 12 MFCC coefficients with an additional energy coefficient. Thus, we have a final vector of 39 x 13 = 507 dimensions. Standardization and normalization were applied on the MFCC features.</p> <p>To construct the multimodal digits dataset, we associated written and spoken digits of the same class respecting the initial partitioning in [2] and [3] for the training and test subsets. Since we have less samples for the spoken digits, we duplicated some random samples to match the number of written digits and have a multimodal digits database of 70000 samples (60000 for training and 10000 for test).</p> <p>The dataset is provided in six files as described below. Therefore, if a shuffle is performed on the training or test subsets, it must be performed in unison with the same order for the written digits, spoken digits and labels.</p> <p> </p> <p><strong>Files:</strong></p> <ul> <li>data_wr_train.npy: 60000 samples of 784-dimentional written digits for training;</li> <li>data_sp_train.npy: 60000 samples of 507-dimentional spoken digits for training;</li> <li>labels_train.npy: 60000 labels for the training subset;</li> <li>data_wr_test.npy: 10000 samples of 784-dimentional written digits for test;</li> <li>data_sp_test.npy: 10000 samples of 507-dimentional spoken digits for test;</li> <li>labels_test.npy: 10000 labels for the test subset.</li> </ul> <p> </p> <p><strong>References:</strong></p> <ol> <li>Khacef, L. et al. (2020), "Brain-Inspired Self-Organization with Cellular Neuromorphic Computing for Multimodal Unsupervised Learning".</li> <li>LeCun, Y. & Cortes, C. (1998), “MNIST handwritten digit database”.</li> <li>Warden, P. (2018), “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition”.</li> </ol>
Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)
<p>Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine. However, most publicly available datasets on human body movements cannot be used to study both problems in an out-of-the-lab movement acquisition setting. The objective of the VIDIMU dataset is to pave the way towards affordable patient tracking solutions for remote daily life activities recognition and kinematic analysis. </p> <p>The VIDIMU dataset includes 54 healthy young adults that were recorded on video and 16 of them were simultaneously recorded using custom IMUs. For each subject, 13 activities were registered using a low-resolution video camera and five Inertial Measurement Units (IMUs). Inertial sensors were placed in the lower or the upper limbs of the subject, respectively for activities that involve movement with the lower or the upper body. Video recordings were postprocessed using the state-of-the-art pose estimator <em>BodyTrack</em> (similar to OpenPose, and included in NVIDIA Maxine-AR-SDK) to provide a sequence of 3D joint positions for each movement. Raw IMU recordings were post-processed to compute joint angles by inverse kinematics with <em>OpenSim</em>. For recordings including simultaneous acquisition of video and IMU data types, these signals were used for data file synchronization. Collected data can be further used in applications related to human activity recognition and biomechanics related experiments in simulated home-like settings.</p> <p> </p> <p> </p>
Supporting data for "Entanglement between a Telecom Photon and an On-Demand Multimode Solid-State Quantum Memory"
<p>This repository contains the data supporting the article "Entanglement between a Telecom Photon and an On-Demand Multimode Solid-State Quantum Memory" by Jelena V. Rakonjac, Dario Lago-Rivera, Alessandro Seri, Margherita Mazzera, Samuele Grandi and Hugues de Riedmatten, Phys Rev Lett 2021.</p> <p>The data files used for the figures in the main text are included here, as well as a version of the final article submission.</p>
MEWL: Few-shot multimodal word learning with referential uncertainty
<p><strong>Dataset Release for <a href="https://arxiv.org/abs/2306.00503">MEWL: Few-shot multimodal word learning with referential uncertainty (ICML 2023) </a></strong></p> <p><strong>GitHub:</strong> <a href="https://github.com/jianggy/MEWL">https://github.com/jianggy/MEWL</a></p> <p><strong>Abstract: </strong>Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human's core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children's core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents' performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines.</p>
Synthetic Multimodal Dataset for Daily Life Activities
<p><strong>Outline</strong></p> <ul> <li>This dataset is originally created for the <a href="https://challenge.knowledge-graph.jp/2022/">Knowledge Graph Reasoning Challenge for Social Issue</a>s (KGRC4SI)</li> <li>Video data that simulates daily life actions in a virtual space from Scenario Data.</li> <li>Knowledge graphs, and transcriptions of the Video Data content ("who" did what "action" with what "object," when and where, and the resulting "state" or "position" of the object).</li> <li>Knowledge Graph Embedding Data are created for reasoning based on machine learning </li> <li>This data is open to the public as open data</li> </ul> <p><strong>Details</strong></p> <ul> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Movie">Videos</a></p> <ul> <li>mp4 format</li> <li>203 action scenarios</li> <li>For each scenario, there is a character rear view (file name ending in 0), an indoor camera switching view (file name ending in 1), and a fixed camera view placed in each corner of the room (file name ending in 2-5). Also, for each action scenario, data was generated for a minimum of 1 to a maximum of 7 patterns with different room layouts (scenes). A total of 1,218 videos</li> <li>Videos with slowly moving characters simulate the movements of elderly people.</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/RDF">Knowledge Graphs</a></p> <ul> <li>RDF format</li> <li>203 knowledge graphs corresponding to the videos</li> <li>Includes schema and location supplement information</li> <li>The schema is described below</li> <li><a href="http://kgrc4si.ml:7200/sparql">SPARQL endpoints</a> and <a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/tree/kgrc4si#%E3%83%8A%E3%83%AC%E3%83%83%E3%82%B8%E3%82%B0%E3%83%A9%E3%83%95%E3%81%AE%E4%BD%BF%E7%94%A8%E6%96%B9%E6%B3%95">query examples</a> are available</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Program">Script Data</a></p> <ul> <li>txt format</li> <li>Data provided to VirtualHome2KG to generate videos and knowledge graphs</li> <li>Includes the action title and a brief description in text format.</li> </ul> </li> <li>Embedding <ul> <li>Embedding Vectors in TransE, ComplEx, and RotatE. Created with DGL-KE (<a href="https://dglke.dgl.ai/doc/">https://dglke.dgl.ai/doc/</a>)</li> <li>Embedding Vectors created with jRDF2vec (<a href="https://github.com/dwslab/jRDF2Vec">https://github.com/dwslab/jRDF2Vec</a>).</li> </ul> </li> </ul> <p><strong>Specification of Ontology</strong></p> <ul> <li>Please refer to the specification for descriptions of all classes, instances, and properties: <a href="https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.html">https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.htm</a></li> </ul> <p><strong>Related Resources</strong></p> <ul> <li><a href="https://www.youtube.com/watch?v=Ajbn8hNXiZ8&list=PLHaRK-B0LUwjvrPgmIBTrf3DsPhmdnFTW">KGRC4SI Final Presentations with automatic English subtitles (YouTube)</a></li> <li><a href="https://github.com/aistairc/VirtualHome2KG">VirtualHome2KG (Software)</a></li> <li><a href="https://github.com/aistairc/virtualhome_unity_aist">VirtualHome-AIST (Unity</a>)</li> <li><a href="https://github.com/aistairc/virtualhome_aist">VirtualHome-AIST (Python API</a>)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_visualization">Visualization Tool</a> (Software)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_generation">Script Editor</a> (Software)</li> </ul>
Raw dataset for: High-throughput multimodal wide-field Fourier-transform Raman microscope
<p>This is the Raw spectral dataset of the data published in 10.1364/OPTICA.488860</p> <p>Data are arranged as follows:</p> <p>wavenumber [Nx1]</p> <p>Hyperspectrum_cube [Nx2, A, B]: hyperspectral datacube, where: Hyperspectrum_cube (1:N, :, :) is the real part; Hyperspectrum_cube (N+1:2N, :, :) is the imaginary part</p> <p>maximum [1x1]</p> <p>minimum [1x1].</p> <p>N: number of spectral bands</p> <p>A and B: size of the spatial coordinates</p> <p>Spectral amplitudes are obtained by: Hyperspectrum_cube=double(Hyperspectrum_cube)./(2.^16-1).*maximum+minimum</p>
Corticothalamic communication under analgesia, sedation and gradual ischemia: a multimodal model of controlled gradual cerebral ischemia in pig
Open the record for dataset details and reuse information.
JDC2015 - A multimodal dataset from facilitating multi-tabletop lessons in an open-doors day
<p><strong>IMPORTANT NOTE: Two of the files in this dataset are incorrect, see this dataset's errata at https://zenodo.org/record/204063 and https://zenodo.org/record/204819</strong></p> <p>This dataset contains eye-tracking, EEG, accelerometer, indoor location and video coding data from a single subject (a researcher with limited teacher experience), facilitating four maths lessons in a simulated multi-tabletop classroom, with four cohorts of 10-12 year old students, using tangible paper tabletops and a projector. These sessions were recorded in the frame of the MIOCTI project (http://chili.epfl.ch/miocti).</p> <p>This dataset has been used in several scientific works, such a submitted journal paper "Orchestration Load Indicators and Patterns: In-the-wild Studies Using Mobile Eye-tracking", by Luis P. Prieto, Kshitij Sharma, Lukasz Kidzinski & Pierre Dillenbourg (the analysis and usage of this dataset is available publicly at https://github.com/chili-epfl/paper-IEEETLT-orchestrationload)</p>
Dataset of the scientific paper " Multimodal robotic system for upper-limb rehabilitation in physical environment" (Advances in Mechanical Engineering)
<p>There are eight files with the following information:<br> - pos_stateXX.bin, binary file with information of the end effector position of the robot device in meters along the three axis (X, Y, Z) during state XX of the experiment<br> - target_stateXX.bin, binary file with information of the target position for the robot device in meters along the three axis (X, Y, Z) during state XX of the experiment<br> - emg_channelXX.bin, binary file with information of channel 1 of the EMG sensor in mV during during the whole time of the experiment<br> - color_stateXX.bin, binary file with information of color filter information during state XX of the experiment. This information is the percentage of pixels with the correct color (yellow, cyan or magenta) inside the region of interest</p> <p> </p>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
Raw experimental data for `Tailoring the Rotational Memory Effect in Multimode Fibers`
<p>Raw data for the article [**Tailoring the Rotational Memory Effect in Multimode Fibers**](https://arxiv.org/abs/2310.19337)</p><p>Measurement of transmission matrices and rotational memory effect for 4 segments of 50 micron core graded index multimode fibers with a numerical aperture of 0.2.</p>
Section 5.3 "Task Area 3: Multimodal data linking and integration" Figure 10
<p>Figure 10. Data flow to obtain a multimodal data structure (mmDS) with an overarching graph database (MUGDAT).</p> <p>from NFDI Grant Application, "<strong>National Research Data Infrastructure for Microscopy and Bioimage Analysis</strong>" (NFDI4BIOIMAGE)</p>
N20EMv2 dataset for automatic music transcription from multimodal singing
<p>N20EMv2 dataset for multimodal automatic music transcription from multimodal singing, presented in our TOMM 2024 paper, Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing. This dataset contains recordings of two modalities: audio and video. </p> <p>Our paper is available at: https://dl.acm.org/doi/10.1145/3651310.</p> <p>Code is available at: https://github.com/guxm2021/SVT_SpeechBrain</p> <p>Please cite our work as:</p> <pre>@article{gu2024automatic, title={Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing}, author={Gu, Xiangming and Ou, Longshen and Zeng, Wei and Zhang, Jianan and Wong, Nicholas and Wang, Ye}, journal={ACM Transactions on Multimedia Computing, Communications and Applications}, publisher={ACM New York, NY}, year={2024} }</pre>
ROCOv2: Radiology Objects in COntext Version 2, An Updated Multimodal Image Dataset
<p>Recent advances in deep learning techniques have enabled the development of systems for automatic analysis of medical images. These systems often require large amounts of training data with high quality labels, which is difficult and time consuming to generate.</p> <p>Here, we introduce Radiology Object in COntext Version 2 (ROCOv2), a multimodal dataset consisting of radiological images and associated medical concepts and captions extracted from the PubMed Open Access subset. Concepts for clinical modality, anatomy (X-ray), and directionality (X-ray) were manually curated and additionally evaluated by a radiologist. Unlike MIMIC-CXR, ROCOv2 includes seven different clinical modalities.</p> <p>It is an updated version of the ROCO dataset published in 2018, and includes 35,705 new images added to PubMed since 2018, as well as manually curated medical concepts for modality, body region (X-ray) and directionality (X-ray). The dataset consists of 79,789 images and has been used, with minor modifications, in the concept detection and caption prediction tasks of ImageCLEFmedical 2023. The participants had access to the training and validation sets after signing a user agreement.</p> <p>The dataset is suitable for training image annotation models based on image-caption pairs, or for multi-label image classification using the UMLS concepts provided with each image, e.g., to build systems to support structured medical reporting.</p> <p>Additional possible use cases for the ROCOv2 dataset include the pre-training of models for the medical domain, and the evaluation evaluation of deep learning models for multi-task learning.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.