Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
298
datasets available to search
ShareScore release 0.9.0
Dataset results
298 results for “Multi-modal”
EarSet: A Multi-Modal In-Ear Dataset
<p>EarSet aims at providing the research community with a novel, multi-modal, dataset, which, for the first time, will allow studying of the impact of body and head/face movements on both the morphology of the PPG wave captured at the ear, as well as on the vital signs estimation. To accurately collect in-ear PPG data, coupled with a 6 degrees-of-freedom (DoF) motion signature, we prototyped and built a flexible research platform for in-the-ear data collection. The platform is centered around a novel ear-tip design which includes a 3-channel PPG (green, red, infrared) and a 6-axis (accelerometer, gyroscope) motion sensor (IMU) co-located on the same ear-tip. This allows the simultaneous collection of spatially distant (i.e., one tip in the left and one in the right ear) PPG data at multiple wavelengths and the corresponding motion signature, for a total of 18 data streams. <br> Inspired by the Facial Action Coding Systems (FACS), we consider a set of potential sources of motion artifact (MA) caused by natural facial and head movements. Specifically, we gather data on 16 different head and facial motions - head movements (nodding, shaking, tilting), eyes movements (vertical eyes movements, horizontal eyes movements, brow raiser, brow lowerer, right eye wink, left eye wink), and mouth movements (lip puller, chin raiser, mouth stretch, speaking, chewing).<br> We also collect motion and PPG data under activities, of different intensities, which entail the movement of the entire body (walking and running). Together with in-ear PPG and IMU data, we collect several vital signs including, heart rate, heart rate variability, breathing rate, and raw ECG, from a medical-grade chest device.</p> <p>With approximately 17 hours of data from 30 participants of mixed gender and ethnicity (mean age: 28.9 years, standard deviation: 6.11 years), our dataset empowers the research community to analyze the morphological characteristics of in-ear PPG signals with respect to motion, device positioning (left ear, right ear), as well as a set of configuration parameters and their corresponding data quality/power consumption trade-off. We envision such a dataset could open the door to innovative filtering techniques to mitigate, and eventually eliminate, the impact of MA on in-ear PPG. We ran a set of preliminary analyses on the data, considering both handcrafted features, as well as a DNN (Deep Neural Network) approach. Ultimately, we observe statistically significant morphological differences in the PPG signal across different types of motions when compared to a situation where there is no motion. We also discuss a 3-classes classification task and show how full-body motions and head/face motions can be discriminated from a still baseline (and among themselves). These preliminary results represent the first step towards the detection of corrupted PPG segments and show the importance of studying how head/face movements impact PPG signals in the ear. </p> <p>To the best of our knowledge, this is the first in-ear PPG dataset that covers a wide range of full-body and head/facial motion artifacts. Being able to study the signal quality and motion artifacts under such circumstances will serve as a reference for future research in the field, acting as a stepping stone to fully enable PPG-equipped earables.</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from V3C1+V3C2 Dataset
<p>This repository contains a diverse set of features extracted from the V3C1+V3C2 dataset, sourced from the Vimeo Creative Commons Collection. These features were utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] during the latest editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>).</p> <p>The original V3C1+V3C2 dataset, provided by NIST, can be downloaded using the instructions provided at <a href="https://videobrowsershowdown.org/about-vbs/existing-data-and-tools/">https://videobrowsershowdown.org/about-vbs/existing-data-and-tools/</a>.</p> <p>It comprises 7,235 video files, amounting for 2,300h of video content and encompassing 2,508,113 predefined video segments.</p> <p>We subdivided the predefined video segments longer than 10 seconds into multiple segments, with each segment spanning no longer than 16 seconds. As a result, we obtained a total of 2,648,219 segments. For each segment, we extracted one frame, specifically the middle one, and computed several features, which are described in detail below.</p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <pre>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} }</pre> </blockquote> <p> </p> <p>This repository comprises the following files:</p> <ul> <li><strong><em>msb.tar.gz </em></strong> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione"). </li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original V3C videos available.</li> <li><strong><em>features-aladin.tar.gz<sup>†</sup></em> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip-laion.tar.gz<sup>†</sup></strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz<sup>†</sup> </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip2video.tar.gz<sup>†</sup> </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the video segments. <strong> </strong>In particular 1) we concatenate consecutive short segments so to create segments at least 3 seconds long; 2) we downsample the obtained segments to 2.5 fps; 3) we feed the network with the first min(36, n) frames, where n is the number of frames of the segment. Notice that the minimum processed length consists of 7 frames, given that the segment is no shorter than 3s. </li> <li><em><strong>objects-frcnn-oiv4.tar.gz<sup>*</sup> </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]). </li> <li><em><strong>objects-mrcnn-lvis.tar.gz<sup>*</sup></strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz<sup>*</sup></strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong><sup>†</sup>Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the V3C1+V3C2 dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used. Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8514-8523).</p>
Multi-modal Dataset of Human Activities of Daily Living with Ambient Audio, Vibration and Environmental Data
<pre>This dataset provides over 43000 samples of 25 different human activities (e.g. walking, opening/closing a door, sitting down, vacuum cleaning). Each sample is recorded by 5 sensor devices with multiple types of sensors. The main part is audio and vibration. Further metrics were recorded at a low frequency, and are: infrared array, light color, temperature, relative humidity, atmospheric pressure, air quality measure, volatile organic compounds and CO2 equivalent. The data was recorded in supervised sessions to label each sample. The recording environment consisted of a kitchen and dining room. Flawed samples were removed and the different metrics were synchronized, but no further processing or filtering of the data was performed. </pre>
Repetitive Transcranial Magnetic Stimulation and Multi-modality Aphasia Therapy for Post-stroke Non-fluent Aphasia
ClinicalTrials.gov study NCT04102228. IPD Sharing: NO. Countries: 1. Publications: 1.
LCPUFA Supplementation: A Multi-Modality Imaging Study
ClinicalTrials.gov study NCT02076048. IPD Sharing: Not stated. Countries: 1. Publications: 1.
The Carotid Artery Multi-modality Imaging Prognostic (CAMP) Study
ClinicalTrials.gov study NCT04679727. IPD Sharing: Not stated. Countries: 1. Publications: 9.
Multi-Modal Longevity Protocol Including Autologous Cell-Free Conditioned Media
ClinicalTrials.gov study NCT07322224. IPD Sharing: NO. Countries: 1. Publications: 0.
Multi-Modal Monitoring of Disease Symptoms in Myasthenia Gravis
ClinicalTrials.gov study NCT07224386. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Impact on Quality of Life From Multi-modality Lung Cancer
ClinicalTrials.gov study NCT04540757. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Evaluation of Kidney Function by Multi-modal Magnetic Resonance Imaging and Spectroscopy
ClinicalTrials.gov study NCT00575432. IPD Sharing: UNDECIDED. Countries: 1. Publications: 12.
Identification of Multi-modal Bio-markers for Early Diagnosis and Treatment Prediction in Schizophrenia Individuals
ClinicalTrials.gov study NCT03790085. IPD Sharing: NO. Countries: 1. Publications: 8.
Multi-modality Imaging in Acute Myocardial Infarction
ClinicalTrials.gov study NCT02926755. IPD Sharing: NO. Countries: 1. Publications: 1.
Multi-modal Neuroimaging in Alzheimer's Disease
ClinicalTrials.gov study NCT01554202. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Multi-Modal Intervention In Frail And Prefrail Older People With Type 2 Diabetes
ClinicalTrials.gov study NCT01654341. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Multi-modality Prostate Cancer Image Guided Interventions
ClinicalTrials.gov study NCT04009174. IPD Sharing: YES. Countries: 1. Publications: 2.
rTMS and Multi-Modality Aphasia Therapy for Post-Stroke Aphasia
ClinicalTrials.gov study NCT04901156. IPD Sharing: NO. Countries: 1. Publications: 1.
Multi-modality Imaging and Collection of Biospecimen Samples in Understanding Bone Marrow Changes in Patients With Acute Myeloid Leukemia Undergoing TBI and Chemotherapy
ClinicalTrials.gov study NCT03422731. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Standardization of Multi-modal Tumor Ablation Therapy System
ClinicalTrials.gov study NCT03223142. IPD Sharing: Not stated. Countries: 1. Publications: 6.
Early Identification of Mental Disorders: Application of a Multi-modal & Domains System
ClinicalTrials.gov study NCT05939154. IPD Sharing: YES. Countries: 1. Publications: 8.
Systematically Assessing Effects of Colored Light on Humans With a Multi-modal Approach (Substudy 1 of 7)
ClinicalTrials.gov study NCT02882516. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.