Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
156
datasets available to search
ShareScore release 0.9.0
Dataset results
156 results for “Video dataset”
Dataset for Uplink-based Live Session Model for Stalling Prediction in Video Streaming
<p>This dataset presents aggregated YouTube streaming data used for uplink based quality impairment estimation. </p>
WildAvatar: Web-scale In-the-wild Video Dataset for 3D Avatar Creation
Open the record for dataset details and reuse information.
Dataset related to the following study: Humanoid attitudes influence humans in video and live interactions
Open the record for dataset details and reuse information.
Dataset of students' watching video record
Open the record for dataset details and reuse information.
A fMRI dataset in response to large number of short natural dynamic facial expression videos
<p><span>Facial expression is among the most natural methods for human beings to convey their emotional information in daily life. Although </span>the neural mechanism of facial expression has been extensively studied employing lab-controlled images and a <span>small number of</span> lab-controlled video stimuli, how the human brain processes <span>natural</span> facial expressions still needs to be investigated. <span>T</span>o our knowledge<span>,</span> this type of data <span>specifically</span> <span>on </span><span>large number of </span><span>natural</span> <span>facial expression videos</span> is currently missing. <span>W</span>e describe <span>here </span>the <span>natural </span>Facial Expressions Dataset (NFED), a fMRI dataset <span>including </span>responses to 1,320 short (<span>3-second</span>) <span>natural</span> facial expression video clips. <span>These video clips is </span><span>annotated</span> <span>with three types of labels: emotion, gender, and ethnicity</span><span>, along with accompanying metadata</span>. We validate that the dataset has good quality within and across participants and, notably, can capture temporal and spatial stimuli features. NFED provides researchers with fMRI data for understanding of the visual processing of large number of <span>natural </span> <span>facial expression videos.</span></p>
A fMRI dataset in response to large-scale short natural dynamic facial expression videos
<p>Facial expression is among the most natural methods for human beings to convey their emotional information in daily life. Although the neural mechanism of facial expression has been extensively studied employing lab-controlled images and a small number of lab-controlled video stimuli, how the human brain processes natural facial expressions still needs to be investigated. To our knowledge, this type of data specifically on large<span>-scale</span> natural facial expression videos is currently missing. We describe here the natural Facial Expressions Dataset (NFED), a fMRI dataset including responses to 1,320 short (3-second) natural facial expression video clips. These video clips is annotated with three types of labels: emotion, gender, and ethnicity, along with accompanying metadata. We validate that the dataset has good quality within and across participants and, notably, can capture temporal and spatial stimuli features. NFED provides researchers with fMRI data for understanding of the visual processing of large number of natural facial expression videos.</p>
GazeMining: A Dataset of Video and Interaction Recordings on Dynamic Web Pages. Labels of Visual Change, Segmentation of Videos into Stimulus Shots, and Discovery of Visual Stimuli.
<p><strong>Recording setup</strong><br> Recordings have been taken place on 12th March 2019. Gaze data has been recorded with a Tobii 4C eye tracker with Pro license at 90 Hz. Resolution of the viewport was set to 1024x768. The display had a size of 24 inches and a resolution of 1680x1050 pixels. We polled the DOM tree every 50 milliseconds for fixed elements. We recorded the Web browsing of four participants, who followed the protocol as stored under "Dataset_visual_change/Instructions.doc".</p> <p><strong>Description of the dataset</strong><br> The dataset consists of following three subsets.</p> <p><em>1. Dataset_visual_change</em><br> The recordings of each participant p1-p4 on twelve Web sites are in the corresponding directories. For each Web site, there are nine to eleven files:</p> <ul> <li><site>.json: datacast</li> <li><site>.webm: video recording</li> <li><site>.features.csv: computer-vision features per observation</li> <li><site>.features_meta.csv: meta information about features</li> <li><site>.labels-l<X>.csv: labels of observations</li> <li><site>_meta.csv: meta information about recording</li> <li><site>_scroll_cache.csv: cache of estimated scrolling</li> <li><site>_scroll_cache_map.csv: mapping of observations to scroll cache entries</li> <li><site>_times.csv: timestamps of frames in the video recording</li> <li><site>_layer_pixels.csv: first row is the pixel count of root layer, second row is pixel count of all fixed elements</li> </ul> <p><em>2. Dataset_stimuli</em><br> Stimulus shots and visual stimuli computed with the framework. Value-based, edge-based, signal-based, and SIFT-based features have been used. The labels of the first participant's session had been used to train a random forest classifier with 100 trees for visual change classification, using the named features. The discovery has been performed on each Web site from the dataset and<br> the results are placed in the respective directories. Inside each directory, there is one directory for the detected shots and one for the discovered stimuli. In the shots directory, there is one overview as <participant>_<site>.csv file. For each shot, there are four further files:</p> <ul> <li><participant>_<site>_<shot>.png: stitched frame of the stimulus shot</li> <li><participant>_<site>_<shot>-blind.csv: frames from animations that are not contributing to the stitched frame</li> <li><participant>_<site>_<shot>-gaze.csv: gaze data (in stitched frame space)</li> <li><participant>_<site>_<shot>-mouse.csv: mouse data (in stitched frame space)</li> </ul> <p>The shots have been merged to stimuli, which are placed in the stimuli directory. The stimuli are grouped per layer (scrollable, fixed elements, etc.) and meta information is available in <layer_index>-<xpath>-meta.csv files. Furthermore, there are directories per layer, storing the discovered stimuli. Each discovered visual stimulus is represented by four files:</p> <ul> <li><stimulus_id>.png: stitched frame of the visual stimulus</li> <li><stimulus_id>-gaze.csv: gaze data (in stitched frame space)</li> <li><stimulus_id>-mouse.csv: mouse data (in stitched frame space)</li> <li><stimulus_id>-shots.csv: contained stimulus shots</li> </ul> <p><em>3. Dataset_evaluation</em><br> We have performed two evaluations of the visual stimuli discovery. One computational estimating the quality of stimuli. One case-study of an expert's task. There are two respective directories with the annotation data.</p> <p><strong>Changelog</strong><br> [1.0.2] Add counts of layer pixels per participant.<br> [1.0.1] Change to CC0 license.<br> [1.0.1] Add labels of third annotator "l3".<br> [1.0.0] Initial release.</p>
A comprehensive video dataset for Multi-Modal Recognition Systems
<p>A fully-labelled video dataset will act as a unique resource for researchers and analysts in the fields such as machine learning, computer vision and deep learning. The videos contain similar text recited by 67 different subjects. The text contains digits from 1 to 20 recited by 67 different subjects within the same experimental setup.</p>
YM2413-MDB: A Multi-Instrumental FM Video Game Music Dataset with Emotion Annotations
<p>YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations. It includes 669 audio and MIDI files of music from Sega and MSX PC games in the 80s using YM2413, a programmable sound generator based on FM. The collected game music is arranged with a subset of 15 monophonic instruments and one drum instrument. They were converted from binary commands of the YM2413 sound chip. Each song was labeled with 19 emotion tags by two annotators and validated by three verifiers to obtain refined tags</p> <p>For more detailed information about the dataset, please refer to our paper: <a href="https://arxiv.org/abs/2211.07131">YM2413-MDB: A Multi-Instrumental FM Video Game Music Dataset with Emotion Annotations</a>.</p> <p><strong>File Description</strong></p> <p><strong>1) Pure data</strong></p> <p>- original_vgms: crawled vgm files from <a href="https://www.smspower.org/">SMS POWER</a> and <a href="https://vgmrips.net/packs/">VGMRIPs</a></p> <p>- wav: rendered vgm files using <a href="https://github.com/vgmrips/vgmplay">VGMPlay</a></p> <p> </p> <p><strong>2) MIDI data</strong></p> <p>- midi/vgmplay_log_to_midi: converted midi files</p> <p>- midi/adjust_tempo: add postprocessing(metrically aligned using wav_downbeat files) after midi conversion</p> <p>- midi/adjust_tempo_remove_delayed_inst: add postprocessing(metrically aligned using wav_downbeat files, remove delayed instrument) after midi conversion</p> <p> </p> <p><strong>3) Metadata</strong></p> <p>- emotion_annotation/verified_annotation.csv: contains emotion annotation for each songs</p> <p>- tags_kor_eng.txt: Korean <-> English tag dictionary</p> <p> </p> <p><strong>4) Useful middle-time step data</strong></p> <p>- wav_downbeat: extracted downbeat values using TCNBeatTracker of <a href="https://github.com/CPJKU/madmom">madmom</a></p> <p>- vgm_txts: disassembled vgm files as txt using <a href="https://github.com/vgmrips/vgmtools#vgm-text-writer-vgm2txt">vgm2txt</a></p> <p>- ydr: YM2413 Disassembly Raw(YDR). command list of vgm files. generated by reading vgm_txts</p> <p> </p> <p><strong>Update Log</strong></p> <p>- version 1.0.1: Fix ticks per beat value adjust to tempo where tempo values are not 150. Also, madmom downbeat files are updated from DBNBeatTracker(ISMIR, 2015) to TCNBeatTracker(Newer one EUSIPCO, 2019).</p> <p>- version 1.0.2: <strong><a href="https://github.com/jech2/YM2413-MDB/issues/2">Wrong emotion tag issue in the verification annotation file was fixed.</a></strong></p>
video games recommendations dataset crafted by human experts
<p>video games recommendations dataset crafted by human experts</p>
MM-Fit Dataset - physical exercise workout sessions (video)
<p>This repository contains the video recording of 20 physical exercise workout sessions, which is part of the <strong>MM-Fit Dataset. </strong>Further details about our dataset and other time synchronised sensor data modalities (accelerometer, gyroscope, heart rate) collected with wearable devices during these workout sessions can be found on the project page: https://mmfit.github.io/</p> <p>If you find our dataset useful and use it in your work please cite our paper:</p> <p>David Strömbäck, Sangxia Huang, Valentin Radu, <strong>MM-Fit: Multimodal Deep Learning for Automatic Exercise Logging Across Sensing Devices</strong>, ACM Interactive, Mobile, Wearable and Ubiquitous Technologies (IMWUT): Volume 4 Issue 4, December 2020.</p> <pre>@article{stromback2020mm, title={Mm-fit: Multimodal deep learning for automatic exercise logging across sensing devices}, author={Str{\"o}mb{\"a}ck, David and Huang, Sangxia and Radu, Valentin}, journal={Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies}, volume={4}, number={4}, pages={1--22}, year={2020}, publisher={ACM New York, NY, USA} }</pre>
Human embryo time-lapse video dataset
<p>A human embryo time-lapse video dataset.</p>
Arabic words sign language video dataset
<p>The dataset consists of 3000 unaltered videos captured by five volunteers, with each volunteer performing 20 repetitions of 30 signs, resulting in approximately 600 videos per volunteer. These videos were recorded without stabilization tools, reflecting the prevalent use of smartphones with built-in cameras. They encompass diverse resolutions, locations, places, and backgrounds, providing a comprehensive representation of real-life scenarios.</p>
The benchmark datasets for object tracking in satellite videos
<p>A new type of earth observation satellite uses "gaze" method to continuously observe a certain area, and uses "video recording" method to record dynamic information and analyze its instantaneous characteristics. With the development of dynamic acquisition technology for satellite video data, significant progress in target tracking has been made in recent years, which plays an important role in monitoring rapidly changing events. Different from targets in ordinary videos, targets in satellite videos usually demonstrate the phenomenon of a small size occupation pattern (small) and weak feature capture (dim) due to occlusion, illumination variation, and confusion with the surroundings.</p>
DroneZaic Dataset: a robust end-to-end pipeline for mosaicking freely flown aerial video of agricultural fields
Open the record for dataset details and reuse information.
Teahing method analysis - lecture video dataset
<p>Lecture Videos dataset, these videos are collected using CCTV camera, later these are split into 3 seconds clips and stored in folders based on action performed in the clips. each folder represent sest of 2 actions seperated by _ underscore sign.</p>
A video dataset for wooden box assembly
<p>This is a dataset of videos for wooden box assembly. The main strength of this dataset is the design of standard and uniform workflow and the use of multiple cameras capturing videos from different angles. A total of 62 videos of 17 subjects were collected. The duration of the videos is 13.0 hours. The whole workflow is designed into nine steps. The dataset contains the videos of the assembly process and the temporal annotation data for the nine steps in each video. Our dataset could be used to facilitate the studies in different applications such as object recognition, human action classification, intelligent automation etc.</p>
VidTIMIT Audio-Video Dataset
<p>See http://conradsanderson.id.au/vidtimit/ for details.<br> <br> Summary: Video and corresponding audio recordings of 43 people, reciting short sentences. Useful for research on topics such as automatic lip reading, multi-view face recognition, multi-modal speech recognition and person identification.</p> <p> </p>
WebRTC-QoE: A Dataset of Quality of Experience in Audio-Video Communications
<p>In the realm of real-time communications, WebRTC-based multimedia applications are increasingly prevalent as these can be smoothly integrated within Web browsing sessions. The browsing experience is then significantly improved concerning scenarios where browser add-ons and/or plug-ins are used; still, the end user's Quality of Experience (QoE) in WebRTC sessions may be affected by network impairments, such as delays and losses. Due to the variability in user perceptions under different communications scenarios, comprehending and enhancing the resulting service quality is a complex endeavour. To address this, we present a dataset that provides a comprehensive perspective on the conversational quality of a two-party WebRTC-based audiovisual telemeeting service. This dataset was gathered through subjective evaluations involving 20 subjects across 15 different test conditions (TCs). A specialized system was developed to induce controlled network disruptions such as delay, jitter, and packet loss rate, which adversely affected the communication between the parties. This methodology offered insight into user perceptions under various network impairments. The dataset encompasses a blend of objective and subjective data, including ACR (Absolute Category Rating) subjective scores, <em>webrtc-internals</em> parameters, facial expressions features, and speech features. Consequently, it serves as a substantial contribution to the improvement of WebRTC-based video call systems, offering practical and real-world data that can drive the development of more robust and efficient multimedia communication systems, thereby enhancing the user’s experience.</p> <div> <p> </p> <p> In the following, we discuss the details of the provided datasets.</p> <p><strong>Subjective_results_dataset.csv:</strong> This dataset encompasses subjective evaluation results from 20 subjects (users) who assessed the quality of WebRTC-based video calls under 15 distinct test conditions (TCs), which included combinations of 3 network impairments (delay, jitter, packet loss) to disturb the communication. The single discrete Absolute Category Rating (ACR) scale with five category labels (1-Bad, 2-Poor, 3-Fair, 4-Good, and 5-Excellent) was used by the users to rate the perceived QoE. A total of 300 ACR scores were obtained (20 participants x 15 TCs). The size of this dataset is 4.00 KB.</p> <p>The significance of each column is explained as follows:</p> <ul> <li>Test Condition (TC): It enumerates the TC numbers, which span from 1 to 15.</li> <li>Delay [ms]: It refers to the time it takes for a signal to travel from one point to another and is represented at three levels: 0 ms (no delay), 500 ms (moderate delay), and 1000 ms (significant delay).</li> <li>Jitter [ms]: It refers to the variability in delay and is represented at two levels: 0 ms (no jitter) and 500 ms (moderate jitter).</li> <li>Packet Loss Rate [%]: It refers to the loss of data packets during transmission and is represented at three levels: 0% (no packet loss), 15% (moderate packet loss), and 30% (significant packet loss).</li> <li>User: The users, identified from 1 to 20, participated in 15 video calls and evaluated the quality of these calls on a scale from 1 to 5.</li> </ul> <p><strong>Webrtc_internals_dataset.zip:</strong> This dataset contains text files collected using the webrtc-internals tool during the video calls. The zip is organized into 20 distinct folders, each labeled from ‘User1’ to ‘User20’. Each user folder contains 15 text files (.txt), each named following the pattern ‘webrtc_internals_dump-TCx_y-z-t.txt’. In this naming convention, ‘x’ denotes the TC number, ranging from 1 to 15. The ‘y-z-t’ segment varies with each TC, where ‘y’ signifies the delay value (0, 500, or 1000), ‘z’ indicates the jitter value (0 or 500), and ‘t’ represents the packet loss rate (0, 15, or 30). Each text file includes application-level data concerning WebRTC sessions' statistics in a JSON format. A total of 300 webrtc-internals dump text files were obtained (20 participants x 15 TCs). The size of this zip dataset is 28.7 MB (396 MB uncompressed). </p> <p><strong>Facial_expression_features_dataset.zip:</strong> This dataset contains facial expression features extracted from the recorded videos with face images using the OpenFace toolkit. The zip is organized into 20 distinct folders, each labelled from ‘User1’ to ‘User20’. Each user folder contains 15 .csv files, which are named following the pattern ‘TCx_y-z-t.csv’. In this naming convention, ‘x’ denotes the TC, ranging from 1 to 15. The ‘y-z-t’ segment varies with each TC, where ‘y’ signifies the delay value (0, 500, or 1000), ‘z’ indicates the jitter value (0 or 500), and ‘t’ represents the packet loss rate (0, 15, or 30). A total of 300 facial expression feature files in .csv format were obtained (20 participants x 15 TCs). For each face image of each TC, the OpenFace outputs 6 gaze direction features, 280 eye region landmarks and 35 Action Units (AUs). The size of this zip dataset is 547 MB (2.14 GB uncompressed). </p> <p>Each column carries a specific significance, which is elaborated as follows:</p> <ul> <li>frame: the frame number in the context of sequences.</li> <li>face_id: the identifier assigned to each face when multiple faces are present.</li> <li>timestamp: the elapsed time in seconds during the processing of a video sequence.</li> <li>confidence: the level of confidence the tracker has in the current landmark detection estimate.</li> <li>success: a face has been detected in the frame and it has been tracked accurately.</li> <li>gaze_0_x, gaze_0_y, gaze_0_z: the normalized eye gaze direction vector in world coordinates for eye 0, which is the eye on the left in the image.</li> <li>gaze_1_x, gaze_1_y, gaze_1_z: the normalized eye gaze direction vector in world coordinates for eye 1, which is the eye on the right in the image.</li> <li>gaze_angle_x, gaze_angle_y: the eye gaze direction, averaged for both eyes and expressed in world coordinates in radians, is converted into a format that is easier to use than gaze vectors.</li> <li>eye_lmk_x_0, eye_lmk_x_1, ..., eye_lmk_x55, eye_lmk_y_1, ... eye_lmk_y_55: the pixel coordinates of 2D landmarks in the eye region.</li> <li>eye_lmk_X_0, eye_lmk_X_1, ..., eye_lmk_X55, eye_lmk_Y_0, ..., eye_lmk_Z_55: the position of landmarks in the eye region in 3D space, measured in millimeters.</li> <li>17 AUr: detect the activation intensity (from 1 to 5) of a particular facial muscle. These are: AU01_r, AU02_r, AU04_r, AU05_r, AU06_r, AU07_r, AU09_r, AU10_r, AU12_r, AU14_r, AU15_r, AU17_r, AU20_r, AU23_r, AU25_r, AU26_r, AU45_r.</li> <li>18 AUc: Identify the activation of a particular muscle and note its presence (0 for absent, 1 for present). These are: AU01_c, AU02_c, AU04_c, AU05_c, AU06_c, AU07_c, AU09_c, AU10_c, AU12_c, AU14_c, AU15_c, AU17_c, AU20_c, AU23_c, AU25_c, AU26_c, AU28_c, AU45_c.</li> </ul> <p><strong>Speech_features_dataset.csv:</strong> This dataset is a robust assembly of speech features extracted from the recorded audio files using the OpenSMILE toolkit. The dataset is further enhanced with the integration of Absolute Category Rating (ACR) scores, which were assigned by each subject for every TC. These scores are embedded within the speech features of each subject. The dataset is exhaustive, encompassing a total of 1,911,900 speech features (calculated as 15 TCs x 20 subjects x 6373 speech features per subject). The total size of this dataset is 18.9 MB. </p> <p>Each column carries a specific significance, which is elaborated as follows:</p> <ul> <li>file: It presents the audio files from which the speech features were extracted.</li> <li>Users: The users, identified from 1 to 20, participated in 15 video calls and evaluated the quality of these calls on a scale from 1 to 5.</li> <li>Test Condition (TC): It enumerates the Test Condition numbers, which span from 1 to 15.</li> <li>Delay [ms]: It refers to the time it takes for a signal to travel from one point to another and is represented at three levels: 0 ms (no delay), 500 ms (moderate delay), and 1000 ms (significant delay).</li> <li>Jitter [ms]: It refers to the variability in delay and is represented at two levels: 0 ms (no jitter) and 500 ms (moderate jitter).</li> <li>Packet Loss Rate [%]: It refers to the loss of data packets during transmission and is represented at three levels: 0% (no packet loss), 15% (moderate packet loss), and 30% (significant packet loss).</li> <li>OpenSmile Speech Features Columns: This presents a detailed view of the speech features. Specifically, from column G to column IKI, it enumerates 6373 distinct speech features for each user under each TC of each processed audio file in .wav.</li> <li>ACR Score: The Absolute Category Rating (ACR) scale ranges from 1, representing the lowest quality, to 5, indicating the highest quality evaluated by the subjects.</li> </ul> <p> </p> <p><strong>If you make use of this dataset, please consider citing the following publication:</strong></p> <p>Bingol G., Porcu, S., Floris, A., & Atzori, L. (2024). WebRTC-QoE: A dataset of QoE assessment of subjective scores, network impairments, and facial & speech features. Computer Networks, 244, 110356, doi: 10.1016/j.comnet.2024.110356.</p> <p>BibTex format:</p> <p>@article{bingol2024datasetwebrtc, title={WebRTC-QoE: A dataset of QoE assessment of subjective scores, network impairments, and facial & speech features}, author={Bingol, Gulnaziye and Porcu, Simone and Floris, Alessandro and Atzori, Luigi}, journal={Computer Networks}, volume={244}, pages={110356}, year={2024}, publisher={Elsevier}, doi = {https://doi.org/10.1016/j.comnet.2024.110356} }</p> </div>
TACDEC: Dataset of Tackle Events in Soccer Game Videos
<p>TACDEC is a dataset of tackle events in soccer game videos. Recognizing the gap in existing open datasets that predominantly focus on official soccer events such as goals and cards, TACDEC targets a comprehensive analysis of tackles — a critical aspect of soccer that combines technical skills, tactical decision-making, and physical engagement. By leveraging video data from the Norwegian Eliteserien league across multiple seasons, we annotated 425 videos with 4 types of tackle events, categorized into "tackle-live", "tackle-replay", "tackle-live-incomplete", and "tackle-replay-incomplete", yielding a total of 836 event annotations. The dataset offers an unprecedented resource for the development and testing of machine learning models aimed at understanding and analyzing soccer game dynamics.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.