Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

156

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

156 results for “Video dataset”

Learn how ShareScore rates datasets ↗
zenodo48/100

Disinformation on YouTube: A dataset of YouTube comments on videos related to claims made by Trump and Vance on Haitian immigrants

<div> <div> <div> <div> <div> <p>The corpus&nbsp;contains three files. First, the youtube_haitian_disinformation_videos_meta.csv file includes comments and YouTube video metadata. Data is organized around per video information. The columnar&nbsp;values are:&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>video_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>date (video publication date)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>title&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>description&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>channel_title&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>transcript&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>transcript_str (Video transcript without timestamps)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>views&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>likes&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comments (all comments per video)&nbsp;</p> </li> </ul> </div> <div> <div> <div> <p>Second, the youtube_haitian_disinformation_comment_reply_metadata.csv file includes comments, replies, and comment metadata. Each comment occupies its own row in the spreadsheet. The columnar field are:&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>video_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_date&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_like_count&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>author&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>in_reply_to&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>neg, neu, pos, compound (VADER polarity scores)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>named_entities (spaCy named entity tags with tokens)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>emoji (spaCy Emojis and token spans)&nbsp;</p> </li> </ul> </div> </div> <div> <div> <div> <div> <p>Comments and associated metadata are represented in individual rows.&nbsp;</p> </div> <div> <p>The third file contains the results of the TFIDF analysis described herein. The TFIDF analysis features the top 5000 terms weights for the comments to each video in a .csv file.&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>youtube_disinfo_comments_tfidf_results_per_video.csv&nbsp;</p> </li> </ul> </div> </div> </div>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Videos of the processed microscope images and time series of the petrophysical parameters from image processing and geochemical simulation and of the measured induced polarisation [Video][Dataset]

<p>Supporting Information for the manuscript&nbsp;<em>Microfluidics and&nbsp;spectral induced polarization for direct observation and petrophysical modeling of calcite dissolution</em> published in Geophysical Research Letters</p> <ul> <li><strong>Data Set S1.</strong> Porosity, water saturation, and calcite sample perimeter from image<br>processing.</li> <li><strong>Data Set S2.</strong> Porosity, water conductivity, and pH from geochemical simulation.</li> <li><strong>Data Set S3.</strong> Real and imaginary components of the complex electrical conductivity at<br>2.5 Hz and CEC from petrophysical modeling.</li> <li><strong>Movie S1.</strong> Dissolution of the calcite sample with the detected contour superimposed in<br>white on the grayscale images. Time, length scale, and flow direction are indicated. In<br>case of problems launching the file, we recommend using VLC Media Player software.</li> <li><strong>Movie S2.</strong> Segmented images of the CO2 bubbles produced by the calcite dissolution.<br>Time, length scale, and flow direction are indicated. In case of problems launching the<br>file, we recommend using VLC Media Player software.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (VIDIMU)

<p>Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine. However, most publicly available datasets on human body movements cannot be used to study both problems in an out-of-the-lab movement acquisition setting. The objective of the VIDIMU dataset is to pave the way towards affordable patient tracking solutions for remote daily life activities recognition and kinematic analysis.&nbsp;</p> <p>The VIDIMU dataset includes 54 healthy young adults that were recorded on video and 16 of them were simultaneously recorded using custom IMUs.&nbsp; For each subject, 13 activities were registered using a low-resolution video camera and five Inertial Measurement Units (IMUs). Inertial sensors were placed in the lower or the upper limbs of the subject, respectively for activities that involve movement with the lower or the upper body. Video recordings were postprocessed using the state-of-the-art pose estimator <em>BodyTrack</em> (similar to OpenPose, and&nbsp;included in NVIDIA Maxine-AR-SDK) to provide a sequence of 3D joint positions for each movement. Raw IMU recordings were post-processed to compute joint angles by inverse kinematics with <em>OpenSim</em>. For recordings including simultaneous acquisition of video and IMU data types, these signals were used for data file synchronization. Collected data can be further used in applications related to human activity recognition and biomechanics related experiments in simulated home-like settings.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

AutoML for Video Analytics with Edge Computing - Dataset

<p>Latency and confidence measurements obtained from an edge-assisted object recognition system.</p> <p>The records are obtained by tuning the image encoding rate and Neural Network input layer size and measuring the latency on completing each system operation, i.e. encoding the image, transmiting it wirelessly, decoding and rotating at the server, and performing object recognition with YOLO on the server&#39;s GPU. Moreover, we document the achievable frame rate as a result of the total latency, as well as the object recognition confidence and cumulative confidence for all identified objects of each image.</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Design of an Ontology-Driven Constraint Tester (ODCT) and Application to SAREF & Smart Energy Appliances: Datasets, SHACL Shapes, Demo Video of Web Application, and Detailed Performance Reports

<h2>Description</h2> <p>This repository presents the resources used for validating the compliance of <strong>smart energy appliances</strong> against the <strong>Smart Appliances REFerence (SAREF)</strong> ontology and its extension <strong>SAREF4ENER</strong>, as part of the <strong>Ontology-Driven Constraint Tester (ODCT)</strong> project. The ODCT tool is specifically designed to ensure <strong>semantic interoperability</strong> and adherence to standardized ontological frameworks, which are crucial for integrating smart devices into modern energy management systems.</p> <h2>ODCT Overview</h2> <p>The <strong>Ontology-Driven Constraint Tester (ODCT)</strong> is a robust framework created to validate datasets against ontologies defined by <strong>SAREF</strong> and <strong>SAREF4ENER</strong>, both established under ETSI SmartM2M. This tool has been applied to the <strong>Flexible Start use case</strong> from the <strong>Joint Research Centre&rsquo;s (JRC) Code of Conduct for Energy Smart Appliances</strong>. The ODCT tool ensures that smart devices like energy-efficient washing machines, thermostats, and connected lighting operate in compliance with established ontologies, thereby enhancing their <strong>interoperability</strong> within energy management systems and smart grids.</p> <h2>Repository Contents</h2> <p>This repository contains essential resources used in the ODCT compliance testing process:</p> <ul> <li> <p><strong>Compliant Dataset</strong>: This dataset represents a fully compliant scenario where no errors are present in the smart energy appliances&rsquo; profiles, demonstrating the ODCT&rsquo;s accuracy under ideal conditions.</p> </li> <li> <p><strong>Modified Datasets</strong>: These datasets introduce various types of errors to showcase ODCT&rsquo;s ability to handle diverse compliance scenarios:</p> <ol> <li><strong>Modified Dataset 1</strong>: Introduces type mismatches and spelling errors in key attributes.</li> <li><strong>Modified Dataset 2</strong>: Contains extraneous properties and missing required properties, including details about energy consumption and efficiency class.</li> <li><strong>Modified Dataset 3</strong>: Includes both extraneous and missing properties, and additional priority levels for energy profiles.</li> </ol> </li> <li> <p><strong>SHACL Shapes</strong>: The SHACL shapes used in the compliance testing for both SAREF and SAREF4ENER ontologies are included in this repository to allow reproducibility of the validation process.</p> </li> </ul> <ul> <li> <p><strong>Error Detection Results and Performance Reports</strong>: After conducting compliance tests using ODCT we got the Results and Performance Reports, the repository includes comprehensive reports detailing the results. These reports highlight the types of errors detected and provide a performance analysis of the tool under various scenarios.</p> </li> <li> <p><strong>Demonstration Video</strong>: A video is provided to guide users through the <strong>ODCT web application</strong>, showcasing how the tool detects errors and generates detailed compliance reports based on smart energy appliance datasets.</p> </li> </ul> <h2>Background</h2> <p>The integration of smart energy appliances into modern power grids is key to improving <strong>energy management</strong> and supporting <strong>sustainability goals</strong> like the <strong>European Green Deal</strong>. However, ensuring that these devices communicate effectively and conform to <strong>standardized protocols</strong> is a challenge. The <strong>ODCT</strong> tool addresses this challenge by providing a rigorous, ontology-based validation framework that is both <strong>protocol-agnostic</strong> and <strong>technology-flexible</strong>.</p> <p>This work is grounded in the broader context of <strong>global warming</strong> and the need for <strong>energy efficiency</strong> and <strong>demand-side flexibility</strong> in energy systems. By ensuring compliance with <strong>SAREF</strong> and <strong>SAREF4ENER</strong>, ODCT supports the EU&rsquo;s ambitions for <strong>carbon neutrality</strong> by 2050, contributing to a connected, efficient, and sustainable energy ecosystem.</p> <h2>Methodology</h2> <p>ODCT uses a structured methodology that involves:</p> <ol> <li><strong>Generating relevant datasets</strong> for validation.</li> <li><strong>Defining SHACL shape constraints</strong> based on ontologies.</li> <li><strong>Developing a user-friendly web application</strong> to facilitate compliance testing.</li> <li><strong>Performing compliance tests</strong> that validate datasets against SHACL shapes, ensuring interoperability and adherence to energy management standards.</li> </ol> <h2>Why It Matters</h2> <p>Researchers and developers working on smart energy appliances will benefit from ODCT by:</p> <ul> <li>Ensuring their devices meet standardized ontological requirements for <strong>interoperability</strong>.</li> <li>Reducing <strong>compliance issues</strong> in the development phase, leading to smoother integration into energy management systems.</li> <li>Supporting the <strong>sustainability efforts</strong> by enhancing device communication in <strong>smart grids</strong>.</li> </ul> <p>This repository showcases the potential of ODCT in fostering <strong>data accuracy</strong>, <strong>semantic interoperability</strong>, and <strong>compliance</strong> with essential energy standards. It offers comprehensive resources for furthering research and development in the field of smart energy appliances and energy management.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Dataset on UAV RGB videos acquired over a vineyard property of Bodegas Terras Gauda at an early stage of Botrytis cinerea infection in 2021

<p>The videos were collected in a vineyard owned by Bodegas Terras Gauda, in June 2021. The videos were collected with a DJI Matrice 210 RTK UAV, which had a DJI Zenmuse X5S sensor onboard. A total of 4 rows were recorded with side videos.&nbsp;The flights were carried out on a sunny day with wind velocity lower than 0.5 m/s. Annotations of the grape clusters in the MOTS style are provided.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Dataset of a 5G RTSP video streaming use case

<p><strong>About the project: monitoring 5G RTSP video streaming</strong></p> <p>This dataset collects data from a 5G video streaming use case. A video is streamed by a cvlc server (realized as a Kubernetes pod) through RTSP to a variable number of 5G UE clients that activate according to a daily traffic pattern. The values of the 4 dataset features (number of active UEs, gNB&#39;s downlink bit rate, pod&#39;s outbound traffic, and pod&#39;s CPU usage) are collected by a custom monitoring system deployed in the context of the MONB5G project.</p> <p><strong>Setup/Equipment</strong></p> <p>The Kubernetes cluster, including the server pod, runs in a COTS server. The 5G core and gNB is realized through Amarisoft Callbox Ultimate. The UEs are emulated through Amarisoft Simbox. In order to display the video in ffplay clients, we use the Remote UE from Amarisoft, so traffic from Simbox is forwarded to an external VM with GUI.</p> <p><strong>Video</strong></p> <p>The streamed video is Big Buck Bunny at 30 FPS from <a href="https://peach.blender.org/">https://peach.blender.org/</a>.</p> <table> <tbody> <tr> <td> <p>Video codec&nbsp;</p> </td> <td> <p>Advanced Video Codec (AVC)&nbsp;</p> </td> </tr> <tr> <td> <p>Width&nbsp;</p> </td> <td> <p>1920 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Height&nbsp;</p> </td> <td> <p>1080 pixels&nbsp;</p> </td> </tr> <tr> <td> <p>Display aspect radio&nbsp;</p> </td> <td> <p>16:9&nbsp;</p> </td> </tr> <tr> <td> <p>Duration&nbsp;</p> </td> <td> <p>10 min 34 s&nbsp;</p> </td> </tr> <tr> <td> <p>Max Bitrate&nbsp;</p> </td> <td> <p>16.7 Mb/s&nbsp;</p> </td> </tr> <tr> <td> <p>Frame rate&nbsp;</p> </td> <td> <p>30 FPS&nbsp;</p> </td> </tr> </tbody> </table> <p><strong>What does this Zenodo project contain?</strong></p> <ol> <li>The csv file of the dataset (dataset.csv)</li> <li>A picture displaying an overview of the setup (overview.png)</li> <li>A picture displaying Grafana charts for each featuer (grafana.png)</li> <li>A picture displaying a screenshot of the Remote UE VM with multiple UEs playing the video (ues.png)</li> </ol> <p><strong>Dataset</strong></p> <p>The dataset has 5 columns (time + 4 features). Features:</p> <ol> <li><em>Time</em>: timestamp in epoch format.</li> <li><em>Number of active UEs (N)</em>: number of UEs that are currently downloading more than 100 kbps. No unit.</li> <li><em>gNB&#39;s downlink bit rate (R)</em>: aggregate downlinkg bitrate from the gNB to all the UEs. In Mbps.</li> <li><em>Outbound traffic (O)</em>: outboun traffic at the pod&#39;s interface, transmitting the video(s) packets. In Mbps.</li> <li><em>CPU (C)</em>: CPU usage at the server pod. In millicores (mc). Each iteration represents a whole day, composed of 24 &quot;demand periods&quot;. Each demand period takes 2 minutes and is given by the number of active UEs consuming the video stream (N). N is included for informative reasons. Sampling rate is 10 seconds, but some parameters are refreshed at a lower frequency given monitoring limitations. This means that some parameters repeat the same value in consecutive measurements.</li> </ol>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Kuopio gait dataset: motion capture, inertial measurement and video-based sagittal-plane keypoint data from walking trials

<p>This dataset contains motion capture (3D marker trajectories, ground reaction forces and moments), inertial measurement unit (wearable Movella Xsens MTw Awinda sensors on the pelvis, both thighs, both shanks, and both feet), and sagittal-plane video (anatomical keypoints identified with the OpenPose human pose estimation algorithm) data.<br>The data is from 51 willing participants and collected in the HUMEA laboratory in the University of Eastern Finland, Kuopio, Finland, between 2022 and 2023. All trials were conducted barefoot.</p> <p>The file structure contains an Excel file containing information of the participants, data folders under each subject (numbered 01 to 51), and a MATLAB script.</p> <p>The Excel file has the following data for the participants:</p> <ul> <li><strong>ID</strong>: ID of the participants from 1 to 51</li> <li><strong>Age</strong>: age of the participant in years</li> <li><strong>Gender</strong>: biological sex as M for male, F for female</li> <li><strong>Leg</strong>: the participant's dominant leg, identified by asking which foot the participant would use to kick a football; R for right, L for left</li> <li><strong>Height</strong>: height of the participant in centimeters</li> <li><strong>Invalid_trials</strong>: list of invalid trials in the motion capture data (MOCAP) data, usually classified as such because the participant did not properly step on the middle force plate</li> <li><strong>IAD</strong>: inter-asis distance in millimeters, the distance between palpated left and right anterior superior iliac spine, measured with a caliper</li> <li><strong>Left_knee_width</strong>: width of the left knee from medial epicondyle to lateral epicondyle in millimeters, palpated and measured with a caliper</li> <li><strong>Right_knee_width</strong>: same as above for the right knee</li> <li><strong>Left_ankle width</strong>: width of the left ankle from medial malleolus to lateral malleolus in millimeters, palpated and measured with a caliper</li> <li><strong>Right_ankle_width</strong>: same as above for the right ankle</li> <li><strong>Left_thigh_length</strong>: the distance between the greater trochanter of the left femur and the lateral epicondyle of the left femur in millimeters, palpated and measured with a measuring tape</li> <li><strong>Right_thigh_length</strong>: same as above for the right thigh</li> <li><strong>Left_shank_length</strong>: the distance between the medial epicondyle of the femur and the medial malleolus of the tibia in millimeters, palpated and measured with a measuring tape</li> <li><strong>Right_shank_length</strong>: same as above for the right shank</li> <li><strong>Mass</strong>: mass in kilograms, measured on a force plate just before the walking measurements</li> <li><strong>ICD</strong>: inter-condylar distance of the knee of the dominant leg, measured from low-field MRI</li> <li><strong>Left_knee_width_mocap</strong>: distance between reflective MOCAP markers on the medial and lateral epicondyles of the knee in millimeters, measured from a static standing trial; -1 for missing (subject did not have those markers)</li> <li><strong>Right_knee_width_mocap</strong>: same as above for the right knee</li> </ul> <p>The folders under each subject (folders numbered 01 to 51) are as follows:</p> <ul> <li><strong>imu</strong>: "Raw" inertial measurement unit (IMU) data files that can be read with Xsens Device API (included in Xsens MT Manager 4.6, which may be unavailable these days, not sure). You won't need this if you use the data in the imu_extracted folder.</li> <li><strong>imu_extracted</strong>: IMU data extracted from those data files using the Xsens Device API, so you don't have to. <ul> <li>The data is saved as MATLAB structs where the fields are named as a sensor ID (e.g., "B42D48"). The sensor IDs and their corresponding IMU locations are as follows: <ul> <li>pelvis IMU: B42DA3</li> <li>right femur IMU: B42DA2</li> <li>left femur IMU: B42D4D</li> <li>right tibia IMU: B42DAE</li> <li>left tibia IMU: B42D53</li> <li>right foot IMU: B42D48</li> <li>left foot IMU: B42D51 (except for subjects 01 and 02, where left foot IMU has the ID B42D4E)</li> </ul> </li> <li>Some of the data are just zeros as they couldn't be read from these sensors, but under each sensor, the fields "calibratedAcceleration", "freeAcceleration", "time", "rotationMatrix", and "quaternion" contain usable data. <ul> <li>time: Contains time stamps of the measurement at each frame recorded at 100 Hz, so if you remove the first value from all values in the time vector and divide the result by 100, you will get the time in seconds from the beginning of the walking trial.</li> <li>calibratedAcceleration and freeAcceleration: Contain triaxial acceleration data from the accelerometers of the IMU. freeAcceleration is just calibratedAcceleration without the effect of Earth's gravitational acceleration.</li> <li>rotationMatrix: Orientations of the IMU as rotation matrices.</li> <li>quaternion: Orientations of the IMU as quaternions.</li> </ul> </li> </ul> </li> <li><strong>openpose</strong>: Trajectories of the keypoints identified from sagittal plane video frames, saved as json files. <ul> <li>The keypoints are from the BODY_25 model of OpenPose (https://cmu-perceptual-computing-lab.github.io/openpose/web/html/doc/md_doc_02_output.html).</li> <li>Each frame in the video has its own json file.</li> <li>You can use the function in the script "OpenPose_to_keypoint_table.m" in the root folder to read the keypoint trajectories and confidences of all frames in a walking trial into MATLAB tables. The function takes as argument the path to the folder containing the json files of the walking trial.</li> </ul> </li> <li>Note that some subjects (11, 14, 37, 49) do not have keypoint and IMU data.</li> </ul> <p>The folders under each subject are divided into three ZIP archives with 17 subjects each.</p> <p>The script "OpenPose_to_keypoint_table.m" is a MATLAB script for extracting keypoint trajectories and confidences from JSON files into tables in MATLAB.</p> <p><br><strong>Publication in Data in Brief</strong>: <a href="https://doi.org/10.1016/j.dib.2024.110841" target="_blank" rel="noopener">https://doi.org/10.1016/j.dib.2024.110841</a></p> <p><br><strong>Contact</strong>: Jere Lavikainen, jere.lavikainen@uef.fi</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

How do native and non-native speakers recognize emotions in the instructor's voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners [dataset]

<p>Dataset for the journal article&nbsp;<em>How do native and non-native speakers recognize emotions in the instructor&rsquo;s voice in educational videos? Exploring the first step of the cognitive-affective model of e-learning for international learners.</em></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Open Satellite Video Single Target Tracking Datasets (OpenSatSTTD)

<p>We collect the latest open-source datasets for satellite video single target tracking (SatSTT) and launch the OpenSatSTTD project to promote the sharing of the latest research datasets in&nbsp;the SatSTT field. Satellite videos in the OpenSatSTTD project&nbsp;are collected from different sensors and platforms, and four&nbsp;targets (i.e., vehicles, trains, airplanes and&nbsp;vessels) are annotated&nbsp;by oriented bounding boxes. Users can obtain all satellite videos in the OpenSatSTTD&nbsp;project from links in the files.</p> <p>Source:</p> <p>Zheng, Ying., Zhu, Q., Luo, J., Li, Z., Lin, Z., Huang, X., and Zhang L.:&nbsp;Single Target Tracking in High-Resolution Satellite Videos: A Comprehensive Review (1.0) [Data set]. Zenodo.&nbsp;https://doi.org/10.5281/zenodo.6780820, 2022.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

UAV-based monocular SLAM video datasets in vineyards with RTK ground truth

<p>The dataset provides a UAV-based monocular visual SLAM data, designed to evaluate the potential of using monocular visual SLAM in vineyards. It includes videos in ".mp4" format collected by UAV, and&nbsp; "xlsx" tables which include latitude, longitude, height, speed in x, y and z, comjpass, pitch, roll. The ".xlsx" tables were measured by RTK and can be used as ground truth of UAV trajectory and pose.</p> <p>This dataset can be combined with other datasets to enable a comprehensive view of the vineyards:</p> <p>V&eacute;lez S, Ariza-Sent&iacute;s M, Valente J. EscaYard: Precision viticulture multimodal dataset of vineyards affected by Esca disease consisting of geotagged smartphone images, phytosanitary status, UAV 3D point clouds and Orthomosaics. Data in Brief. 2024 Jun 1;54:110497.&nbsp;<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dib.2024.110497" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.dib.2024.110497</span></a></p> <p><span>Ariza-Sent&iacute;s M, Wang K, Cao Z, V&eacute;lez S, Valente J. GrapeMOTS: UAV vineyard dataset with MOTS grape bunch annotations recorded from multiple perspectives for enhanced object detection and tracking. Data in Brief. 2024 Jun 1;54:110432. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dib.2024.110432" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.dib.2024.110432</a></span></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

MHMisinfo - Video-based Mental Health Misinformation Dataset

<p>MHMisinfo-Gold and MHMisinfo-Large datasets, as described in the paper "Supporters and Skeptics: LLM-based Analysis of Engagement with Mental Health (Mis)Information Content on Video-sharing Platforms" (forthcoming at ICWSM 2025). Videos and comments for each dataset are seperately stored in different .csv files</p> <p><strong>Dataset schema, videos</strong></p> <table> <tbody> <tr> <th><strong>Column Name</strong></th> <th><strong>Description</strong></th> </tr> <tr> <td><strong>video_id</strong></td> <td>ID of the Video, as assigned by their respective platforms</td> </tr> <tr> <td><strong>video_title</strong></td> <td>The title of the video</td> </tr> <tr> <td><strong>video_description</strong></td> <td>The description of the video, given by the video creators</td> </tr> <tr> <td><strong>audio_transcript</strong></td> <td>Text transcription of the video's audio track, as generated by Whisper speech-to-text model</td> </tr> <tr> <td><strong>video_view_count</strong></td> <td>View count of the video, at the time of data collection</td> </tr> <tr> <td><strong>video_like_count</strong></td> <td>Like count of the video, at the time of data collection</td> </tr> <tr> <td><strong>video_comment_count</strong></td> <td>Comment count of the video, at the time of data collection</td> </tr> <tr> <td><strong>label_ioi</strong></td> <td>"Information of Interventions" label of video, annotated by experts. 1 = High-quality information on interventions, -1 = Low-quality information on interventions</td> </tr> <tr> <td><strong>label_ebt</strong></td> <td>"Evidence-based Treatment" label of video, annotated by experts. 1 = Encourages evidence-based treatment, -1 = Discourages evidence-based treatment</td> </tr> <tr> <td><strong>label_aoc</strong></td> <td>"Alignment of Consensus" label of video, annotated by experts, 1 = High Alignment with Consensus, -1 = Low Alignment with Consensus</td> </tr> <tr> <td><strong>label</strong></td> <td>Overall mental health misinformation label of the video. 0 = non-MHMisinfo videos, and -1 = MHMisinfo videos</td> </tr> <tr> <td><strong>platform</strong></td> <td>Platform of the video</td> </tr> </tbody> </table> <p><strong>Dataset schema, comments</strong></p> <table> <tbody> <tr> <th><strong>Column Name</strong></th> <th><strong>Description</strong></th> </tr> <tr> <td><strong>text</strong></td> <td>The raw text of the comment</td> </tr> <tr> <td> <p><strong>commenter_channel_display_name</strong></p> </td> <td>The display name of the user who posted the comment.</td> </tr> <tr> <td> <p><strong>comment_publish_date</strong></p> </td> <td>The time when the comment was orignally published, .</td> </tr> <tr> <td> <p><strong>video_id</strong></p> </td> <td>ID of the Video associated by the platform, as assigned by their respective platforms</td> </tr> <tr> <td> <p><strong>platform</strong></p> </td> <td>Platform of the video associated with the comment</td> </tr> <tr> <td> <p><strong>label</strong></p> </td> <td>Overall mental health misinformation label of the video associated with the comment. 0 = non-MHMisinfo videos, and -1 = MHMisinfo videos</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Deeptangle Dataset: Labelled Experimental and Synthetic Videos of Swimming and Overlapping C. elegans worms

<p>This repository contains the dataset employed in the paper <a href="https://arxiv.org/abs/2301.04460">Fast spline detection in high density microscopy data</a>.</p> <p>Three files are provided:</p> <p>1. <em>videos.zip</em>: raw experimental videos.<br> 2. <em>labeled_data.zip</em>: labelled sections of experimental videos used for evaluation<br> 3. <em>syntehthic_dataset.zip</em>: synthetic dataset used for training</p> <p><strong>Labelled data</strong></p> <p>Sections are named as VIDEONAME_FRAMENUMBER_SECTIONID.<br> All labels are in labels.json and correspond to the middle frame of the clip (05.png).<br> Example plotting script is provided (<em>plot_data.py</em>).</p> <p><strong>Synthetic data</strong></p> <p>Clips are named as NUMBEROFWORMS_ID.<br> Labels (for all frames) are stored in labels.npy.<br> Example plotting script is provided (plot_data.py).</p> <p>&nbsp;</p> <p><strong>Related</strong></p> <p>Paper:&nbsp;<a href="https://arxiv.org/abs/2301.04460">https://arxiv.org/abs/2301.04460</a></p> <p>Deeptangle code:&nbsp;<a href="https://github.com/kirkegaardlab/deeptangle">https://github.com/kirkegaardlab/deeptangle</a></p> <p>Labelling tool:&nbsp;<a href="https://github.com/kirkegaardlab/deeptanglelabel">https://github.com/kirkegaardlab/deeptanglelabel</a></p> <p>---</p> <p>If used, please cite</p> <p><em>Albert Alonso &amp; Julius B. Kirkegaard.&nbsp;Fast spline detection in high density microscopy data. 2023.</em></p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Interlinking dataset of the video game databases Mobygames and MediaArt Database

<p>The goal of the DFG (German Research Foundation) funded project diggr (Databased Infrastructure for Global Games Culture Research) is the integration of information from various heterogeneous, fragmentary online resources of the video game domain with the aim of creating robust research datasets.</p> <p>This dataset provides automatically generated links between the video game databases Mobygames (http://www.mobygames.com) and Media Art DB (https://mediaarts-db.bunka.go.jp/gm/) and additional metadata for each game entry.</p> <p>A full description of the dataset is provided in <em>Dataset Description.md.</em></p>

opencc-zeroJul 2019View details →
zenodo40/100

anTraX: high throughput video tracking of color-tagged insects (benchmark datasets)

<p>Datasets used to benchmark anTraX tracking software. Each dataset contains the raw videos, a configured anTraX session with all parameters required to reproduce the tracking results from the paper, as well as the tracking output for the first video in each dataset.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Long document similarity datasets, Wikipedia excerptions for movies, video games and wine collections

<p>Three&nbsp;corpora in different domains extracted from Wikipedia.</p> <p>For all datasets, the figures and tables have been filtered out, as well as the categories and &quot;see also&quot; sections.</p> <p>The article structure, and&nbsp;particularly the sub-titles and paragraphs are kept in these datasets</p> <p>&nbsp;</p> <p><strong>Wines</strong></p> <p>Wikipedia wines dataset consists of 1635 articles from the wine domain. The extracted dataset consists of a non-trivial mixture of articles, including different wine categories, brands, wineries, grape types, and more. The ground-truth recommendations were crafted by a human sommelier, which annotated 92 source articles with ~10 ground-truth recommendations for each sample. Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Dom P&eacute;rignon - Mo&euml;t &amp; Chandon</li> <li>Pinot Meunier - Chardonnay</li> </ul> <p><strong>Movies</strong></p> <p>The Wikipedia movies dataset consists of 100385 articles describing different movies. The movies&#39; articles may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.<br> For this dataset, we have extracted a test set of ground truth annotations for 50 source articles using the &quot;<a href="https://bestsimilar.com/">BestSimilar</a>&quot;&nbsp;database. Each source articles is associated with a list of ${\scriptsize \sim}12$ most similar movies.<br> Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Schindler&#39;s List - The Pianist</li> <li>Lion King - The Jungle Book</li> </ul> <p><strong>Video games</strong></p> <p>The Wikipedia video games dataset consists of 21,935 articles reviewing video games from all genres and consoles. Each article may consist of a different combination of sections, including summary, gameplay, plot, production, etc. Examples for ground-truth expert-based recommendations are:</p> <ul> <li>Grand Theft Auto - Mafia</li> <li>Burnout Paradise - Forza Horizon 3</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo40/100

An Open Access User Generated Video Dataset from 2016 Edinburgh Festival

<p>A user generated video dataset captured during the 2016 Edinburgh festival. The provided dataset was collected using a smart phone and is available with no post-processing. The videos mainly cover the Edinburgh streets and the festival atmosphere, and do not cover any performances. The dataset can be used for evaluation of various research tools, such as video quality assessment and enhancement.</p>

opencc-by-nc-nd-4.0Aug 2017View details →
zenodo40/100

Datasets, codes and video clips for the laboratory flume tests of granular flow

<p>Datasets, video clips, and codes&nbsp;related to the paper &#39;Insight into granular flow dynamics relying on basal stress measurements: from experimental flume tests&#39;, submitted to the Journal of Geophysical Research: Solid Earth.</p> <p>The datasets provides the raw and processed data for the laboratory flume tests of granular flow including parameters reflecting the granular flow behavior, basal normal stresses measured by a force plate, and deposit parameters of the&nbsp; granular flows.</p> <p><strong>S1_data_granular flow_velocity</strong> provides data of the velocity profiles with a 0.1 second time interval, the depth-averaged velocities, the depth-averaged shear rates,&nbsp;&nbsp;and the solid inertial stresses of the granular flows under different experimental conditions. The original data were calculated through particle image velocimetry (PIV) method. The images for PIV analysis were recorded by a high-speed camera.</p> <p><strong>S2_data_granular flow_stress</strong> provides&nbsp;the raw data of the measured basal normal stresses of the granular flows for all tests. The mean and fluctuating stress components extracted by applying a moving window average filter are also listed in the Table.</p> <p><strong>S3_data_granular flow_flow depth</strong> provides the data of the granular flow depth extracted every 0,02 s through a image processing method based on the high-speed photographs.&nbsp;</p> <p><strong>S4_data_granular flow_deposit</strong> provides the&nbsp;parameters of the granular flow deposits for all tests including the apparent friction coefficient and equivalent friction coefficient.&nbsp;The deposit parameters were calculated based on the digital surface model (DSM) of deposit, which were obtained through a oblique photogrammetry method.</p> <p><strong>S5_data_granular flow_density</strong> gives the data of the dynamic bulk flow densities of the granular flows under all experimental conditions. The dynamic bulk densities were calculated according to the measured and calculated normal stresses.</p> <p>The videos of the granular flows under different experimental conditions during their propagation are provided in <strong>&#39;S6_video_granular flow.zip&#39;</strong> to show the granular flow behavior and its evolution.&nbsp;<strong>S6_video_granular flow</strong>&nbsp;includes the side-view of&nbsp;the granular flows under all experimental conditions and front-view of the IMF-223 granular flow .&nbsp;</p> <p><strong>S7_codes_data analysis</strong> provides the computer codes for the extraction of mean and fluctuating components and the calculation of granular flow depth. The former includes one file for conducting moving average filter. The latter contains four files, which are used for median filter, image erosion, threshold segmentation and floodfill, extracting flow depth.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Hand Washing Video Dataset Annotated According to the World Health Organization's Handwashing Guidelines - METC Subset

<p><strong>Overview:</strong> This is a lab-based dataset with videos recording volunteers (medical students) washing their hands as part of a hand-washing monitoring and feedback experiment. The dataset is collected in the Medical Education Technology Center (METC) of Riga Stradins University, Riga, Latvia. In total, 72 participants took part in the experiments, each washing their hands three times, in a randomized order, going through three different hand-washing feedback approaches (user interfaces of a mobile app). The data was annotated in real time by a human operator, in order to give the experiment participants real-time feedback on their performance. There are 212 hand washing episodes in total, each of which is annotated by a single person. The annotations classify the washing movements according to the World Health Organization&#39;s (WHO) guidelines by marking each frame in each video with a certain movement code.</p> <p>This dataset is part on three dataset series all following the same format:</p> <ul> <li><a href="https://zenodo.org/record/4537209">https://zenodo.org/record/4537209 </a>- data collected in Pauls Stradins Clinical University Hospital</li> <li><a href="https://zenodo.org/record/5808764">https://zenodo.org/record/5808764</a> - data collected in Jurmala Hospital</li> <li><a href="https://zenodo.org/record/5808789">https://zenodo.org/record/5808789</a> - data collected in the&nbsp;Medical Education Technology Center (METC) of Riga Stradins University</li> </ul> <p><strong>Note #1:</strong> we recommend that when using this dataset for machine learning, allowances are made for the reaction speed of the human operator labeling the data. For example, the annotations can be expected to be incorrect a short while after the person in the video switches their washing movements.</p> <p><strong>Application: </strong>The intention of this dataset is to serve as a basis for training machine learning classifiers for automated hand washing movement recognition and quality control.</p> <p><strong>Statistics:</strong></p> <ul> <li>Frame rate: ~16 FPS (slightly variable, as the video are reconstructed from a sequence of jpg images taken with max framerate supported by the capturing devices).</li> <li>Resolution: 640x480</li> <li>Number of videos: 212</li> <li>Number of annotation files: 212</li> </ul> <p>Movement codes (in JSON files):</p> <ul> <li>1: Hand washing movement &mdash; Palm to palm</li> <li>2: Hand washing movement &mdash; Palm over dorsum, fingers interlaced</li> <li>3: Hand washing movement&nbsp;&mdash; Palm to palm, fingers interlaced</li> <li>4: Hand washing movement &mdash; Backs of fingers to opposing palm, fingers interlocked</li> <li>5: Hand washing movement&nbsp;&mdash; Rotational rubbing of the thumb</li> <li>6: Hand washing movement &mdash; Fingertips to palm</li> <li>0: Other hand washing movement</li> </ul> <p><strong>Note #2: </strong>The original dataset of JPG images is available upon request. There are 13 annotation classes in the original dataset: for each of the six washing movements defined by the WHO, &quot;correct&quot; and &quot;incorrect&quot; execution is market with two different labels. In this published dataset, all incorrect executions are marked with code 0, as &quot;other&quot; washing movement.</p> <p><strong>Acknowledgments: </strong>The dataset collection was funded by the Latvian Council of Science project: &quot;Automated hand washing quality control and quality evaluation system with real-time feedback&quot;, No: lzp - Nr. 2020/2-0309.</p> <p><strong>References: </strong>For more detailed information, see this article, describing a similar dataset collected in a different project:</p> <ul> <li> <p>M. Lulla, A. Rutkovskis, A. Slavinska, A. Vilde, A. Gromova, M. Ivanovs, A. Skadins, R. Kadikis, A. Elsts. <em>Hand-Washing Video Dataset Annotated According to the World Health Organization&rsquo;s Hand-Washing Guidelines</em>. Data. 2021; 6(4):38. <a href="https://doi.org/10.3390/data6040038">https://doi.org/10.3390/data6040038</a></p> </li> </ul> <p><strong>Contact information: </strong>atis.elsts@edi.lv</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Hand Washing Video Dataset Annotated According to the World Health Organization's Handwashing Guidelines - Jurmala Hospital Subset

<p><strong>Overview:</strong> This is a large-scale real-world dataset with videos recording medical staff washing their hands as part of their normal job duties in the Jurmala Hospital located in Jurmala, Latvia. There are 2427 hand washing episodes in total, almost all of which are annotated by two persons. The annotations classify the washing movements according to the World Health Organization&#39;s (WHO) guidelines by marking each frame in each video with a certain movement code.</p> <p>This dataset is part on three dataset series all following the same format:</p> <ul> <li><a href="https://zenodo.org/record/4537209">https://zenodo.org/record/4537209</a> - data collected in Pauls Stradins Clinical University Hospital</li> <li><a href="https://zenodo.org/record/5808764">https://zenodo.org/record/5808764</a> - data collected in Jurmala Hospital</li> <li><a href="https://zenodo.org/record/5808789">https://zenodo.org/record/5808789</a> - data collected in the&nbsp;Medical Education Technology Center (METC) of Riga Stradins University</li> </ul> <p><strong>Applications: </strong>The intention of this dataset is twofold: to serve as a basis for training machine learning classifiers for automated hand washing movement recognition and quality control, and to allow to investigate the real-world quality of washing performed by working medical staff.</p> <p><strong>Statistics:</strong></p> <ul> <li>Frame rate: 30 FPS</li> <li>Resolution: 320x240 and 640x480</li> <li>Number of videos: 2427</li> <li>Number of annotation files: 4818</li> </ul> <p>Movement codes (both in CSV and JSON files):</p> <ul> <li>1: Hand washing movement &mdash; Palm to palm</li> <li>2: Hand washing movement &mdash; Palm over dorsum, fingers interlaced</li> <li>3: Hand washing movement&nbsp;&mdash; Palm to palm, fingers interlaced</li> <li>4: Hand washing movement &mdash; Backs of fingers to opposing palm, fingers interlocked</li> <li>5: Hand washing movement&nbsp;&mdash; Rotational rubbing of the thumb</li> <li>6: Hand washing movement &mdash; Fingertips to palm</li> <li>7: Turning off the faucet with a paper towel</li> <li>0: Other hand washing movement</li> </ul> <p><strong>Acknowledgments: </strong>The dataset collection was funded by the Latvian Council of Science project: &quot;Automated hand washing quality control and quality evaluation system with real-time feedback&quot;, No: lzp - Nr. 2020/2-0309.</p> <p><strong>References: </strong>For more detailed information, see this article, describing a similar dataset collected in a different project:</p> <ul> <li> <p>M. Lulla, A. Rutkovskis, A. Slavinska, A. Vilde, A. Gromova, M. Ivanovs, A. Skadins, R. Kadikis, A. Elsts. <em>Hand-Washing Video Dataset Annotated According to the World Health Organization&rsquo;s Hand-Washing Guidelines</em>. Data. 2021; 6(4):38. <a href="https://doi.org/10.3390/data6040038">https://doi.org/10.3390/data6040038</a></p> </li> </ul> <p><strong>Contact information: </strong>atis.elsts@edi.lv</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record