Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
118
datasets available to search
ShareScore release 0.9.0
Dataset results
118 results for “Anomaly detection”
Detecting anomalies in melt-extruded 3D printed parts using in situ data
<p>The data in this repository was gathered from a study to collect real-time, in situ data from polymer melt extrusion (ME) 3D printing, using a set of sensors to non-destructively identfy printed parts that contain defects. The data underwent variance analysis to determine an "acceptable" range of filament diameters and non-destructivley identify spatial regions of printed cylinders in multi-part builds that contain defects.</p> <p>The data consists of two folders and a log meant to track procedural adherence for each cylinder printed, the introduced defects, or lack thereof, and the pressurization of the part. The "Final Build Logs" spreadsheet contains information regarding the two locations of the deformations along the 56 meters of filament needed to have no more than three anomalous cylinders out of the six printed cylinders, the date of the applied deformations to the filament, the initials of the researcher applying the deformations, the date that the build was printed along with the initials of the researcher who printed it, the part number, researcher initials, and date of the pressurization test for each cylinder within the build, and a comment describing any deviations from the procedure that play into the random error of the statistical analysis for each cylinder. </p> <p>The "Pressure Test Data" folder contains a folder for each build. Within these folders are .tdms files containing metadata on the measurement system in the header and tab-delimeted values for the columns. The columns of interest to the study are X_Value, representing time elapsed, and pressure, which we evaluated on the values' exponential decay rate. The files also contain supplemental information such as a column for temperature (celsius), and the flow rate (SLPM Normalized). The "Build Data" folder contains in situ data from the sensor-equipped printer in a .csv file, the STL file for the build, the gcode file from the applied slicer settings, the AMRP file stores printer settings, and a .pdf file for the setup specifications.</p>
Visualizing histopathologic deep learning classification and anomaly detection using nonlinear feature space dimensionality reduction
<p>Representative Testing/Validation WSIs used in the manuscript "Visualizing histopathologic deep learning classification and anomaly detection using nonlinear feature space dimensionality reduction"</p>
Visualizing histopathologic deep learning classification and anomaly detection using nonlinear feature space dimensionality reduction
<p>Training image dataset used in the manuscript "Visualizing histopathologic deep learning classification and anomaly detection using nonlinear feature space dimensionality reduction"</p>
FRGADB - FIRST Radio Galaxy Anomaly Detection Benchmark
<p>This dataset is a combination of samples from the MiraBest, FRGMRC and LRG catalogues. It is intended to serve as a benchmark for models' performance with respect to anomalous source detection in radio astronomy.</p>
MACHINE LEARNING ALGORITHMS FOR ANOMALY DETECTION IN PUBLIC DATA USING GITHUB AS AN EXAMPLE
<p>This study explores the application of machine learning algorithms for detecting anomalies in GitHub data to enhance the evaluation of technological projects. The research aims to develop a robust methodology for identifying data anomalies, such as artificial activity spikes, that can distort project assessments. Methods such as Isolation Forest, One-Class SVM, and advanced deep learning techniques like autoencoders and GANs are employed to analyze and identify irregular patterns in GitHub repositories. The findings demonstrate that these algorithms effectively detect both obvious and subtle anomalies, offering reliable insights into project authenticity. The proposed conceptual model integrates these methods into a scalable system, enhancing transparency and accuracy in technological project evaluation. The novelty of this work lies in its comprehensive approach to analyzing GitHub data, combining traditional and deep learning techniques to improve the reliability of assessments, making it a significant contribution to the field.</p>
M100 dataset: time-aggregated data for anomaly detection
<p>This entry is a part of a larger data set collected from the most recent Tier-0 supercomputer hosted at CINECA (Marconi100, <a href="https://www.hpc.cineca.it/hardware/marconi100">https://www.hpc.cineca.it/hardware/marconi100</a>). The data covers the entirety of the system, ranging from the computing nodes (980+ computing nodes) internal information such as core loads, temperatures, frequencies, memory write/read operations, CPU power consumption, fan speed, GPU usage details, etc., to the system-wide information, including the liquid cooling infrastructure, the air conditioning system, the power supply units, workload manager statistics, and job-related information, system status alerts, and weather forecast. <br> It comprises hundreds of metrics measured on each computing node, in addition to hundreds of other metrics gathered from sensors monitored along all system components.</p> <p>This particular dataset is made for anomaly detection purposes, it contains the same data as the main dataset but aggregated over time, with one Parquet file for each node. The data is distributed in tarballs, each one including all the files relative to the nodes contained in a given rack. For each file, the rows represent periods of 15 minutes, with the columns being aggregated values (average, standard deviation, min, max) over all the IPMI metrics that are available for the node; an additional column contains anomaly labels from Nagios.</p> <p>More details can be found in the companion repository: <a href="https://gitlab.com/ecs-lab/exadata">https://gitlab.com/ecs-lab/exadata</a>, including the spatial distribution of the nodes in the room.</p>
Dataset Artifact for Prodigy: Towards Unsupervised Anomaly Detection in Production HPC Systems
<p>The dataset contains a small set of application runs from Eclipse supercomputer. The applications run with and without synthetic HPC performance anomalies. More detailed information regarding synthetic anomalies can be found at: https://github.com/peaclab/HPAS.</p> <p>We have chosen four applications, namely LAMMPS, sw4, sw4Lite, and ExaMiniMD, to encompass both real and proxy applications. We have executed each application five times on four compute nodes without introducing any anomalies. To showcase our experiment, we have specifically selected the "memleak" anomaly as it is one of the most commonly occurring types. Additionally, we have also executed each application five times with the chosen anomaly. The dataset we have collected consists of a total of 160 samples, with 80 samples labeled as anomalous and 80 samples labeled as healthy. For the details of applications please refer to the paper.</p> <p>The applications were run on Eclipse, which is situated at Sandia National Laboratories. Eclipse comprises 1488 compute nodes, each equipped with 128GB of memory and two sockets. Each socket contains 18 E5-2695 v4 CPU cores with 2-way hyperthreading, providing substantial computational power for scientific and engineering applications.</p>
Detecting anomalies in melt-extruded 3D printed parts using in situ data
Open the record for dataset details and reuse information.
Data for "Wave anomaly detection in wave buoy measurements" - Phase-Resolving Time Series
<p>The datasets contain extreme time series obtained from the post-processed 3D wave fields simulated using HOS-Ocean, a high-order spectral model (HOSM) that solves the deterministic propagation of nonlinear wave fields in deep water (Ducrozet et al., 2016).</p> <p>Voermans. (2020). Data for "Wave anomaly detection in wave buoy measurements" - Phase-Resolving Time Series [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4028014</p> <p> </p>
Dataset for Anomaly Detection Using Inter-Arrival Curves for Real-time Systems
<p>The dataset shows the input files and detailed results for the experiments discussed in the paper. A README file provide more details on the data.</p>
EIRSAT-1 Test Campaign and Flight Dataset for Anomaly Detection
<p><span>We have curated a unique dataset derived from EIRSAT-1, Ireland's inaugural domestically produced satellite, as a testing and validation resource for these ML models and the future development cycle of AI-enabled small satellites. This dataset consists of a training set developed during ground testing and containing artificial anomalies induced to train satellite operators, a validation dataset containing real anomalies encountered during the qualification campaign, and an early flight test dataset collected since the satellite was launched on December 1<sup>st</sup>, 2023. This paper presents an in-depth analysis of the efficacy of these ML techniques when applied to the EIRSAT-1 dataset, offering insights into their potential to revolutionize the domain of satellite operations through enhanced autonomy and responsiveness. This study not only showcases the capabilities of these ML techniques in an operational environment but also sets the stage for future research and development in autonomous satellite systems.</span></p>
ComplexVAD Video Anomaly Detection Dataset
<p><strong>Introduction</strong></p> <p>The ComplexVAD dataset consists of 104 training and 113 testing video sequences taken from a static camera looking at a scene of a two-lane street with sidewalks on either side of the street and another sidewalk going across the street at a crosswalk. The videos were collected over a period of a few months on the campus of the University of South Florida using a camcorder with 1920 x 1080 pixel resolution. Videos were collected at various times during the day and on each day of the week. Videos vary in duration with most being about 12 minutes long. The total duration of all training and testing videos is a little over 34 hours. The scene includes cars, buses and golf carts driving in two directions on the street, pedestrians walking and jogging on the sidewalks and crossing the street, people on scooters, skateboards and bicycles on the street and sidewalks, and cars moving in the parking lot in the background. Branches of a tree also move at the top of many frames.</p> <p>The 113 testing videos have a total of 118 anomalous events consisting of 40 different anomaly types.</p> <p>Ground truth annotations are provided for each testing video in the form of bounding boxes around each anomalous event in each frame. Each bounding box is also labeled with a track number, meaning each anomalous event is labeled as a track of bounding boxes. A single frame can have more than one anomaly labeled.</p> <p><strong>At a Glance</strong></p> <ul> <li>The size of the unzipped dataset is ~39GB</li> <li>The dataset consists of Train sequences (containing only videos with normal activity), Test sequences (containing some anomalous activity), a ground truth annotation file for each Test sequence, and a README.md file describing the data organization and ground truth annotation format.</li> <li>The zip files contain a Train directory, a Test directory, an annotations directory, and a README.md file.</li> </ul> <p><strong>License</strong></p> <p>The ComplexVAD dataset is released under <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>.</p> <p>All data:</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2024 SPDX-License-Identifier: CC-BY-SA-4.0</code></pre>
Long-Tailed Anomaly Detection (LTAD) Dataset
<p><strong>Introduction</strong></p> <p>Anomaly detection (AD) aims to identify defective images and localize their defects (if any). Ideally, AD models should be able to: detect defects over many image classes; not rely on hard-coded class names that can be uninformative or inconsistent across datasets; learn without anomaly supervision; and be robust to the long-tailed distributions of real-world applications. To address these challenges, we formulate the problem of long-tailed AD by introducing several datasets with different levels of class imbalance for performance evaluation.</p> <p>To encourage more follow up works on long-tailed AD, we are publicly releasing the dataset split used in our paper (“Long-Tailed Anomaly Detection with Learnable Class Names” by Chih-Hui Ho, Kuan-Chuan Peng, and Nuno Vasconcelos, CVPR 2024).</p> <p>Files in the unzipped folder:</p> <p>1. ./README.md: This Markdown file</p> <p>2. ./dataset_split: Folder contains long-tail splits from three datasets. See below for details.</p> <p><strong> </strong></p> <p><strong>At a Glance</strong></p> <ul> <li>The size of the unzipped dataset is ~16MB</li> <li>Three datasets are used in this project, including [MVTec](https://www.mvtec.com/company/research/datasets/mvtec-ad), [VisA](https://github.com/amazon-science/spot-diff) and [DAGM](https://www.kaggle.com/datasets/mhskjelvareid/dagm-2007-competition-dataset-optical-inspection). Please download the datasets from their original repositories.</li> <li>The dataset split provided in this folder is organized as follows:<br>```<br>dataset_split<br>|---dagm_lt<br>|---mvtec_lt<br>|---visa_lt<br>|-----|-- exp<br>|-----|-----|----- 100<br>|-----|-----|----- |-----test.json<br>|-----|-----|----- |-----train.json<br>|-----|-----|----- 200<br>|-----|-- step<br>|-----|-- ...<br>```</li> <li>Each long-tailed dataset split contains a subfolder ``imbalance_type/imbalance_factor", where imbalance type can be [exponential (exp), step, reverse exponential (exp_reverse), reverse step (step_reverse)]. The definition of imbalance type and imbalance factor can be found in our paper. Each subfolder contains two json files, one for training and the other for testing.</li> <li>Each entry in the json file contains the meta information of an image and is similar to<br>```<br>{"filename": "candle/test/bad/000.JPG", "label": 1, "label_name": "defective", "clsname": "candle", "maskname": "candle/ground_truth/bad/000.png"}<br>```<br>- filename: location of the input image in the dataset<br>- label: indicates whether the input image is normal (labeled as 0) or defective (labeled as 1)<br>- label name: can be "good" or "defective"<br>- clsname: class name of the input image<br>- maskname (optional): location of the binary image that indicates the defect region. This is only available for test.json, because there is no defect image during training.</li> </ul> <p><strong>Citation</strong></p> <p>If you use the LTAD dataset in your research, please cite our contribution:</p> <pre><code>@InProceedings{Ho_2024_CVPR, author = {Ho, Chih-Hui and Peng, Kuan-Chuan and Vasconcelos, Nuno}, title = {Long-Tailed Anomaly Detection with Learnable Class Names}, booktitle = {The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, month = {June}, year = {2024} } </code></pre> <p><strong>License</strong></p> <p>The LTAD dataset is released under <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>. For the images in the MVTec, VisA, and DAGM datasets, please refer to their websites for their copyright and license terms.</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2023-2024 SPDX-License-Identifier: CC-BY-SA-4.0</code></pre>
Street Scene Video Anomaly Detection Dataset
<p><strong><span>Introduction</span></strong></p> <p><span>The Street Scene dataset consists of 46 training video sequences and 35 testing video sequences taken from a static USB camera looking down on a scene of a two-lane street with bike lanes and pedestrian sidewalks.<span> </span>Videos were collected from the camera at various times during two consecutive summers.<span> </span>All of the videos were taken during the daytime.<span> </span>The dataset is challenging because of the variety of activities taking place such as cars driving, turning, stopping and parking; pedestrians walking, jogging and pushing strollers; and bikers riding in bike lanes. In addition, the videos contain changing shadows, and moving background such as a flag and trees blowing in the wind.</span></p> <p><span>There are a total of 202,545 color video frames (56,135 for training and 146,410 for testing) each of size 1280 x 720 pixels. The frames were extracted from the original videos at 15 frames per second.</span></p> <p><span>The 35 testing sequences have a total of 205 anomalous events consisting of 17 different anomaly types. A complete list of anomaly types and the number of each in the test set can be found in our paper.</span></p> <p><span>Ground truth annotations are provided for each testing video in the form of bounding boxes around each anomalous event in each frame. Each bounding box is also labeled with a track number, meaning each anomalous event is labeled as a track of bounding boxes. Track lengths vary from tens of frames to 5200 which is the length of the longest testing sequence. A single frame can have more than one anomaly labeled.</span></p> <p><span>NOTE: This version of the dataset differs slightly with the original made available in 2020.<span> </span>Some anomalies were found in a few of the normal training sequences.<span> </span>These training frames were deleted from the dataset.<span> </span>Specifically, the following frames were removed:</span></p> <p><span>Train026: frames 1-184 (car taking a u-turn)</span></p> <p><span>Train027: frames 1-229 (jay walkers)</span></p> <p><span>Train031: frames 1-299 (jay walkers, illegally parked car)</span></p> <p><strong><span>At a Glance</span></strong></p> <ul> <li><span>The size of the unzipped dataset is ~46GB</span></li> <li><span>The dataset consists of Train sequences (containing only videos with normal activity), Test sequences (containing some anomalous activity) along with ground truth annotations, and a README.md file describing the data organization and ground truth annotation format.</span></li> <li><span>The zip file contains a Train directory, a Test directory and a README.md file.</span></li> </ul> <p><strong><span>Other Resources</span></strong></p> <p><span>None</span></p> <p><strong><span>Citation</span></strong></p> <p><span>If you use the Street Scene dataset in your research, please cite our contribution:</span></p> <pre><code>@inproceedings{ramachandra2020street, title={Street Scene: A new dataset and evaluation protocol for video anomaly detection}, author={Ramachandra, Bharathkumar and Jones, Michael}, booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision}, pages={2569--2578}, year={2020} } </code></pre> <p><strong><span>License</span></strong></p> <p><span>The Street Scene dataset is released under </span><a href="https://creativecommons.org/licenses/by-sa/4.0/"><span>CC-BY-SA-4.0 license</span></a><span>.</span></p> <p><span>All data:</span></p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2023 SPDX-License-Identifier: CC-BY-SA-4.0 </code></pre>
Datasets for Simulation-based Anomaly Detection for Multileptons at the LHC
<p>The simulated background and signal data used for a signal model agnostic machine learning search. This search examined the decay of the Higgs boson to leptons working off of LHC data from the Atlas experiment. Details are provided in the paper entitled "Simulation-based Anomaly Detection for Multileptons at the LHC". </p>
Dataset for Sound-based Anomalies Detection in Agricultural Robotics Application
<p>This data set contains data related to a Mowing Intelligent Tool (MowIT).</p> <p>Two different microphones were used to collect the sound samples, recording the audio with just one single channel, with a sampling rate of 44100 Hz and 16 bits resolution.</p> <p>The data provided by an inertial measurement unit (IMU) was also recorded since that was already integrated into the MowIT.</p> <p>Two different data collections were performed in different open-air environments with grass to cut.</p> <p>In each collection, eight different sample sets were made, five with the machine cutting using a trimmer line and the other three using the blades. Various combinations were used in each set, and tools were or were not placed on each of the three cutting axes of the MowIT. For each group, the acquisitions were designated from 0 to 7.</p> <p>Each folder of the first collection is a combination containing two audio files, one for each microphone used, the IMU data and a photograph of the lower part of the MowIT to understand the configuration used.</p> <p>In the second collection, to improve the variety of data, three distinct sub-sets were performed for combination: the first with the MowIT turned on but not cutting grass and the next two cutting grass. </p> <p>In samples 4 and 7, there is one audio where the MowIT cuts but stops due to motor stress. In sample 6, the initial recording was not made without cutting grass, and only the two recordings were made cutting grass.</p> <p> </p> <p> </p> <p> </p>
Dataset for Non-resonant Anomaly Detection with Background Extrapolation
<p>These are the datasets used in the journal version of the Non-resonant Anomaly Detection with Background Extrapolation paper. The datasets are simulated using MadGraph5 aMC@NLO, Pythia 8.310, and Delphes. There are 0.2M signal events of semi-visible jets in five sets of parameters (invisible-ratio, Z' mass) = { (1/3, 4 TeV), (1/3, 2 TeV), (1/3, 3 TeV), (0, 4 TeV), (2/3, 4 TeV) }, 18.6M background events of SM QCD jets (including background, ideal AD background, and simulated background) for training, and 21.4M background events for testing. The detailed breakdown of number of events after selections in different regions is listed in Table1 of the paper. The input parameter cards used for generating background and signal events are also included.</p>
Sound database of Industrial Machine for Audio Anomaly Detection
<p><span>Audio anomaly detection(AAD) can seamlessly determinefaults in industrial machines and improve the efficiency of predictive maintenance systems. However, the unavailability of audio sound recordings of real industrial machines operating in their actual industrial setup has limited the efficacy of detection systems. Many different audio databases exist having collections of sounds from dummy (or real) systems operating in controlled environments but a collection of audio sounds from actual industrial machines is missing. Therefore, audio sound recordings of an Air compressor machine working in its natural industrial environment are presented. Only real sounds of an actual machine are captured. Synthetic mixing of sounds is avoided. Damaging the machine to create an anomalous state is avoided. Yet fourteen different unhealthy states are identified and their audio recordings are presented. Dataset with varied values of SNRs is also presented. Spectrograms are plotted and spectral shape parameter values of the developed corpus are calculated. The findings demonstrate the divergence in the developed database and its usefulness in building an effective AAD system for a real industrial machine.</span> </p>
Comprehensive Dataset for Detecting Road Anomalies in Diverse Real-World Situations
<p>In Smart Cities, technologies are playing an important role in efficiently managing the rapid growth of the world's industrialization today. The deployment of surveillance cameras has proliferated to improve public safety and security. Many Closed-Circuit Television (CCTV) cameras have been installed to monitor and safeguard public spaces efficiently within the cities. Despite advancements in technology, video and image processing still largely rely on manual observation. This manual analysis is time-consuming, prone to missing critical details, and costly in terms of labor and resources. Nevertheless, monitoring large video feeds for long periods indicates fatigue, demise of focus, and errors, particularly when video surveillance is a necessity. <br>Road anomaly detection is one of the prominent computer vision issues that researchers have investigated to guarantee public safety. Road anomaly identification is increasingly difficult and complex due to the variety and complexity of abnormalities. <br>Deep learning algorithms must be efficient but also need a large dataset to train to recognize road anomalies in different environments. We proposed a custom real-world data set containing road anomaly images and videos that are made available to the public and private surveillance systems. Primary data were collected from diverse sites in Pakistan, and the data were gathered by recording videos and capturing images by using mobile and surveillance cameras The dataset encompasses five major categories of road anomaly effects.: vehicle accidents, vehicle fire, fighting, snatching(gunpoint), and potholes that classification modeling while promoting improvement in both scientific research and realistic application. The dataset also encompasses annotations with You Only Look Once (YOLO) based bounding boxes and class label files in text format for every image. <br>The researchers can utilize data to train and validate their anomaly detection algorithms and models, thus increasing public security and safety. This dataset focuses on natural environment scenes with a detailed examination of safe transportation and impacts on broader environmental knowledge. Data can give to the liable and ethical arrangement of Artificial Intelligence technologies in surveillance security system</p>
EagleEye: A general purpose density anomaly detection method
<p>The Air2m_northern_DJF.npy and Air2m_northern_JJA.npy datasets is derived from the <em>NCEP-NCAR Reanalysis 1 data provided by the NOAA PSL, Boulder, Colorado, USA, from their website at <a href="https://psl.noaa.gov/">https://psl.noaa.gov</a></em>., a robust atmospheric dataset that includes a wide range of climatic measurements essential for comprehensive climate analysis. The original data can be accessed at the NOAA Physical Sciences Laboratory website: https://psl.noaa.gov/data/gridded/data.ncep.reanalysis.html.</p> <p>The LHC Olympics R&D dataset used in the article can be downloaded from https://zenodo.org/records/4536377 .</p> <pre> </pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.