Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
118
datasets available to search
ShareScore release 0.9.0
Dataset results
118 results for “Anomaly Detection”
Anomaly Detection dataset for the ISS Panel mockup
<p>If you use the dataset, please cite:</p> <p><em>Siddhant Shete, Dennis Mronga</em></p> <p><strong>"Adaptive Online Anomaly Detection using Transfer Learning"</strong></p> <p>About the dataset: The dataset is basically used for anomaly detection in the ISS(International Space Station) panel. This dataset was captured from the mockup used for experiments at the institute. The panel replicates the curcuit and control boards at ISS. The dataset is segregated in two parts Normal data and Anomalous data.</p> <p>Contents of <em><strong> ISSPanelDataset.zip </strong></em></p> <ol> <li>Nomal</li> <li>Anomaly</li> </ol> <p>The Anomaly folder has data with different scenarios where the led lights are on, the fan panel cover is missing or some parts are missaligned.</p> <p> </p> <p><em>This dataset is provided by the Robotics Innivation Center, DFKI GmbH.</em></p> <p><em>The grant was provided by Federal Ministry for Economic Affairs and Climate Action </em></p> <p><em>Grant number: 20W1922F</em></p>
Atrial Anomalies Predict Silent Atrial Fibrillation Detected by Implantable Cardiac Monitor in Cryptogenic Stroke
ClinicalTrials.gov study NCT06542770. IPD Sharing: YES. Countries: 1. Publications: 15.
Diagnostic Accuracy Of Forced Oscillation Technique To Detect Lung Function Anomalies
ClinicalTrials.gov study NCT04006964. IPD Sharing: Not stated. Countries: 2. Publications: 17.
Using Bursty Announcements for Detecting BGP Routing Anomalies
<p>This dataset contains all the require data to reproduce the Indonesia incident. This is provided to facilitate reproducibility of results presented in the following paper:</p> <p>Pablo Moriano, Raquel Hill, and L. Jean Camp. "Using bursty announcements for detecting BGP routing anomalies." vol. 188, 107835, 2021. DOI: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.comnet.2021.107835" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.comnet.2021.107835</span></a></p> <p>These data are provided for non-commercial purposes only. If you use this dataset for research, please be sure to cite the above paper.</p>
GECCO Industrial Challenge 2018 Dataset: A water quality dataset for the 'Internet of Things: Online Anomaly Detection for Drinking Water Quality' competition at the Genetic and Evolutionary Computation Conference 2018, Kyoto, Japan.
<p>Dataset of the 'Internet of Things: Online Anomaly Detection for Drinking Water Quality' competition hosted at The Genetic and Evolutionary Computation Conference (GECCO) July 15th-19th 2018, Kyoto, Japan</p> <p> </p> <p>The task of the competition was to develop an anomaly detection algorithm for a water- and environmental data set.</p> <p> </p> <p>Included in zenodo: </p> <p>- dataset of water quality data</p> <p>- additional material and descriptions provided for the competition</p> <p> </p> <p>The competition was organized by:</p> <p>F. Rehbach, M. Rebolledo, S. Moritz, S. Chandrasekaran, T. Bartz-Beielstein (TH Köln)</p> <p> </p> <p>The dataset was provided by:</p> <p>Thüringer Fernwasserversorgung and IMProvT research project</p> <p> </p> <p>GECCO Industrial Challenge: 'Internet of Things: Online Anomaly Detection for Drinking Water Quality'</p> <p>Description:</p> <p>For the 7th time in GECCO history, the SPOTSeven Lab is hosting an industrial challenge in cooperation with various industry partners. This years challenge, based on the 2017 challenge, is held in cooperation with "Thüringer Fernwasserversorgung" which provides their real-world data set. The task of this years competition is to develop an anomaly detection algorithm for the water- and environmental data set. Early identification of anomalies in water quality data is a challenging task. It is important to identify true undesirable variations in the water quality. At the same time, false alarm rates have to be very low.<br> Additionally to the competition, for the first time in GECCO history we are now able to provide the opportunity for all participants to submit 2-page algorithm descriptions for the GECCO Companion. Thus, it is now possible to create publications in a similar procedure to the Late Breaking Abstracts (LBAs) directly through competition participation!</p> <p> </p> <p>Accepted Competition Entry Abstracts<br> - Online Anomaly Detection for Drinking Water Quality Using a Multi-objective Machine Learning Approach (Victor Henrique Alves Ribeiro and Gilberto Reynoso Meza from the Pontifical Catholic University of Parana)<br> - Anomaly Detection for Drinking Water Quality via Deep BiLSTM Ensemble (Xingguo Chen, Fan Feng, Jikai Wu, and Wenyu Liu from the Nanjing University of Posts and Telecommunications and Nanjing University)<br> - Automatic vs. Manual Feature Engineering for Anomaly Detection of Drinking-Water Quality (Valerie Aenne Nicola Fehst from idatase GmbH)</p> <p>Official webpage:</p> <p><a href="http://www.spotseven.de/gecco/gecco-challenge/gecco-challenge-2018/">http://www.spotseven.de/gecco/gecco-challenge/gecco-challenge-2018/</a></p>
Data for "Wave anomaly detection in wave buoy measurements"
<p>Data used for the production of figures in the manuscript "Wave anomaly detection in wave buoy measurements" submitted for review.</p>
Datasets for ICSE'25 submission: "Scalable and Adaptive Log-based Anomaly Detection: A Synergistic Approach"
Open the record for dataset details and reuse information.
Domain-independent anomalies datasets (adaptions of the MVTec Anomaly Detection dataset)
<p>An adaption of the MVTec Anomaly Detection dataset, presented in the paper <em><a href="https://arxiv.org/abs/2407.02910">Domain-independent detection of known anomalies</a></em></p> <p>The datasets may be used to evaluate approaches on the hybrid task of detecting known anomalies across different, previously unseen objects.</p> <p>The source code for training and testing models on theses datasets can be found at: <a href="https://doi.org/10.5281/zenodo.11924708">https://doi.org/10.5281/zenodo.11924708</a></p>
Anomaly Detection on Dynamic Knowledge Graphs
Open the record for dataset details and reuse information.
EAD: Effortless Anomalies Detection, A deep learning based approach for detecting outliers in textual data
<p>Xiuzhe Wang used this data set for his project</p>
Filming the sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms
<p>This repository makes available the dataset related to the paper:</p> <p>Zafer Çınar, Alessandro Russo, Matteo Spanio, Niccolò Pretto, and Sergio Canazza, <em>Filming the Sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms</em>, IAI4CH, Bozen, 2024.</p> <p>The dataset and the experiment are described in the publication above.</p> <p>This repository contains two main directories (<strong>bold</strong> indicates directory names):</p> <ul> <li><strong>video samples</strong>: the actual videos used in the paper's experiment. This folder contains four subdirectories - 3.75 ips, 7.5 ips, 15 ips, and 30 ips - each representing a different playback speed (in inches per second). Within each subdirectory are several MP4 files, recorded on an A810 Studer open reel recorder, documenting the playback of magnetic audio tapes. The files follow the naming convention “Xips (Y).mp4,” where <em>X</em> represents the tape playback speed and <em>Y</em> is a serial number identifier for each video.</li> <li><strong>irregularities</strong>: the metadata for each video with timestamp and type of irregularity. The folder includes four CSV files - 3.75.csv, 7.5.csv, 15.csv, and 30.csv - corresponding to the playback speeds of the video samples. Each CSV file provides handmade annotations for its respective videos, with three columns: <ul> <li><em>video_id</em>: name of the video file in the format “Xips (Y).mp4,” where <em>X</em> is the tape speed and <em>Y</em> is the ID number.</li> <li><em>time_label</em>: timestamp indicating the irregularity, formatted as HH:MM:SS.mls.</li> <li><em>irregularity_type</em>: category of the detected anomaly, which may be one of the following: “splice,” “shadow,” “end-of-tape,” or “annotation.”</li> </ul> </li> </ul>
ADS-B anomaly detection in the surveillance of low-altitude aircrafts
Open the record for dataset details and reuse information.
Data Anomaly Detection in Cyber-Physical Energy Systems
<p>The data is created in an agent-system, used for controlling distributed energy systems. The agents follow the Lightweight Power Exchange Protocol to negotiate whenever an agent detects a planning problem. The protocol is explained in "A lightweight distributed software agent for automatic demand—supply calculation in smart grids" by Veith, Steinbach and Windeln. Different kinds of anomalies were created by manipulating an agent accordingly: anomalies in the values of the exchanged messages and anomalies in the communication behavior. The implementation of the power exchange used for the data generation can be found here: https://gitlab.com/mango-agents/mango-library/-/tags/Integration_of_the_LPEP.</p> <p>Anomalies in the values are manipulated by 500 % of the original value (dataset 500_p.csv).</p> <p>Anomalies in the communication behavior are anomalously started every minute for 8996 seconds in the future (dataset 1m_8996_50p.csv) and every 15 minutes for 2012 seconds in the future (15m_2012_50p.csv).</p> <p>For anomalies in the communication topology, an agent was chosen to which selected agent does not send any messages, although the agent is part of the neighborhood (Topology anomalies/agent_removed.csv). Furthermore, an agent was manipulated to send messages to an agent which is normally not part of its neighborhood (Topologie anomalies/agent_added.csv).</p> <p>Normal data is also given.</p>
Replication package for: Anomaly Detection Through Container Testing: A Survey of Company Practices
<p>The file contains the survey questionnaire used to collect data for our research.</p>
Data set for Exploring Machine Learning-Based Methods for anomalies detection: Evidence from cryptocurrencies
<p><strong>Exploring Machine Learning-Based Methods for anomalies detection: Evidence from cryptocurrencies</strong></p>
MRI Versus Four Dimensional Ultrasound in Detection of CNS Fetal Congenital Anomalies
ClinicalTrials.gov study NCT03888794. IPD Sharing: UNDECIDED. Countries: 0. Publications: 3.
Data from: Detecting the anomaly zone in species trees and evidence for a misleading signal in higher-level skink phylogeny (Squamata: Scincidae)
Open the record for dataset details and reuse information.
Machine Learning for Anomaly Detection in Cyanobacterial Fluorescence Signals
<p>Excel files containing chlorophyll a and phycocyanin fluorescence data imported from <a href="https://www.glerl.noaa.gov/res/HABs_and_Hypoxia/habTracker.html">https://www.glerl.noaa.gov/res/HABs_and_Hypoxia/habTracker.html</a> for buoys WE2, WE4, WE8, and WE13. The Python code used to manipulate the data is also included.</p>
CADeSH Dataset: Collaborative Anomaly Detection for Smart Homes
<p>Dataset used for quantitative evaluation in the paper:</p> <p>Y. Meidan, D. Avraham, H. Libhaber and A. Shabtai, "CADeSH: Collaborative Anomaly Detection for Smart Homes," in IEEE Internet of Things Journal, 2022, doi: 10.1109/JIOT.2022.3194813.</p> <p> </p> <p>This is a table of flow-level traffic data which was continuously captured during a period of 21 days from five real home networks which were subscribed to a smart home security service, and from our lab at Ben-Gurion University of The Negev. This security service provider shared with us these network traffic flows, plus the related DNS requests and responses, and reputation intelligence of the destination IP addresses. Each instance in this dataset represents an outbound network traffic flow (in the form of an IPFIX) which emanated from an instance of the IoT model streamer.Amazon.Fire_TV_Gen_3.</p> <p>In our lab, we infected our streamer.Amazon.Fire_TV_Gen_3 with a cryptominer and executed cryptomining from this device. To imitate a scanning activity typically performed by some botnets, we also scanned the network using Nmap. In accordance, we labeled these malicious activities as (1) `is executing cryptomining,' or (2) `being scanned by Nmap.' All of the remaining IPFIXs captured in our lab or on the home networks were labeled as `assumed benign'.</p> <p>The multitude of real home networks, and the multitude of identical source devices, enable using this dataset for quantitative evaluation of (collaborative) anomaly/attack detection methods, especially for the IoT.</p>
AnoShift: A distribution shift benchmark for unsupervised anomaly detection
<p>Analyzing the distribution shift of data is a growing research direction in nowadays Machine Learning (ML), leading to emerging new benchmarks that focus on providing a suitable scenario for studying the generalization properties of ML models. The existing benchmarks are focused on supervised learning, and to the best of our knowledge, there is none for unsupervised learning. Therefore, we introduce an unsupervised anomaly detection benchmark with data that shifts over time, built over Kyoto-2006+, a traffic dataset for network intrusion detection. This type of data meets the premise of shifting the input distribution: it covers a large time span (10 years), with naturally occurring changes over time (eg users modifying their behavior patterns, and software updates). We first highlight the non-stationary nature of the data, using a basic per-feature analysis, t-SNE, and an Optimal Transport approach for measuring the overall distribution distances between years. Next, we propose AnoShift, a protocol splitting the data in IID, NEAR, and FAR testing splits. We validate the performance degradation over time with diverse models, ranging from classical approaches to deep learning. Finally, we show that by acknowledging the distribution shift problem and properly addressing it, the performance can be improved compared to the classical training which assumes independent and identically distributed data (on average, by up to 3% for our approach).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.