Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
72
datasets available to search
ShareScore release 0.9.0
Dataset results
72 results for “anomaly dataset”
GECCO Industrial Challenge 2018 Dataset: A water quality dataset for the 'Internet of Things: Online Anomaly Detection for Drinking Water Quality' competition at the Genetic and Evolutionary Computation Conference 2018, Kyoto, Japan.
<p>Dataset of the 'Internet of Things: Online Anomaly Detection for Drinking Water Quality' competition hosted at The Genetic and Evolutionary Computation Conference (GECCO) July 15th-19th 2018, Kyoto, Japan</p> <p> </p> <p>The task of the competition was to develop an anomaly detection algorithm for a water- and environmental data set.</p> <p> </p> <p>Included in zenodo: </p> <p>- dataset of water quality data</p> <p>- additional material and descriptions provided for the competition</p> <p> </p> <p>The competition was organized by:</p> <p>F. Rehbach, M. Rebolledo, S. Moritz, S. Chandrasekaran, T. Bartz-Beielstein (TH Köln)</p> <p> </p> <p>The dataset was provided by:</p> <p>Thüringer Fernwasserversorgung and IMProvT research project</p> <p> </p> <p>GECCO Industrial Challenge: 'Internet of Things: Online Anomaly Detection for Drinking Water Quality'</p> <p>Description:</p> <p>For the 7th time in GECCO history, the SPOTSeven Lab is hosting an industrial challenge in cooperation with various industry partners. This years challenge, based on the 2017 challenge, is held in cooperation with "Thüringer Fernwasserversorgung" which provides their real-world data set. The task of this years competition is to develop an anomaly detection algorithm for the water- and environmental data set. Early identification of anomalies in water quality data is a challenging task. It is important to identify true undesirable variations in the water quality. At the same time, false alarm rates have to be very low.<br> Additionally to the competition, for the first time in GECCO history we are now able to provide the opportunity for all participants to submit 2-page algorithm descriptions for the GECCO Companion. Thus, it is now possible to create publications in a similar procedure to the Late Breaking Abstracts (LBAs) directly through competition participation!</p> <p> </p> <p>Accepted Competition Entry Abstracts<br> - Online Anomaly Detection for Drinking Water Quality Using a Multi-objective Machine Learning Approach (Victor Henrique Alves Ribeiro and Gilberto Reynoso Meza from the Pontifical Catholic University of Parana)<br> - Anomaly Detection for Drinking Water Quality via Deep BiLSTM Ensemble (Xingguo Chen, Fan Feng, Jikai Wu, and Wenyu Liu from the Nanjing University of Posts and Telecommunications and Nanjing University)<br> - Automatic vs. Manual Feature Engineering for Anomaly Detection of Drinking-Water Quality (Valerie Aenne Nicola Fehst from idatase GmbH)</p> <p>Official webpage:</p> <p><a href="http://www.spotseven.de/gecco/gecco-challenge/gecco-challenge-2018/">http://www.spotseven.de/gecco/gecco-challenge/gecco-challenge-2018/</a></p>
Datasets for ICSE'25 submission: "Scalable and Adaptive Log-based Anomaly Detection: A Synergistic Approach"
Open the record for dataset details and reuse information.
Domain-independent anomalies datasets (adaptions of the MVTec Anomaly Detection dataset)
<p>An adaption of the MVTec Anomaly Detection dataset, presented in the paper <em><a href="https://arxiv.org/abs/2407.02910">Domain-independent detection of known anomalies</a></em></p> <p>The datasets may be used to evaluate approaches on the hybrid task of detecting known anomalies across different, previously unseen objects.</p> <p>The source code for training and testing models on theses datasets can be found at: <a href="https://doi.org/10.5281/zenodo.11924708">https://doi.org/10.5281/zenodo.11924708</a></p>
Dataset for "Electrical Conductivity of H2O-rich Silicate Melt: Implications for Subduction Zone Magnetotelluric Anomalies"
Open the record for dataset details and reuse information.
JPL GRACE/GRACE-FO Gridded-AOD1B Water-Equivalent-Thickness Surface-Mass Anomaly RL06.3 dataset for Tellus Level-3 mascon 0.5-degree grid
GRACE non-tidal high-frequency atmospheric and oceanic mass variation models are routinely generated at GFZ as so-called Atmosphere and Ocean De-aliasing Level-1B (AOD1B) products (in terms of corresponding spherical harmonic geopotential coefficients) to be added to the background static gravity model during GRACE monthly gravity field determination. AOD1B products are 3-hourly series of spherical harmonic coefficients up to degree and order 180 which are routinely provided to the GRACE Science Data System and the user community with only a few days time delay. These products reflect spatio-temporal mass variations in the atmosphere and oceans deduced from an operational atmospheric model and corresponding ocean dynamics provided by an ocean model. The variability is derived by subtraction of a long-term mean of vertical integrated atmospheric mass distributions and a corresponding mean of ocean bottom pressure as simulated with the ocean model.<br><br>The Gridded AOD1B data sets provided here contain the monthly mean AOD1B data in geolocated gridded form, smoothed or spatially aggregated to be consistent with the GRACE and GRACE-FO Tellus Level-3 data products of land and/or ocean mass anomalies. With these gridded AOD1B Level-3 products, users can remove or add the effects of the modeled mean monthly atmospheric and ocean bottom pressure change (e.g., to compare different models).
JPL GRACE/GRACE-FO Gridded-AOD1B Water-Equivalent-Thickness Surface-Mass Anomaly RL06.3 dataset for Tellus Level-3 1.0-degree grid
GRACE non-tidal high-frequency atmospheric and oceanic mass variation models are routinely generated at GFZ as so-called Atmosphere and Ocean De-aliasing Level-1B (AOD1B) products (in terms of corresponding spherical harmonic geopotential coefficients) to be added to the background static gravity model during GRACE monthly gravity field determination. AOD1B products are 3-hourly series of spherical harmonic coefficients up to degree and order 180 which are routinely provided to the GRACE Science Data System and the user community with only a few days time delay. These products reflect spatio-temporal mass variations in the atmosphere and oceans deduced from an operational atmospheric model and corresponding ocean dynamics provided by an ocean model. The variability is derived by subtraction of a long-term mean of vertical integrated atmospheric mass distributions and a corresponding mean of ocean bottom pressure as simulated with the ocean model.<br><br>The Gridded AOD1B data sets provided here contain the monthly mean AOD1B data in geolocated gridded form, smoothed or spatially aggregated to be consistent with the GRACE and GRACE-FO Tellus Level-3 data products of land and/or ocean mass anomalies. With these gridded AOD1B Level-3 products, users can remove or add the effects of the modeled mean monthly atmospheric and ocean bottom pressure change (e.g., to compare different models).
RARE: A Labeled Dataset for Cloud-Native Memory Anomalies
<p>The dataset is linked to the paper RARE: A Labeled Dataset for Cloud-Native Memory Anomalies (<a href="https://doi.org/10.1145/3416505.3423560">https://doi.org/10.1145/3416505.3423560</a>)</p> <p>This dataset has been generated using a microservice for injecting artificial byte stream in order to overload the nodes, provoking memory anomalies,</p> <p>It includes 2 files:</p> <p>- List_of_anomalies.csv includes the details on the anomalies injected</p> <p>- RARE.csv is the actual dataset.</p> <p>There are a total of 30 anomalies injected in the dataset. The full dataset comprise 10K labelled time-series, each with 7062 metrics.</p>
CADeSH Dataset: Collaborative Anomaly Detection for Smart Homes
<p>Dataset used for quantitative evaluation in the paper:</p> <p>Y. Meidan, D. Avraham, H. Libhaber and A. Shabtai, "CADeSH: Collaborative Anomaly Detection for Smart Homes," in IEEE Internet of Things Journal, 2022, doi: 10.1109/JIOT.2022.3194813.</p> <p> </p> <p>This is a table of flow-level traffic data which was continuously captured during a period of 21 days from five real home networks which were subscribed to a smart home security service, and from our lab at Ben-Gurion University of The Negev. This security service provider shared with us these network traffic flows, plus the related DNS requests and responses, and reputation intelligence of the destination IP addresses. Each instance in this dataset represents an outbound network traffic flow (in the form of an IPFIX) which emanated from an instance of the IoT model streamer.Amazon.Fire_TV_Gen_3.</p> <p>In our lab, we infected our streamer.Amazon.Fire_TV_Gen_3 with a cryptominer and executed cryptomining from this device. To imitate a scanning activity typically performed by some botnets, we also scanned the network using Nmap. In accordance, we labeled these malicious activities as (1) `is executing cryptomining,' or (2) `being scanned by Nmap.' All of the remaining IPFIXs captured in our lab or on the home networks were labeled as `assumed benign'.</p> <p>The multitude of real home networks, and the multitude of identical source devices, enable using this dataset for quantitative evaluation of (collaborative) anomaly/attack detection methods, especially for the IoT.</p>
Dataset for "Using Large-Scale Anomaly Detection on Code to Improve Kotlin Compiler"
<p>Dataset used in "Using Large-Scale Anomaly Detection on Code to Improve Kotlin Compiler". <br> The data is based on open source code once publicly available on GitHub.</p>
Synthetic basement depth, gravity anomalies, density and observation points training dataset to train deep learning model
<p>Contains 200000 training data for our DNN model.</p>
Training Dataset and Trained Model for Supervised Anomaly Diagnosis
<ul> <li>train_data.hdf: This is the training dataset used to train the Random Forest Classifier model.</li> <li>train_label.hdf: These are the anomaly labels corresponding to each application run in the training dataset.</li> <li>parameters.json: This JSON file helps to construct all possible features (feature extraction version).</li> <li>normal_graph_search_dictionary.json: This file contains a dictionary that helps find the healthy node based on the specific application type users select.</li> <li>search_dataframe.csv: This CSV file contains the raw data of the healthy nodes provided in the normal_graph_search_dictionary.json file.</li> <li>sample_data_2.csv: This is a sample dataset used for the sample data option.</li> <li>node_anoms_sample_2.csv: This file contains the prediction results for the dataset in the sample_data_2.csv.</li> <li>test_model.sav: This is a trained Random Forest Classifier model that was trained offline.</li> </ul> <p> </p> <p> </p>
Anomaly Detection in Semiconductor Wafer Fabrication Using Stream Processing Systems: A Case Study - Dataset
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.