Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
118
datasets available to search
ShareScore release 0.9.0
Dataset results
118 results for “Anomaly Detection”
High-dimensional Anomaly Detection with Radiative Return in e+e- Collisions
<p>Numpy files of Pythia + Delphes simulated e+e- collisions used as DNN/PFN training inputs.</p>
Data for the research article: "Detecting Seismo-ionospheric Anomalies Possibly Associated with the 2019 Ridgecrest (California) Earthquakes by GNSS, CSES and Swarm Observations"
<p>The zip archives contain 10 .mat files (MATLAB readable). Each .mat file can be loaded into the MATLAB workspace using the <em>load</em> command.</p>
Anomaly Detection and Machine Learning
<p>The datasets were preprocessed. Correlated features were removed.</p> <p>Related papers:</p> <p><strong>[1]</strong> Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization”, 4th International Conference on Information Systems Security and Privacy (ICISSP), Portugal, January 2018</p> <p><strong>[2]</strong> Nour Moustafa, October 16, 2019, "UNSW_NB15 dataset", IEEE Dataport, doi: https://dx.doi.org/10.21227/8vf7-s525.</p> <p><strong>[3]</strong> “Sebastian Garcia, Agustin Parmisano, & Maria Jose Erquiaga. (2020). IoT-23: A labeled dataset with malicious and benign IoT network traffic (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4743746”</p> <p><strong>[4]</strong> A. D. Kent, “Comprehensive, Multi-Source Cybersecurity Events,” Los Alamos National Laboratory, http://dx.doi.org/10.17021/1179829, 2015.</p> <p> </p>
Dataset for Quantum anomaly detection in the latent space of proton collision events at the LHC
<p>Dataset used for https://arxiv.org/abs/2301.10780. The initial dataset is compressed to a low-dimensional latent space using a deep autoencoder. Files with compressed data are provided here in HDF5 format. Different sets of files are given, for different choices of dimensionality for the latent space. A description of the dataset is provided in https://arxiv.org/abs/2301.10780</p>
Log-based anomaly detection datasets
<p>Dataset for the ICSE'22 paper: Log-based Anomaly Detection with Deep Learning: How Far Are We?</p> <p>If you find the data useful for your research, please cite the following paper:</p> <pre>@inproceedings{le2022log, title={Log-based anomaly detection with deep learning: How far are we?}, author={Le, Van-Hoang and Zhang, Hongyu}, booktitle={Proceedings of the 44th international conference on software engineering}, pages={1356--1367}, year={2022} }</pre>
Pattern Recognition and Anomaly Detection in Fetal Morphology Using Deep Learning and Statistical Learning
ClinicalTrials.gov study NCT05738954. IPD Sharing: YES. Countries: 1. Publications: 4.
Performance Anomaly Detection in Microservice Architectures under Continuous Change
<p>Supplementary material for the master's thesis<br> <strong>"Performance Anomaly Detection in Microservice Architectures under Continuous Change"</strong><br> of<br> <strong>Thomas F. Düllmann</strong></p> <p><br> The folder for the supplemental material is structured as follows:.</p> <ul> <li><strong>abstract-de.txt</strong> (abstract in german)</li> <li><strong>abstract-en.txt</strong> (abstract in english)</li> <li><strong>01-MicroserviceMetamodel</strong> (Eclipse project containing the Ecore meta model and the Xtend generation template)</li> <li><strong>02-AnomalyDetectionImplementation</strong> (Eclipse project containing the implementation of the customized RanCorr approach and the EAR approach) <ul> <li><strong>EARExperimentSetup</strong> (the Evaluation Setup that is run based on the input data from the experiment and the EAR implementation)</li> <li><strong>Kieker</strong> (Kieker and the customized RanCorr approach)</li> </ul> </li> <li><strong>03-ExperimentTools</strong> (supplementary microservices and scripts that were used for the experiment setup. The generated services need to be placed next to these files and have to have the folder prefix "gen-") <ul> <li><strong>jmeter</strong> (the microservice that is used for generating load)</li> <li><strong>jmsserver</strong> (the microservice that runs ActiveMQ to bundle the monitoring logs from the services)</li> <li><strong>monitoringserver</strong> (the microservice that collects the monitoring data from the jmsserver microservice and stores them)</li> <li><strong>registry</strong> (the microservice that is responsible for managing the delays that should be injected)</li> <li><strong>copyResults.sh</strong> (simple bash script that masks the scp command to copy the monitoring data from the monitoringserver to the local file system) </li> <li><strong>deployPackage.sh</strong> (bash script that uploads the docker images to defined remote systems via ssh and initiates the start of the microservices on a Kubernetes cluster using the said docker images)</li> <li><strong>dockerinit.sh</strong> (bash script that goes into the microservice folders to compile and package them and create the corresponding Docker images (useful if running minikube for example))</li> <li><strong>kubeinit.sh</strong> (bash script that starts the microservices on a Kubernetes cluster that is associated with the kubectl Kubernetes tool)</li> <li><strong>kubeclean.sh</strong> (bash script that removes the microservices from a Kubernetes cluster that is associated with the kubectl Kubernetes tool) </li> </ul> </li> <li><strong>04-EvaluationData</strong> <ul> <li><strong>RawData </strong> (the raw data that was extracted from the experiment environment) <ul> <li>kieker-monitoring-data (the Kieker monitoring data obtained from the experiment setup)</li> <li><strong>anomalies.log</strong> (the log data that shows the injected real anomalies)</li> <li><strong>events.log</strong> (log file that contains the timesstamps, the scope and the type of event injections)</li> <li><strong>registry.log</strong> (registry microservice log showing the injections of real and change anomalies)</li> </ul> </li> <li><strong>Results</strong> (results for the evaluation with different thresholds) <ul> <li><strong>*folders*</strong> (contain the results of the anomaly detection using the threshold represented by the folder name)</li> <li><strong>results.csv</strong> (the calulated results in terms of TP/FN/FP/FN for every approach with every threshold)</li> <li><strong>results-calculated.csv</strong> (further metrics that were calculated based on the TP/FN/FP/FN values)</li> </ul> </li> </ul> </li> <li><strong>05-EvaluationServices</strong> (the folders containing the microservices that were used for the evaluation)</li> </ul> <p> </p>
Multi-Domain Dataset for Robots (MDDRobots) - Multi-Domain Indoor Dataset for Visual Place Recognition and Anomaly Detection by Mobile Robots
<h2><strong>License</strong></h2> <p>The MDDRobots dataset is made available under the CC BY 4.0 license <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a>.</p> <h2><strong>Summary</strong></h2> <p>The Multi-Domain Dataset for Robots (MDDRobots) contains data for computer vision problems, indoor visual place recognition, and anomaly detection. The recorded images are from different cameras and indoor environmental conditions. </p> <p>It is obligatory to cite the following paper in every work that uses the dataset: <br><strong>Wozniak, P., Krzeszowski, T. & Kwolek, B. Multi-Domain Indoor Dataset for Visual Place Recognition and Anomaly Detection by Mobile Robots. <em>Sci Data</em> 12, 817 (2025). https://doi.org/10.1038/s41597-025-05124-3</strong></p> <h2><strong>Data description</strong></h2> <p>The data are divided into five sets (containing data for different cameras), which have further subsets. Each of the subsets: Training, Test 1, Test 2, and Test 3 consists of nine image sequences. A total of 89,550 three-channel RGB color images in PNG format are organized into 20 zip folders with a whole size of 34.3 GB. Each image in the sequence has a label that represents a room. The number of images for each subset differs due to the split into training and testing data. The difference also results from different methods of recording the image sequences. In order to have balanced data in the subsets, each room in the sequence has the same number of images. Different environmental changes were introduced in each subset. The data from Test 1 are closest to those from the training set. The differences between the sequences are mainly due to changes in the route, robot, and recording equipment. The rooms are well lighted, but not overexposed. The sequences from Test 3 present changed conditions, such as a different time of day, a changed lighting system, and intensive layout changes. The key change is the different paths of the human and the robot. This means a different perspective from previously recorded scenes. The Test 2 sequences pose the most difficult challenge because they contain various recorded activities performed by people moving around rooms. People can occlude important parts of the scene and pass in front of the camera. The images were anonymized by manually blurring the faces of observed people.</p> <h2><strong>Dataset structure<br></strong></h2> <ul> <li>RobotPiCamera_DataSet <ul> <li>DataSet_RobotPiCamera_RGB_train</li> <li>DataSet_RobotPiCamera_RGB_test1</li> <li>DataSet_RobotPiCamera_RGB_test2</li> <li>DataSet_RobotPiCamera_RGB_test3</li> </ul> </li> <li> Xtion_DataSet <ul> <li>DataSet_XTION_RGB_train</li> <li>DataSet_XTION_RGB_test1</li> <li>DataSet_XTION_RGB_test2</li> <li>DataSet_XTION_RGB_test3</li> </ul> </li> <li> GOPRO_DataSet <ul> <li>DataSet_GOPRO_RGB_train</li> <li>DataSet_GOPRO_RGB_test1</li> <li>DataSet_GOPRO_RGB_test2</li> <li>DataSet_GOPRO_RGB_test3</li> </ul> </li> <li>iPhone_DataSet <ul> <li>DataSet_IPHONE_RGB_train</li> <li>DataSet_IPHONE_RGB_test1</li> <li>DataSet_IPHONE_RGB_test2</li> <li>DataSet_IPHONE_RGB_test3</li> </ul> </li> <li>P40PRO_DataSet <ul> <li>DataSet_P40PRO_RGB_train</li> <li>DataSet_P40PRO_RGB_test1</li> <li>DataSet_P40PRO_RGB_test2</li> <li>DataSet_P40PRO_RGB_test3</li> </ul> </li> </ul> <p><em>Example folder content: DataSet_P40PRO_RGB_train\Corridor1_RGB - 00000000.png, 00000001.png, 00000002.png, 00000003.png, ... 00000599.png.</em></p> <p>Total Images (Images per Place)</p> <table> <tbody> <tr> <td>Subset</td> <td>Mounted</td> <td>Training</td> <td>Test 1</td> <td>Test 2</td> <td>Test 3</td> </tr> <tr> <td>Pi Camera</td> <td>Robot</td> <td>7200 (800)</td> <td>5400 (600)</td> <td>5400 (600)</td> <td>5400 (600)</td> </tr> <tr> <td>Xtion</td> <td>Robot</td> <td>7200 (800) </td> <td>1800 (200) </td> <td>1800 (200)</td> <td>1800 (200) </td> </tr> <tr> <td>GoPro</td> <td>Hand</td> <td>5400 (600)</td> <td>4500 (500)</td> <td>4500 (500)</td> <td>4500 (500)</td> </tr> <tr> <td>iPhone</td> <td>Hand</td> <td>5400 (600) </td> <td>4500 (500)</td> <td>4500 (500)</td> <td>4500 (500) </td> </tr> <tr> <td>P40Pro</td> <td>Hand</td> <td>5400 (600)</td> <td>4050 (450)</td> <td>3150 (350) </td> <td>3150 (350) </td> </tr> </tbody> </table> <h2><br>Further information</h2> <p>For any questions, comments or other issues please contact Piotr Woźniak <p.wozniak@prz.edu.pl>.</p>
Research Artifacts - An Anomaly-based Approach for Detecting Modularity Violations on Method Placement
<p>This record contains research artifacts of our paper.</p> <p>The description of each file is as follows.</p> <ul> <li>case_study.xlsx: a spreadsheet that contains a list of our detected cases and results of our manual inspection</li> <li>java-large+.tar.gz: a preprocessed dataset used for training our detection model</li> <li>model_parameters.ckpt.gz: trained parameters and a word dictionary of our neural network model</li> <li>src.tar.gz: source code of our neural network model written in Python and a dataset preprocessor written in Java</li> </ul>
syslrn: Learning What to Monitor for Efficient Anomaly Detection [Dataset]
<p>This repository includes the dataset for the paper:</p> <p><em><a href="http://doi.org/10.1145/3517207.3526979">D. Sanvito, G. Siracusano, S. Santhanam, R. Gonzalez, R. Bifulco</a></em><br> <strong><em><a href="http://doi.org/10.1145/3517207.3526979">syslrn: Learning What to Monitor for Efficient Anomaly Detection </a></em></strong><br> <em><a href="http://doi.org/10.1145/3517207.3526979">ACM EuroMLSys 2022</a></em></p> <p>The dataset contains two directories at the root level:</p> <ul> <li><em><strong>raw_dataset</strong></em></li> <li><strong><em>processed_dataset</em></strong></li> </ul> <p>Each folder in the <strong><em>raw_dataset</em> </strong>directory contains the raw monitoring data used to generate the graph associated to a single experiment together with additional metadata files.<br> Each folder in the <strong><em>processed_dataset</em> </strong>directory contains the graph associated to a single experiment as a set of three CSV files: two for the graph edges (<em>pid_childof_pid_df.csv</em> and <em>pid_speakswith_pid_df.csv</em>) and one for the graph nodes (<em>proc_df.csv</em>).<br> We provide below a code snippet to parse a graph from <strong><em>processed_dataset</em> </strong>directory.</p> <p>In both folders the name of each sub-folder is based on the following schema: <strong><em>[SCENARIO]_[W]wl/test_[TEST_ID]</em></strong> where:</p> <ul> <li><em>[SCENARIO]</em> reports the target component for the failure injection (<em>cinder_failure</em>, <em>neutron_failure</em>, <em>nova_failure</em>). <em>ff</em> indicates instead a failure-free execution</li> <li><em>[W]</em> reports the number of concurrent workloads</li> <li><em>[TEST_ID] </em>reports the ID of the specific failure scenario injected (same ID selected by the OpenStack failure injection framework [1] )</li> </ul> <p>Each experiment includes the following data in the <strong><em>raw_dataset</em></strong> sub-folders:</p> <ul> <li><em>audit_raw_logs_[TEST_ID]/</em>: raw audit monitoring data</li> <li><em>bpf_tools_[TEST_ID]/</em>: raw ebpf tools monitoring data</li> <li><em>instance-[INSTANCE_ID]/</em>: workload-specific metadata files, e.g. stdout/stderr (generated by the OpenStack failure injection framework [1] )</li> <li><em>logs_workload_[TEST_ID]/:</em> OpenStack application logs</li> <li><em>perf_tools_[TEST_ID]/</em>: raw perf tools monitoring data</li> <li><em>audit_filtered_[TEST_ID].log:</em> audit data pre-processed by <em>ausearch</em> (e.g. numerical entities are resolved to symbols)</li> <li><em>failure_[TEST_ID].info</em>: metadata information about the specific failure scenario (generated by the OpenStack failure injection framework [1] )</li> <li><em>timestamps_[TEST_ID]:</em> timing information</li> </ul> <p><em>[1] D. Cotroneo, L. De Simone, P. Liguori, R. Natella, N. Bidokhti - How Bad Can a Bug Get? An Empirical Analysis of Software Failures in the OpenStack Cloud Computing Platform [ACM ESEC/FSE 2019]</em></p> <p> </p> <p>Example: parsing a graph from <strong><em>processed_dataset</em> </strong>directory</p> <pre><code class="language-python">import pandas as pd import networkx as nx def parse_csv(path): processes_df = pd.read_csv('%sproc_df.csv' % path, index_col=0).reset_index(drop=True) speakswith_edges_df = pd.read_csv('%spid_speakswith_pid_df.csv' % path, index_col=0) speakswith_edges_df['type'] = 'speaksWith' childof_edges_df = pd.read_csv('%spid_childof_pid_df.csv' % path, index_col=0) childof_edges_df['type'] = 'childOf' return processes_df, pd.concat([speakswith_edges_df, childof_edges_df], ignore_index=True) def make_graph(nodes_df, edges_df): G = nx.MultiGraph() for _, node in nodes_df.iterrows(): G.add_node(node.pid, **node) for _, edge in edges_df.iterrows(): G.add_edge(edge.pid1, edge.pid2, type=edge.type) return G PATH = 'processed_dataset/ff_1wl/test_1/' nodes_df, edges_df = parse_csv(PATH) G = make_graph(nodes_df, edges_df) nx.draw_networkx(G, node_size=10, with_labels=False)</code></pre> <p> </p> <p>If you use this dataset for your research, please cite the following paper:</p> <pre><code>@inproceedings{sanvito2022syslrn, title={syslrn: Learning What to Monitor for Efficient Anomaly Detection}, author={Sanvito, Davide and Siracusano, Giuseppe and Santhanam, Sharan and Gonzalez, Roberto and Bifulco, Roberto}, booktitle={2nd European Workshop on Machine Learning and Systems (EuroMLSys '22)}, year={2022}, address = {Rennes, France}, publisher = {ACM}, month = apr, } </code></pre>
Dataset used in Can process mining help in anomaly-based intrusion detection?
<p>This is the dataset used in the paper Can process mining help in anomaly-based intrusion detection?</p>
Detecting anomalies in system logs with a compact convolutional transformer - Data
<p><strong>Detecting anomalies in system logs with a compact convolutional transformer - Data</strong></p> <p>Preprocessed data and a pre-trained model for the Larisch, Vitay, Hamker (2022) publication.</p> <p>The <em>data</em> directory contains the Blue Gene/L data set, after tokenization, shuffling, and splitting in a training and test set (BGL_masked_Xtrain.npy and BGL_masked_Xtest.npy, respectively) and the corresponding labels.<br> Additionally, the BGL_masked_Xtest_uniq.npy and BGL_masked_Ytest_uniq.npy contains the test data, where samples from the training set are removed.</p> <p>The <em>model</em> directory contains a pre-trained compact convolutional transformer (CCT) model. The CCT is trained on the proposed BGL training data and uses a 4x4 convolutional kernel.</p>
IMAD-DS: A Dataset for Industrial Multi-Sensor Anomaly Detection Under Domain Shift Conditions
<p>IMAD-DS is a dataset developed for multi-rate multi-sensor anomaly detection (AD) in industrial environments, that considers varying operational and environmental conditions known as domain shifts.</p> <p><strong>Dataset Overview:</strong></p> <p>This dataset includes data from two scaled industrial machines: a robotic arm and a brushless motor.</p> <p>It includes both normal and abnormal data recorded under various operating conditions to account for domain shifts. These shifts are categorized into:</p> <p>Robotic Arm: The robotic arm is a scaled version of a robotic arm used to move silicon wafers in a factory. Anomalies are created by removing bolts at the nodes of the arm, resulting in an imbalance in the machine.<br>Brushless Motor: The brushless motor is a scaled representation of an industrial brushless motor. Two anomalies are introduced: first, a magnet is moved closer to the motor load, causing oscillations by interacting with two symmetrical magnets on the load; second, a belt that rotates in unison with the motor shaft is tightened, creating mechanical stress.</p> <p>The following domain shifts are included in the dataset:</p> <p>Operational Domain Shifts: Variations caused by changes in machine conditions (e.g., load changes for the robotic arm and speed changes for the brushless motor).</p> <p>Environmental Domain Shifts: Variations due to changes in background noise levels.</p> <p>Combinations of operating and environmental conditions divide each machine's dataset into two subsets: the <em>source domain</em> and the <em>target domain</em>. The source domain has a large number of training examples. The target domain, instead, has limited training data. This discrepancy highlights a common issue in the industry where sufficient training data is often unavailable for the target domain, as machine data is collected under controlled environments that do not fully represent the deployment environments.</p> <p> </p> <p><strong>Data Collection and Processing:</strong></p> <p>Data is collected using the STEVAL-STWINBX1 IoT Sensor Industrial Node. The sensor used to record the dataset are the following.</p> <p>· Analog Microphone (16 kHz)</p> <p>· 3-axis Accelerometer (6.7 kHz)</p> <p>· 3-axis Gyroscope (6.7 kHz)</p> <p>Recordings are conducted in an anechoic chamber to control acoustic conditions precisely</p> <p><strong>Data Format:</strong><strong><br></strong>Files are already divided into train and test sets. Inside each folder, each sensor's data is stored in a separate '.parquet' file.</p> <p>Sensor files related to the <em>same</em> segment of machine data share a unique ID. The mapping of each machine data segment to the sensor files is given in .csv files inside the train and test folders. Those .csv files also contain metadata denoting the operational and environmental conditions of a specific segment.</p> <p> </p> <p> </p> <p> </p>
An Autonomous Drone Swarm for Detecting and Tracking Anomalies among Dense Vegetation
<p><strong>Abstract: </strong></p> <p>Swarms of drones offer an increased sensing aperture, and having them mimic behaviors of natural swarms enhances sampling by adapting the aperture to local conditions. We demonstrate that such an approach makes detecting and tracking heavily occluded targets practically feasible. While object classification applied to conventional aerial images generalizes poorly the randomness of occlusion and is therefore inefficient even under lightly occluded conditions, anomaly detection applied to synthetic aperture integral images is robust for dense vegetation, such as forests, and is independent of pre-trained classes. Our autonomous swarm searches the environment for occurrences of the unknown or unexpected, tracking them while continuously adapting its sampling pattern to optimize for local viewing conditions. We achieved an average positional accuracy of 0.39 m with an average precision of 93.2% and an average recall of 95.9%. Here, adapted particle swarm optimization considers detection confidences and predicted target appearance. We show that sensor noise can effectively be included in the synthetic aperture image integration process, removing the need for a computationally costly optimization of high-dimensional parameter spaces. Finally, we present a complete hard- and software framework that supports low-latency transmission and fast processing of extensive video and telemetry data.</p>
Replication Package: Anomaly detection via runtime monitoring data for structural equation modeling
<p>This replication package contains the following information:</p> <ul> <li><strong>Data extraction from literature & interviews: </strong><em>Generation Structural & Measurement Model via literature and interviews.xlsx</em> - here you can find the mapping of the extracted phrases to inductively summarise information regarding the structural and measurement models.</li> <li><strong>Dataset</strong> of runtime monitoring data extracted from TrainTicket via EvoMaster: <br> <ul> <li><em>TrainTicket faults classification.xlxs:</em> Describes the datasets and their faults, in which microservice the fault is injected for better explainability of the obtained results</li> <li><em>IndicatorDescriptionbasedonAnomalyDetectionToolsInterviews.xlsx:</em> description and mapping of selected indicators to the defined parameters from <a href="https://arxiv.org/abs/2408.07816" target="_blank" rel="noopener">previous work </a></li> <li>Unfortunately, the size of the datasets generated via EvoMaster and their injected faults are too big to upload here, thus, they will be available here: <a href="https://uibkacat-my.sharepoint.com/:f:/g/personal/monika_steidl_uibk_ac_at/EjLMt8SYWwtJtp2YuSaqavcBKJoCQ3b5H_l_OY0ifbVRCA?e=fatKyD" target="_blank" rel="noopener">Datasets with injected anomalies</a><br> <ul> <li>the error description can be found <a href="https://github.com/FudanSELab/train-ticket/wiki/Fault-Description" target="_blank" rel="noopener">here</a></li> <li>the datasets are named ts-error-<em>indicatorOfError</em>-reset.zip because the databases are getting reset so that no anomalies are introduced with wrong database entries</li> </ul> </li> </ul> </li> <li><strong>Code</strong> for handling and transforming data to extract indicators describing the whole system's and microservices' behavior from the collected runtime monitoring data collected from TrainTicket:<br> <ul> <li><a href="https://github.com/moniSt13/ConTest-Parsing" target="_blank" rel="noopener">link to the Github repository</a></li> </ul> </li> <li><strong>reports</strong> regarding the established PLS-SEM model using previously handled and transformed runtime monitoring data. Please be aware that opening the reports can leas to out of memory due to their size: <ul> <li><em>Assessment of Measurement Model: MeasurementModel_TrainTicket_erorcleaned.zip & MeasurementModel_Bootstrap_ALL_TrainTicket_errorcleaned.zip</em> </li> <li><em>Assessment of Structural Model: StructuralModel_TrainTicket_errorcleaned.zip & StructuralModel_Bootstrap_ALL_TrainTicket_errorcleaned.zip</em></li> </ul> </li> </ul> <p><br><br>---------------------------------------</p> <p><em>Future work </em>not elaborated in the associated paper due to space restrictions:</p> <ul> <li><strong>reports regarding F5 error</strong>: PLS-SEM model results without interpretation and further mediating effects between microservices included: F5_error.zip</li> </ul> <p> </p> <p> </p> <p> </p>
Anonymised Phone Call Dataset for Anomaly Detection
<p>The dataset provides anonymized information related to phone calls, including the following details:</p> <p>1. Origin Numbers (A-Numbers)<br>2. Destination Numbers (B-Numbers)<br>3. Timestamp of the call<br>4. Call Result, indicating whether the call was blacklisted (coded as 001) or not (coded as 000)</p> <p>The dataset is divided into two subsets with the following characteristics:</p> <p>Dataset 1<br>- Collection Period: 24th July 2018 to 21st October 2018<br>- Duration: 89 days<br>- Total Records: 83,366,367 examples<br>- Unique A-Numbers: 9,006,011<br>- Unique B-Numbers: 2,387,932</p> <p>Dataset 2<br>- Collection Period: 1st June 2019 to 30th June 2019<br>- Duration: 29 days<br>- Total Records: 32,879,670 examples<br>- Unique A-Numbers: 3,217,069<br>- Unique B-Numbers: 1,380,235</p>
Enhanced Anomalies Detection
Open the record for dataset details and reuse information.
Anomaly Detection Algorithm Performance
<p>Anomaly detection algorithms performance metrics: AUC and Average precision; two sets of 298 + 13 algorithms; 9315 datasets.</p>
Analysis of Anomaly Detection for Artificial Intelligence of Things: A Systematic Literature Mapping
<p>This data set contains the characteristics extracted from each of the works selected from the literary review carried out following the PRISMA methodology.</p>
Anomaly Detection dataset for the fuselage of an aircraft
<p>If you use the dataset, please cite:</p> <p><em>Siddhant Shete, Dennis Mronga</em></p> <p><strong>"Adaptive Online Anomaly Detection using Transfer Learning"</strong></p> <p>About the dataset: The dataset is basically used for anomaly detection in the fuselage of an aircraft manufacturing company. We captured the data on the mockup of the fuselage with several iterations at different distances away from the mockup. The dataset is basically the scans of mockup from top to bottom with and without anomalies. The dataset has been segregated into two panels.</p> <p>Contents of <em><strong> AircraftFuselageMockupDataset.zip </strong></em></p> <ol> <li>Nomal_panel1 </li> <li>Nomal_panel2</li> <li>Anomaly_panel1</li> <li>Anomaly_panel2</li> </ol> <p>Every folder has data at 3 distances 15cm, 25cm, 35cm.</p> <p> </p> <p><em>This dataset is provided by the Robotics Innivation Center, DFKI GmbH.</em></p> <p><em>The grant was provided by Federal Ministry for Economic Affairs and Climate Action </em></p> <p><em>Grant number: 20W1922F</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.