Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
72
datasets available to search
ShareScore release 0.9.0
Dataset results
72 results for “anomaly dataset”
Comprehensive Dataset for Detecting Road Anomalies in Diverse Real-World Situations
<p>In Smart Cities, technologies are playing an important role in efficiently managing the rapid growth of the world's industrialization today. The deployment of surveillance cameras has proliferated to improve public safety and security. Many Closed-Circuit Television (CCTV) cameras have been installed to monitor and safeguard public spaces efficiently within the cities. Despite advancements in technology, video and image processing still largely rely on manual observation. This manual analysis is time-consuming, prone to missing critical details, and costly in terms of labor and resources. Nevertheless, monitoring large video feeds for long periods indicates fatigue, demise of focus, and errors, particularly when video surveillance is a necessity. <br>Road anomaly detection is one of the prominent computer vision issues that researchers have investigated to guarantee public safety. Road anomaly identification is increasingly difficult and complex due to the variety and complexity of abnormalities. <br>Deep learning algorithms must be efficient but also need a large dataset to train to recognize road anomalies in different environments. We proposed a custom real-world data set containing road anomaly images and videos that are made available to the public and private surveillance systems. Primary data were collected from diverse sites in Pakistan, and the data were gathered by recording videos and capturing images by using mobile and surveillance cameras The dataset encompasses five major categories of road anomaly effects.: vehicle accidents, vehicle fire, fighting, snatching(gunpoint), and potholes that classification modeling while promoting improvement in both scientific research and realistic application. The dataset also encompasses annotations with You Only Look Once (YOLO) based bounding boxes and class label files in text format for every image. <br>The researchers can utilize data to train and validate their anomaly detection algorithms and models, thus increasing public security and safety. This dataset focuses on natural environment scenes with a detailed examination of safe transportation and impacts on broader environmental knowledge. Data can give to the liable and ethical arrangement of Artificial Intelligence technologies in surveillance security system</p>
Dataset for Quantum anomaly detection in the latent space of proton collision events at the LHC
<p>Dataset used for https://arxiv.org/abs/2301.10780. The initial dataset is compressed to a low-dimensional latent space using a deep autoencoder. Files with compressed data are provided here in HDF5 format. Different sets of files are given, for different choices of dimensionality for the latent space. A description of the dataset is provided in https://arxiv.org/abs/2301.10780</p>
Log-based anomaly detection datasets
<p>Dataset for the ICSE'22 paper: Log-based Anomaly Detection with Deep Learning: How Far Are We?</p> <p>If you find the data useful for your research, please cite the following paper:</p> <pre>@inproceedings{le2022log, title={Log-based anomaly detection with deep learning: How far are we?}, author={Le, Van-Hoang and Zhang, Hongyu}, booktitle={Proceedings of the 44th international conference on software engineering}, pages={1356--1367}, year={2022} }</pre>
Dataset: timeseries of temperatures and anomalies for the city of Paris (France) for Climate 101 Galaxy training
<p>Dataset is originally downloaded from <a href="https://knmi-ecad-assets-prd.s3.amazonaws.com/ensembles/data/Grid_0.1deg_reg_ensemble/tg_ens_mean_0.1deg_reg_v20.0e.nc">https://knmi-ecad-assets-prd.s3.amazonaws.com/ensembles/data/Grid_0.1deg_reg_ensemble/tg_ens_mean_0.1deg_reg_v20.0e.nc</a> </p> <p>Then 3 single locations have been extracted: </p> <ul> <li>Paris (France): latitude=48.85341,longitude=2.3488</li> <li>Freiburg (Germany): latitude=47.996894,longitude=7.841431</li> <li>Oslo (Norway): latitude=59.911491,longitude=10.75793</li> </ul> <p>Climatologies and anomalies have been computed using <a href="https://code.mpimet.mpg.de/">cdo</a></p> <p>This dataset is meant to be used for teaching purposes only.</p>
ESA Anomaly Dataset
<p>ESA Anomaly Dataset is the first large-scale, real-life satellite telemetry dataset with curated anomaly annotations originated from three ESA missions. We hope that this unique dataset will allow researchers and scientists from academia, research institutes, national and international space agencies, and industry to benchmark models and approaches on a common baseline as well as research and develop novel, computational-efficient approaches for anomaly detection in satellite telemetry data.</p> <p>The dataset results from the work of an 18-month project carried by an industry Consortium composed of Airbus Defence and Space, KP Labs and the European Space Agency’s European Space Operations Centre. The project, funded by the European Space Agency (ESA), is part of the Artificial Intelligence for Automation (A²I) Roadmap (De Canio et al., 2023), a large endeavour started in 2021 to automate space operations by leveraging artificial intelligence.</p> <p>Further details can be found on the <a href="https://arxiv.org/abs/2406.17826">arXiv</a> and <a href="https://github.com/kplabs-pl/ESA-ADB">Github</a>.</p> <p><em>References</em><br>De Canio, G. et al. (2023) Development of an actionable AI roadmap for automating mission operations. In, 2023 SpaceOps Conference. American Institute of Aeronautics and Astronautics, Dubai, United Arab Emirates.</p>
Multi-Domain Dataset for Robots (MDDRobots) - Multi-Domain Indoor Dataset for Visual Place Recognition and Anomaly Detection by Mobile Robots
<h2><strong>License</strong></h2> <p>The MDDRobots dataset is made available under the CC BY 4.0 license <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a>.</p> <h2><strong>Summary</strong></h2> <p>The Multi-Domain Dataset for Robots (MDDRobots) contains data for computer vision problems, indoor visual place recognition, and anomaly detection. The recorded images are from different cameras and indoor environmental conditions. </p> <p>It is obligatory to cite the following paper in every work that uses the dataset: <br><strong>Wozniak, P., Krzeszowski, T. & Kwolek, B. Multi-Domain Indoor Dataset for Visual Place Recognition and Anomaly Detection by Mobile Robots. <em>Sci Data</em> 12, 817 (2025). https://doi.org/10.1038/s41597-025-05124-3</strong></p> <h2><strong>Data description</strong></h2> <p>The data are divided into five sets (containing data for different cameras), which have further subsets. Each of the subsets: Training, Test 1, Test 2, and Test 3 consists of nine image sequences. A total of 89,550 three-channel RGB color images in PNG format are organized into 20 zip folders with a whole size of 34.3 GB. Each image in the sequence has a label that represents a room. The number of images for each subset differs due to the split into training and testing data. The difference also results from different methods of recording the image sequences. In order to have balanced data in the subsets, each room in the sequence has the same number of images. Different environmental changes were introduced in each subset. The data from Test 1 are closest to those from the training set. The differences between the sequences are mainly due to changes in the route, robot, and recording equipment. The rooms are well lighted, but not overexposed. The sequences from Test 3 present changed conditions, such as a different time of day, a changed lighting system, and intensive layout changes. The key change is the different paths of the human and the robot. This means a different perspective from previously recorded scenes. The Test 2 sequences pose the most difficult challenge because they contain various recorded activities performed by people moving around rooms. People can occlude important parts of the scene and pass in front of the camera. The images were anonymized by manually blurring the faces of observed people.</p> <h2><strong>Dataset structure<br></strong></h2> <ul> <li>RobotPiCamera_DataSet <ul> <li>DataSet_RobotPiCamera_RGB_train</li> <li>DataSet_RobotPiCamera_RGB_test1</li> <li>DataSet_RobotPiCamera_RGB_test2</li> <li>DataSet_RobotPiCamera_RGB_test3</li> </ul> </li> <li> Xtion_DataSet <ul> <li>DataSet_XTION_RGB_train</li> <li>DataSet_XTION_RGB_test1</li> <li>DataSet_XTION_RGB_test2</li> <li>DataSet_XTION_RGB_test3</li> </ul> </li> <li> GOPRO_DataSet <ul> <li>DataSet_GOPRO_RGB_train</li> <li>DataSet_GOPRO_RGB_test1</li> <li>DataSet_GOPRO_RGB_test2</li> <li>DataSet_GOPRO_RGB_test3</li> </ul> </li> <li>iPhone_DataSet <ul> <li>DataSet_IPHONE_RGB_train</li> <li>DataSet_IPHONE_RGB_test1</li> <li>DataSet_IPHONE_RGB_test2</li> <li>DataSet_IPHONE_RGB_test3</li> </ul> </li> <li>P40PRO_DataSet <ul> <li>DataSet_P40PRO_RGB_train</li> <li>DataSet_P40PRO_RGB_test1</li> <li>DataSet_P40PRO_RGB_test2</li> <li>DataSet_P40PRO_RGB_test3</li> </ul> </li> </ul> <p><em>Example folder content: DataSet_P40PRO_RGB_train\Corridor1_RGB - 00000000.png, 00000001.png, 00000002.png, 00000003.png, ... 00000599.png.</em></p> <p>Total Images (Images per Place)</p> <table> <tbody> <tr> <td>Subset</td> <td>Mounted</td> <td>Training</td> <td>Test 1</td> <td>Test 2</td> <td>Test 3</td> </tr> <tr> <td>Pi Camera</td> <td>Robot</td> <td>7200 (800)</td> <td>5400 (600)</td> <td>5400 (600)</td> <td>5400 (600)</td> </tr> <tr> <td>Xtion</td> <td>Robot</td> <td>7200 (800) </td> <td>1800 (200) </td> <td>1800 (200)</td> <td>1800 (200) </td> </tr> <tr> <td>GoPro</td> <td>Hand</td> <td>5400 (600)</td> <td>4500 (500)</td> <td>4500 (500)</td> <td>4500 (500)</td> </tr> <tr> <td>iPhone</td> <td>Hand</td> <td>5400 (600) </td> <td>4500 (500)</td> <td>4500 (500)</td> <td>4500 (500) </td> </tr> <tr> <td>P40Pro</td> <td>Hand</td> <td>5400 (600)</td> <td>4050 (450)</td> <td>3150 (350) </td> <td>3150 (350) </td> </tr> </tbody> </table> <h2><br>Further information</h2> <p>For any questions, comments or other issues please contact Piotr Woźniak <p.wozniak@prz.edu.pl>.</p>
Dataset for the article Artificial intelligence for earthquake prediction: a preliminary system based on periodically trained neural networks using ionospheric anomalies
<p>Training and validation data sets along with the corresponding trained convolutional neural network in the article "Artificial intelligence for earthquake prediction: a preliminary system based on periodically trained neural networks using ionospheric anomalies" by Sergio Baselga published in <em>Appl. Sci.</em> <strong>2024</strong>, <em>14</em>(23), 10859; https://doi.org/10.3390/app142310859</p>
syslrn: Learning What to Monitor for Efficient Anomaly Detection [Dataset]
<p>This repository includes the dataset for the paper:</p> <p><em><a href="http://doi.org/10.1145/3517207.3526979">D. Sanvito, G. Siracusano, S. Santhanam, R. Gonzalez, R. Bifulco</a></em><br> <strong><em><a href="http://doi.org/10.1145/3517207.3526979">syslrn: Learning What to Monitor for Efficient Anomaly Detection </a></em></strong><br> <em><a href="http://doi.org/10.1145/3517207.3526979">ACM EuroMLSys 2022</a></em></p> <p>The dataset contains two directories at the root level:</p> <ul> <li><em><strong>raw_dataset</strong></em></li> <li><strong><em>processed_dataset</em></strong></li> </ul> <p>Each folder in the <strong><em>raw_dataset</em> </strong>directory contains the raw monitoring data used to generate the graph associated to a single experiment together with additional metadata files.<br> Each folder in the <strong><em>processed_dataset</em> </strong>directory contains the graph associated to a single experiment as a set of three CSV files: two for the graph edges (<em>pid_childof_pid_df.csv</em> and <em>pid_speakswith_pid_df.csv</em>) and one for the graph nodes (<em>proc_df.csv</em>).<br> We provide below a code snippet to parse a graph from <strong><em>processed_dataset</em> </strong>directory.</p> <p>In both folders the name of each sub-folder is based on the following schema: <strong><em>[SCENARIO]_[W]wl/test_[TEST_ID]</em></strong> where:</p> <ul> <li><em>[SCENARIO]</em> reports the target component for the failure injection (<em>cinder_failure</em>, <em>neutron_failure</em>, <em>nova_failure</em>). <em>ff</em> indicates instead a failure-free execution</li> <li><em>[W]</em> reports the number of concurrent workloads</li> <li><em>[TEST_ID] </em>reports the ID of the specific failure scenario injected (same ID selected by the OpenStack failure injection framework [1] )</li> </ul> <p>Each experiment includes the following data in the <strong><em>raw_dataset</em></strong> sub-folders:</p> <ul> <li><em>audit_raw_logs_[TEST_ID]/</em>: raw audit monitoring data</li> <li><em>bpf_tools_[TEST_ID]/</em>: raw ebpf tools monitoring data</li> <li><em>instance-[INSTANCE_ID]/</em>: workload-specific metadata files, e.g. stdout/stderr (generated by the OpenStack failure injection framework [1] )</li> <li><em>logs_workload_[TEST_ID]/:</em> OpenStack application logs</li> <li><em>perf_tools_[TEST_ID]/</em>: raw perf tools monitoring data</li> <li><em>audit_filtered_[TEST_ID].log:</em> audit data pre-processed by <em>ausearch</em> (e.g. numerical entities are resolved to symbols)</li> <li><em>failure_[TEST_ID].info</em>: metadata information about the specific failure scenario (generated by the OpenStack failure injection framework [1] )</li> <li><em>timestamps_[TEST_ID]:</em> timing information</li> </ul> <p><em>[1] D. Cotroneo, L. De Simone, P. Liguori, R. Natella, N. Bidokhti - How Bad Can a Bug Get? An Empirical Analysis of Software Failures in the OpenStack Cloud Computing Platform [ACM ESEC/FSE 2019]</em></p> <p> </p> <p>Example: parsing a graph from <strong><em>processed_dataset</em> </strong>directory</p> <pre><code class="language-python">import pandas as pd import networkx as nx def parse_csv(path): processes_df = pd.read_csv('%sproc_df.csv' % path, index_col=0).reset_index(drop=True) speakswith_edges_df = pd.read_csv('%spid_speakswith_pid_df.csv' % path, index_col=0) speakswith_edges_df['type'] = 'speaksWith' childof_edges_df = pd.read_csv('%spid_childof_pid_df.csv' % path, index_col=0) childof_edges_df['type'] = 'childOf' return processes_df, pd.concat([speakswith_edges_df, childof_edges_df], ignore_index=True) def make_graph(nodes_df, edges_df): G = nx.MultiGraph() for _, node in nodes_df.iterrows(): G.add_node(node.pid, **node) for _, edge in edges_df.iterrows(): G.add_edge(edge.pid1, edge.pid2, type=edge.type) return G PATH = 'processed_dataset/ff_1wl/test_1/' nodes_df, edges_df = parse_csv(PATH) G = make_graph(nodes_df, edges_df) nx.draw_networkx(G, node_size=10, with_labels=False)</code></pre> <p> </p> <p>If you use this dataset for your research, please cite the following paper:</p> <pre><code>@inproceedings{sanvito2022syslrn, title={syslrn: Learning What to Monitor for Efficient Anomaly Detection}, author={Sanvito, Davide and Siracusano, Giuseppe and Santhanam, Sharan and Gonzalez, Roberto and Bifulco, Roberto}, booktitle={2nd European Workshop on Machine Learning and Systems (EuroMLSys '22)}, year={2022}, address = {Rennes, France}, publisher = {ACM}, month = apr, } </code></pre>
Dataset used in Can process mining help in anomaly-based intrusion detection?
<p>This is the dataset used in the paper Can process mining help in anomaly-based intrusion detection?</p>
Dataset related to publication: Landcover-categorized fires respond distinctly to precipitation anomalies in the South-Central United States
<p>Landcover-categorized fires respond distinctly to precipitation anomalies in the South-Central United States</p> <p>Kátia Fernandes and Sen g. Young</p> <p>doi: 10.3389/fenvs.2024.1433920</p> <p>Abstract</p> <p>Satellite detection of active fires have contributed to advance our understanding of fire ecology, fire and climate dynamics, fire emissions and how to better manage the use of fires as a tool. In this study we use 12 years (2012-2023) of active fire data combined with landcover information in the South-Central United States to derive a monthly, <strong>open access dataset of categorized fires.</strong> This is done by calculating a fire predominance index used to rank fire prone land covers, which are then grouped into four main landscapes: grassland, forest, wildland and crop fires. County level aggregated analyses reveal spatial distributions, climatologies, and peak fire months that are particular to each fire type. Using the Standardized Precipitation Index (SPI), it is found that during climatological fire peak-month, SPI and fires exhibit an inverse relationship in forests and crops, whereas grassland and wildland fires show less consistent inverse or even direct relationship with SPI. This varied behavior is discussed in the context of landscapes’ responses to anomalies in precipitation, and fire management practices, such as prescribed fires and crop residue burning. In a case study of Osage County (OK) we find that large wildfires, known to be closely related to climate anomalies, occur where forest fires are located in the county and absent in areas of grassland fires. Weaker grassland fires response to precipitation anomalies can be attributed to the use of prescribed burning, which are normally planned under environmental conditions that facilitate control and thus avoided during droughts. Crop fires on the other hand, are set to efficiently burn residue and practiced more intensely in drier years than in wetter, explaining the consistently strong inverse correlation between fires and precipitation anomalies. In our increasingly volatile climate, understanding how fires, vegetation, and precipitation interact has become imperative to prevent hazardous fire conflagrations and to better manage ecosystems.</p>
Dataset: 2023 GPS Anomalies, NOTAMs, and Aircraft Traffic
<h1><strong>Dataset: 2023 GPS Anomalies, NOTAMs, and Aircraft Traffic</strong></h1> <p>The dataset "2023 GPS Anomalies, NOTAMs, and Aircraft Traffic" was collected and generated for the paper "Detecting GPS Anomalies in Aviation Using ADS-B: Correlating Coordinate Gaps and GPS Deviations with NOTAM Warnings."</p> <p>This dataset provides a collection of geospatial and temporal data necessary for analyzing potential GPS anomalies in aviation. The data sources include NOTAMs received from the FAA, and the aircraft traffic and GPS information calculated and extracted from the OpenSky Trino ADS-B database.</p> <p>The FAA_and_ICAO_locations file includes 21,382 records with identifiers, coordinates, and detailed facility information. This dataset serves as a reference for analyzing the geographical distribution of aviation facilities. The Flights_per_Hour_per_Grid file, with 74,219,036 records, provides hourly flight movement counts within specified grids, offering insights into air traffic patterns and potential disruptions. The GPS_Jumps_from_Routes file, comprising 5,878,275 records, documents deviations in flight paths, capturing metrics such as distances, speeds, and timestamps. This data is crucial for identifying potential GPS spoofing incidents by analyzing unusual jumps between consecutive data points.</p> <p>The GPS_Missing_Coordinates file, with 53,232 records, highlights periods of missing GPS signals, indicating possible GPS jamming events. This file includes start and end times, distances between known coordinates, and Navigation Integrity Category (NIC) values to assess data quality during null periods. The NOTAM_ICAO_GPS and NOTAM_USA files, with 30,160 and 234,205 records respectively, provide detailed information on NOTAM areas, including geographic areas, active periods, and categories. This allows for an analysis of the spatial and temporal correlation between NOTAM warnings and GPS anomalies, facilitating a better understanding of the impact of GPS disruptions on aviation safety and operations.</p> <h1><strong>Summary Table</strong></h1> <table> <tbody> <tr> <td> <p><strong>Category</strong></p> </td> <td> <p><strong>File Names</strong></p> </td> <td> <p><strong>Total Records</strong></p> </td> <td> <p><strong>Columns</strong></p> </td> </tr> <tr> <td> <p><strong>FAA and ICAO Locations</strong></p> </td> <td> <p>FAA_and_ICAO_locations.csv</p> <p>FAA_and_ICAO_locations.dpkg</p> </td> <td> <p>21,382</p> </td> <td> <p>WKT, id, fid, Location_ID, ICAO_ID, IATA_ID, FAA_Location_Code, Facility_Type, Facility_Name, FAA_New_Location_Code, Coordinates, lat, lon, Region, Country_Code, Country, State_Id, State_Name, City, Location, Effective_Date, Site_Id, ADO, ARTCC_Id, ARTCC_Computer_ID, ARTCC_Name, Tie_In_FSS_Id, Tie_In_FSS_Name, NOTAM_Facility_Id, NOTAM_Service</p> </td> </tr> <tr> <td> <p><strong>Flights per Hour per Grid</strong></p> </td> <td> <p>Flights_per_Hour_per_Grid-2023.csv</p> <p>Flights_per_Hour_per_Grid-2023.dpkg</p> </td> <td> <p>74,219,036</p> </td> <td> <p>grid_id, hour, movement_count, geometry</p> </td> </tr> <tr> <td> <p><strong>GPS Jumps from Routes</strong></p> <p><strong>(possible spoofing)</strong></p> </td> <td> <p>GPS_Jumps_from_Routes-2023.csv</p> <p>GPS_Jumps_from_Routes-2023.dpkg</p> </td> <td> <p>5,878,275</p> </td> <td> <p>WKT, id, fid, icao24, callsign, time_before_spoofing, time_of_spoofing, distance, time_difference, speed_m_s, time_start, time_end</p> </td> </tr> <tr> <td> <p><strong>GPS Missing Coordinates</strong></p> <p><strong>(possible jamming)</strong></p> </td> <td> <p>GPS_Missing_Coordinates-2023.csv</p> <p>GPS_Missing_Coordinates-2023.dpkg</p> </td> <td> <p>53,232</p> </td> <td> <p>WKT, id, icao24, callsign, null_start_time, null_end_time, time_of_previous_not_null_coords, time_of_next_not_null_coords, between_coords_distance_m, null_duration_seconds, between_coords_duration_seconds, avg_nic, min_nic, max_nic, start_time, end_time, start_y, end_x, end_y, start_x</p> </td> </tr> <tr> <td> <p><strong>NOTAM ICAO GPS</strong></p> </td> <td> <p>NOTAM_ICAO_GPS-2023.csv</p> <p>NOTAM_ICAO_GPS-2023.dpkg</p> </td> <td> <p>30,160</p> </td> <td> <p>WKT, id, fid, notam_id, category_name, coordinates_center, radius_nm, radius_mod_nm, notam_number, accountability, location_id, icao_id, domestic_text, icao_text, type, category_id, time_start, time_end</p> </td> </tr> <tr> <td> <p><strong>NOTAM USA</strong></p> </td> <td> <p>NOTAM_USA-2023.csv</p> <p>NOTAM_USA-2023.dpkg</p> </td> <td> <p>234,205</p> </td> <td> <p>WKT, id, fid, notam_id, category_name, is_circle, coordinates_polygon, coordinates_center, radius_nm, faa_location_code, is_faa_location, location_id, is_restricted_area, restricted_area_id, restricted_area_code, category_id, message, notam_number, notam_accountability, moa, type, time_start, time_end</p> </td> </tr> <tr> <td> <p> </p> </td> <td> <p> </p> </td> <td> <p> </p> </td> <td> <p> </p> </td> </tr> </tbody> </table> <h1><strong>Details</strong></h1> <h2><strong>1. FAA_and_ICAO_locations.csv </strong>and <strong>FAA_and_ICAO_locations.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 21,382</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>WKT</strong>: Well-Known Text representation of a point in the CSV file, or a geometry field in the DPKG file.</li> <li><strong>id</strong>: Unique identifier for each record.</li> <li><strong>fid</strong>: Feature identifier.</li> <li><strong>Location_ID</strong>: Identifier for the location.</li> <li><strong>ICAO_ID</strong>: ICAO (International Civil Aviation Organization) identifier.</li> <li><strong>IATA_ID</strong>: IATA (International Air Transport Association) identifier.</li> <li><strong>FAA_Location_Code</strong>: FAA location code.</li> <li><strong>Facility_Type</strong>: Type of facility (e.g., airport, heliport).</li> <li><strong>Facility_Name</strong>: Name of the facility.</li> <li><strong>FAA_New_Location_Code</strong>: New location code by FAA.</li> <li><strong>Coordinates</strong>: Coordinates of the location.</li> <li><strong>lat</strong>: Latitude of the location.</li> <li><strong>lon</strong>: Longitude of the location.</li> <li><strong>Region</strong>: Geographical region of the location.</li> <li><strong>Country_Code</strong>: Country code of the location.</li> <li><strong>Country</strong>: Country name of the location.</li> <li><strong>State_Id</strong>: State identifier.</li> <li><strong>State_Name</strong>: Name of the state.</li> <li><strong>City</strong>: City name.</li> <li><strong>Location</strong>: General location information.</li> <li><strong>Effective_Date</strong>: Effective date of the record.</li> <li><strong>Site_Id</strong>: Site identifier.</li> <li><strong>ADO</strong>: Airport District Office.</li> <li><strong>ARTCC_Id</strong>: ARTCC (Air Route Traffic Control Center) identifier.</li> <li><strong>ARTCC_Computer_ID</strong>: ARTCC computer identifier.</li> <li><strong>ARTCC_Name</strong>: Name of the ARTCC.</li> <li><strong>Tie_In_FSS_Id</strong>: Tie-in Flight Service Station identifier.</li> <li><strong>Tie_In_FSS_Name</strong>: Name of the Tie-in Flight Service Station.</li> <li><strong>NOTAM_Facility_Id</strong>: NOTAM (Notice to Airmen) facility identifier.</li> <li><strong>NOTAM_Service</strong>: Indicates if NOTAM service is available (Y/N).</li> </ul> </ul> <h2><strong>2. Flights_per_Hour_per_Grid-2023.csv </strong>and <strong>Flights_per_Hour_per_Grid-2023.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 74,219,036</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>grid_id</strong>: Identifier for the grid.</li> <li><strong>hour</strong>: Timestamp for the hour.</li> <li><strong>movement_count</strong>: Number of flights in each 0.5x0.5 degree grid during each hour of year 2023.</li> <li><strong>geometry</strong>: Well-Known Text representation of a polygon in the CSV file, or a geometry field in the DPKG file.</li> </ul> </ul> <h2><strong>3. GPS_Jumps_from_Routes-2023.csv </strong>and <strong>GPS_Jumps_from_Routes-2023.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 5,878,275</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>WKT</strong>: Well-Known Text representation of a linestring in the CSV file, or a geometry field in the DPKG file.</li> <li><strong>id</strong>: Unique identifier for each record.</li> <li><strong>fid</strong>: Feature identifier.</li> <li><strong>icao24</strong>: ICAO 24-bit aircraft address.</li> <li><strong>callsign</strong>: Callsign of the aircraft.</li> <li><strong>time_before_spoofing</strong>: Timestamp before the spoofing event.</li> <li><strong>time_of_spoofing</strong>: Timestamp of the spoofing event.</li> <li><strong>distance</strong>: Distance of the jump in meters.</li> <li><strong>time_difference</strong>: Time difference between two coordinates in seconds.</li> <li><strong>speed_m_s</strong>: Speed in meters per second.</li> <li><strong>time_start</strong>: Start time of the record.</li> <li><strong>time_end</strong>: End time of the record.</li> </ul> </ul> <h2><strong>4. GPS_Missing_Coordinates-2023.csv </strong>and <strong>GPS_Missing_Coordinates-2023.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 53,232</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>WKT</strong>: Well-Known Text representation of a linestring in the CSV file, or a geometry field in the DPKG file.</li> <li><strong>id</strong>: Unique identifier for each record.</li> <li><strong>icao24</strong>: ICAO 24-bit aircraft address.</li> <li><strong>callsign</strong>: Callsign of the aircraft.</li> <li><strong>null_start_time</strong>: Start time of missing GPS coordinates.</li> <li><strong>null_end_time</strong>: End time of missing GPS coordinates.</li> <li><strong>time_of_previous_not_null_coords</strong>: Time of the last known good GPS coordinates before the null period.</li> <li><strong>time_of_next_not_null_coords</strong>: Time of the first known good GPS coordinates after the null period.</li> <li><strong>between_coords_distance_m</strong>: Distance between the previous and next known good coordinates in meters.</li> <li><strong>null_duration_seconds</strong>: Duration of the null period in seconds.</li> <li><strong>between_coords_duration_seconds</strong>: Duration between the previous and next known good coordinates in seconds.</li> <li><strong>avg_nic</strong>: Average Navigation Integrity Category (NIC) during the period.</li> <li><strong>min_nic</strong>: Minimum NIC during the period.</li> <li><strong>max_nic</strong>: Maximum NIC during the period.</li> <li><strong>start_time</strong>: Human-readable start time of the null period.</li> <li><strong>end_time</strong>: Human-readable end time of the null period.</li> <li><strong>start_y</strong>: Latitude of the start point.</li> <li><strong>end_x</strong>: Longitude of the end point.</li> <li><strong>end_y</strong>: Latitude of the end point.</li> <li><strong>start_x</strong>: Longitude of the start point.</li> </ul> </ul> <h2><strong>5. NOTAM_ICAO_GPS-2023.csv </strong>and <strong>NOTAM_ICAO_GPS-2023.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 30,160</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>WKT</strong>: Well-Known Text representation of a polygon in the CSV file, or a geometry field in the DPKG file.</li> <li><strong>id</strong>: Unique identifier for each record.</li> <li><strong>fid</strong>: Feature identifier.</li> <li><strong>notam_id</strong>: NOTAM identifier.</li> <li><strong>category_name</strong>: Name of the NOTAM category.</li> <li><strong>coordinates_center</strong>: Center coordinates of the NOTAM area.</li> <li><strong>radius_nm</strong>: Radius in nautical miles.</li> <li><strong>radius_mod_nm</strong>: Modified radius in nautical miles.</li> <li><strong>notam_number</strong>: NOTAM number.</li> <li><strong>accountability</strong>: Accountability of the NOTAM.</li> <li><strong>location_id</strong>: Location identifier.</li> <li><strong>icao_id</strong>: ICAO identifier.</li> <li><strong>domestic_text</strong>: Text of the NOTAM for domestic purposes.</li> <li><strong>icao_text</strong>: Text of the NOTAM for ICAO purposes.</li> <li><strong>type</strong>: Type of NOTAM.</li> <li><strong>category_id</strong>: Identifier for the NOTAM category.</li> <li><strong>time_start</strong>: Start time of the NOTAM.</li> <li><strong>time_end</strong>: End time of the NOTAM.</li> </ul> </ul> <h2><strong>6. NOTAM_USA-2023.csv </strong>and <strong>NOTAM_USA-2023.dpkg</strong></h2> <ul> <li><strong>Total Records</strong>: 234,205</li> <li><strong>Columns</strong>:</li> <ul> <li><strong>WKT</strong>: Well-Known Text representation of a polygon in the CSV file, or a geometry field in the DPKG file.</li> <li><strong>id</strong>: Unique identifier for each record.</li> <li><strong>fid</strong>: Feature identifier.</li> <li><strong>notam_id</strong>: NOTAM identifier.</li> <li><strong>category_name</strong>: Name of the NOTAM category.</li> <li><strong>is_circle</strong>: Indicates if the NOTAM area is a circle (1) or not (0).</li> <li><strong>coordinates_polygon</strong>: Coordinates of the polygon vertices.</li> <li><strong>coordinates_center</strong>: Center coordinates of the NOTAM area.</li> <li><strong>radius_nm</strong>: Radius in nautical miles.</li> <li><strong>faa_location_code</strong>: FAA location code.</li> <li><strong>is_faa_location</strong>: Indicates if it is an FAA location (1) or not (0).</li> <li><strong>location_id</strong>: Location identifier.</li> <li><strong>is_restricted_area</strong>: Indicates if it is a restricted area (1) or not (0).</li> <li><strong>restricted_area_id</strong>: Restricted area identifier.</li> <li><strong>restricted_area_code</strong>: Code for the restricted area.</li> <li><strong>category_id</strong>: Identifier for the NOTAM category.</li> <li><strong>message</strong>: NOTAM message.</li> <li><strong>notam_number</strong>: NOTAM number.</li> <li><strong>notam_accountability</strong>: NOTAM accountability.</li> <li><strong>moa</strong>: Military Operations Area (MOA) identifier.</li> <li><strong>type</strong>: Type of NOTAM.</li> <li><strong>time_start</strong>: Start time of the NOTAM.</li> <li><strong>time_end</strong>: End time of the NOTAM.</li> </ul> </ul> <p> </p> <p>Note that all the files are zipped as CSV. The GPKG version is also available where geographical information is present.</p> <p>Due to the size of the file <strong>Flights_per_Hour_per_Grid-2023</strong>, the command line program ogr2ogr may be the best choice to upload the data into a database. Below is an example of a SQL script and a command to upload this file into a PostgreSQL table. Update the placeholders your_DB_table, your_DB_name, your_DB_user, your_DB_password, your_DB_hostname with the actual information.</p> <p>CREATE TABLE your_DB_table (grid_id TEXT, date TIMESTAMP, movement_count INTEGER, geometry GEOMETRY(POLYGON, 4326));</p> <p>"C:\Program Files\QGIS 3.36.2\bin\ogr2ogr" -f "PostgreSQL" PG:"dbname=your_DB_name user=your_DB_user password=your_DB_password host=your_DB_hostname port=5432" C:\Flights_per_Hour_per_Grid.gpkg -nln your_DB_table -a_srs EPSG:4326 -dim 2 -progress -append</p> <h1><strong>Author</strong></h1> <p>Eugene Pik</p> <p><a href="https://orcid.org/0000-0001-6296-919X">https://orcid.org/0000-0001-6296-919X</a></p> <p><a href="https://www.linkedin.com/in/eugene/">https://www.linkedin.com/in/eugene/</a></p> <p>eugene.pik@mevocopter.com</p> <h1><strong>DOI</strong></h1> <p><a href="https://doi.org/10.5281/zenodo.11411991">https://doi.org/10.5281/zenodo.11411991</a></p> <h1><strong>References</strong></h1> <p><strong>Following references were used to update NOTAMs with WKT polygons:</strong></p> <p>airport-data.com. (n.d.). <em>USA airports by FAA code</em>. https://www.airport-data.com/usa-airports/faa-code/A.html</p> <p>DoD. (2019). <em>Flight information publication area planning special use airspace</em>. NATIONAL GEOSPATIAL-INTELLIGENCE AGENCY. https://www.cnatra.navy.mil/assets-global/docs/area-planning-1A-20190815.pdf</p> <p>FAA. (n.d.-a). <em>Airport Data and Information Portal</em>. https://adip.faa.gov/agis/public/#/airportSearch/advanced</p> <p>FAA. (n.d.-b). <em>US ICAO location finder</em>. https://www.notams.faa.gov/common/icao/USA.html</p> <p>FAA. (2017a, September 30). <em>Encodes/decodes—Aeronautical data</em> [Template]. https://www.faa.gov/air_traffic/flight_info/aeronav/aero_data/loc_id_search/Encodes_Decodes/</p> <p>FAA. (2017b, October 12). <em>Aeronautical information manual—Official guide to basic flight Information and ATC procedures</em>. https://www.faa.gov/air_traffic/publications/media/AIM_Basic_dtd_10-12-17.pdf#page=142</p> <p>FAA. (2021, July 26). <em>Order JO 7350.9Z - Location identifiers</em> [Template]. https://www.faa.gov/regulations_policies/orders_notices/index.cfm/go/document.information/documentID/1040529</p> <p>FAA. (2023a). <em>Pilot’s handbook of aeronautical knowledge</em>. https://www.faa.gov/regulations_policies/handbooks_manuals/aviation/phak</p> <p>FAA. (2023b, June 1). <em>Airport Data</em> [Template]. https://www.faa.gov/air_traffic/flight_info/aeronav/aero_data/Airport_Data/</p> <p>FAA. (2024, February 16). <em>Order JO 7400.10F - Special Use Airspace</em> [Template]. https://www.faa.gov/documentLibrary/media/Order/Order_7400.10F_2024_-_final_-signed.pdf</p> <p>ICAO. (2022). <em>North Atlantic (NAT) air navigation plan Volume I (Doc 9634)</em>. https://www.icao.int/EURNAT/EUR%20and%20NAT%20Documents/NAT%20Documents/_eANP%20NAT%20Doc9634/Doc9634%20NAT%20eANP%20Vol%20I.pdf</p> <p>ProAirPilot.com. (2024). <em>Complete list of NOTAM abbreviations</em>. https://proairpilot.com/notam-abbreviations.html</p> <p>SkyVector. (n.d.). <em>Search for Airports by ICAO ID or name</em>. https://skyvector.com/airports</p> <p> </p> <p><strong>Below is the reference to our source of the aircraft traffic and GPS anomalies data, the OpenSky ADS-B database.</strong></p> <p>Schäfer, M., Strohmeier, M., Lenders, V., Martinovic, I., & Wilhelm, M. (2014). Bringing up OpenSky: A large-scale ADS-B sensor network for research. <em>IPSN-14 Proceedings of the 13th International Symposium on Information Processing in Sensor Networks</em>, 83–94. https://doi.org/10.1109/IPSN.2014.6846743</p> <p> </p>
IMAD-DS: A Dataset for Industrial Multi-Sensor Anomaly Detection Under Domain Shift Conditions
<p>IMAD-DS is a dataset developed for multi-rate multi-sensor anomaly detection (AD) in industrial environments, that considers varying operational and environmental conditions known as domain shifts.</p> <p><strong>Dataset Overview:</strong></p> <p>This dataset includes data from two scaled industrial machines: a robotic arm and a brushless motor.</p> <p>It includes both normal and abnormal data recorded under various operating conditions to account for domain shifts. These shifts are categorized into:</p> <p>Robotic Arm: The robotic arm is a scaled version of a robotic arm used to move silicon wafers in a factory. Anomalies are created by removing bolts at the nodes of the arm, resulting in an imbalance in the machine.<br>Brushless Motor: The brushless motor is a scaled representation of an industrial brushless motor. Two anomalies are introduced: first, a magnet is moved closer to the motor load, causing oscillations by interacting with two symmetrical magnets on the load; second, a belt that rotates in unison with the motor shaft is tightened, creating mechanical stress.</p> <p>The following domain shifts are included in the dataset:</p> <p>Operational Domain Shifts: Variations caused by changes in machine conditions (e.g., load changes for the robotic arm and speed changes for the brushless motor).</p> <p>Environmental Domain Shifts: Variations due to changes in background noise levels.</p> <p>Combinations of operating and environmental conditions divide each machine's dataset into two subsets: the <em>source domain</em> and the <em>target domain</em>. The source domain has a large number of training examples. The target domain, instead, has limited training data. This discrepancy highlights a common issue in the industry where sufficient training data is often unavailable for the target domain, as machine data is collected under controlled environments that do not fully represent the deployment environments.</p> <p> </p> <p><strong>Data Collection and Processing:</strong></p> <p>Data is collected using the STEVAL-STWINBX1 IoT Sensor Industrial Node. The sensor used to record the dataset are the following.</p> <p>· Analog Microphone (16 kHz)</p> <p>· 3-axis Accelerometer (6.7 kHz)</p> <p>· 3-axis Gyroscope (6.7 kHz)</p> <p>Recordings are conducted in an anechoic chamber to control acoustic conditions precisely</p> <p><strong>Data Format:</strong><strong><br></strong>Files are already divided into train and test sets. Inside each folder, each sensor's data is stored in a separate '.parquet' file.</p> <p>Sensor files related to the <em>same</em> segment of machine data share a unique ID. The mapping of each machine data segment to the sensor files is given in .csv files inside the train and test folders. Those .csv files also contain metadata denoting the operational and environmental conditions of a specific segment.</p> <p> </p> <p> </p> <p> </p>
Anonymised Phone Call Dataset for Anomaly Detection
<p>The dataset provides anonymized information related to phone calls, including the following details:</p> <p>1. Origin Numbers (A-Numbers)<br>2. Destination Numbers (B-Numbers)<br>3. Timestamp of the call<br>4. Call Result, indicating whether the call was blacklisted (coded as 001) or not (coded as 000)</p> <p>The dataset is divided into two subsets with the following characteristics:</p> <p>Dataset 1<br>- Collection Period: 24th July 2018 to 21st October 2018<br>- Duration: 89 days<br>- Total Records: 83,366,367 examples<br>- Unique A-Numbers: 9,006,011<br>- Unique B-Numbers: 2,387,932</p> <p>Dataset 2<br>- Collection Period: 1st June 2019 to 30th June 2019<br>- Duration: 29 days<br>- Total Records: 32,879,670 examples<br>- Unique A-Numbers: 3,217,069<br>- Unique B-Numbers: 1,380,235</p>
Dataset for "Lithospheric structure above the Northern Appalachian Anomaly: Initial results from the NEST array"
<p>The preprocessed seismic traces used in the receiver function analysis and resulting negative velocity gradient (NVG) depths. This dataset was used in, "Lithospheric structure above the Northern Appalachian Anomaly: Initial results from the NEST array" by Kimberly Espinal, Maureen Long, Paul Karabinos, and James R. Bourke. The manuscript will soon be available.</p> <p>The receiver functions were processed using a version of the software by Jeffrey Park and Vadim Levin (https://seiscode.iris.washington.edu/projects/rfsyn). </p>
Dataset of "Impactor material records the ancient lunar magnetic field in antipodal anomalies"
<p>This dataset contains input files for iSALE-3D for the paper "Impactor material records the ancient lunar magnetic field in antipodal anomalies" by S. Wakita et al.<br> <br> Please note that usage of the iSALE-3D code is restricted to those who have contributed to the development of iSALE-2D, and iSALE-2D is distributed on a case-by-case basis to academic users in the impact community. It requires a registration from the iSALE webpage (http://www.isale-code.de), and usage of iSALE-2D and computational requirements are also shown there. Please also note that pySALEPlot in the current stable release of iSALE-2D (Dellen) would not work for the data from iSALE-3D.</p>
A dataset of Korean weather with anomaly score from 2010 to 2020
<p>This dataset describes the weather data of 64 cities in Korea for each day and the weather anomaly scores for each day from 2010 to 2020. The dataset includes city name, dates, temperature, humidity, vapor pressure, dew point temperature, sea level pressure, ground pressure, ground temperature, LOF anomaly score, IF anomaly score, COPOD anomaly score, ABOD anomaly score, HBOS anomaly score, SOD anomaly score and ROD anomaly score. In the dataset, the weather data and the weather anomaly score of each day for 64 Korean cities from 2010 to 2020 are stroed into 64 csv files. Each csv file in the dataset represents each city. The 64 cities include Seoul, the capital of Korea, and the 6 metropolitan cities of Busan, Daegu, Incheon, Gwangju, Daejeon, and Ulsan. In addition, the weather data and weather anomaly scores for 19 coastal cities and 4 islands in Korea are included in the dataset.</p>
A dataset of Korean weather with anomaly score from 2010 to 2020
<p>This dataset describes the weather data of 64 cities in Korea for each day and the weather anomaly scores for each day from 2010 to 2020. The dataset includes city name, dates, temperature, humidity, vapor pressure, dew point temperature, sea level pressure, ground pressure, ground temperature, LOF anomaly score, IF anomaly score, COPOD anomaly score, ABOD anomaly score, HBOS anomaly score, SOD anomaly score and ROD anomaly score. In the dataset, the weather data and the weather anomaly score of each day for 64 Korean cities from 2010 to 2020 are stored into 64 csv files. Each csv file in the dataset represents each city. The 64 cities include Seoul, the capital of Korea, and the 6 metropolitan cities of Busan, Daegu, Incheon, Gwangju, Daejeon, and Ulsan. In addition, the weather data and weather anomaly scores for 19 coastal cities and 4 islands in Korea are included in the dataset.</p>
Theoretical synthesis datasets of submarine cable magnetic anomalies.
<p>Theoretical synthesis datasets of submarine cable magnetic anomalies, including training, validation and testing, for end-to-end deep learning. A total of 140000 samples and its coresponding labels.</p>
Anomaly Detection dataset for the fuselage of an aircraft
<p>If you use the dataset, please cite:</p> <p><em>Siddhant Shete, Dennis Mronga</em></p> <p><strong>"Adaptive Online Anomaly Detection using Transfer Learning"</strong></p> <p>About the dataset: The dataset is basically used for anomaly detection in the fuselage of an aircraft manufacturing company. We captured the data on the mockup of the fuselage with several iterations at different distances away from the mockup. The dataset is basically the scans of mockup from top to bottom with and without anomalies. The dataset has been segregated into two panels.</p> <p>Contents of <em><strong> AircraftFuselageMockupDataset.zip </strong></em></p> <ol> <li>Nomal_panel1 </li> <li>Nomal_panel2</li> <li>Anomaly_panel1</li> <li>Anomaly_panel2</li> </ol> <p>Every folder has data at 3 distances 15cm, 25cm, 35cm.</p> <p> </p> <p><em>This dataset is provided by the Robotics Innivation Center, DFKI GmbH.</em></p> <p><em>The grant was provided by Federal Ministry for Economic Affairs and Climate Action </em></p> <p><em>Grant number: 20W1922F</em></p>
Anomaly Detection dataset for the ISS Panel mockup
<p>If you use the dataset, please cite:</p> <p><em>Siddhant Shete, Dennis Mronga</em></p> <p><strong>"Adaptive Online Anomaly Detection using Transfer Learning"</strong></p> <p>About the dataset: The dataset is basically used for anomaly detection in the ISS(International Space Station) panel. This dataset was captured from the mockup used for experiments at the institute. The panel replicates the curcuit and control boards at ISS. The dataset is segregated in two parts Normal data and Anomalous data.</p> <p>Contents of <em><strong> ISSPanelDataset.zip </strong></em></p> <ol> <li>Nomal</li> <li>Anomaly</li> </ol> <p>The Anomaly folder has data with different scenarios where the led lights are on, the fan panel cover is missing or some parts are missaligned.</p> <p> </p> <p><em>This dataset is provided by the Robotics Innivation Center, DFKI GmbH.</em></p> <p><em>The grant was provided by Federal Ministry for Economic Affairs and Climate Action </em></p> <p><em>Grant number: 20W1922F</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.