Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
625
datasets available to search
ShareScore release 0.9.0
Dataset results
625 results for “anomaly”
Density Anomalies in White Pine at Harvard Forest 2019
The density of wood is a primary determinant of the amount of carbon sequestered in forests. For boreal and Mediterranean ecosystems, anomalies from the typical intra-annual density increase in radial growth of conifers are predominantly related to drought. We took 41 wood samples at breast height, ten additional samples near branches, and seven samples from the top of 41 white pines to determine the spatial distribution of density anomalies throughout the stem. We measured the ring width, density anomaly presence, position within a ring, and arc of density anomalies at multiple heights. Even in a mesic forest density anomalies in white pine are predominantly occurring during drier growing seasons. Moreover, we examined the spatial extent of density anomalies within the stems of white pines and discovered density anomalies are more likely to occur near branches, at the top of the tree, and in wider rings. Furthermore, the position of the density anomalies within the rings was remarkably consistent with the anomalies occurring roughly 80% into the fully formed ring at all heights for high-frequency years only as well as all years Density anomalies seem to be triggered by exogenous factors but their distribution within the stem varies along endogenous gradients, thus better understanding these systematic anomalies can help us to understand how wood formation has reacted and will respond to the environment and changes therein.
Tropical Pacific SST and wind anomalies generated by a Nonlinear Inverse Model
<p>Tropical Pacific (40S-40N; 120E-50W) sea surface temperature (SST), zonal wind (U) and meridional wind (V) anomalies generated by the Nonlinear Inverse Model described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5). The data consists in 99 realizations (<a href="../api/records/10411023/draft/files/NLIM_output_085.nc/content" target="_blank" rel="noopener noreferrer">NLIM_output_XXX.nc</a>) of 1,000yrs each emulating SST, U, and V monthly anomalies conditions during 1980-2020 (<a href="../api/records/10411023/draft/files/Monthly_obs_1980_2020.nc/content" target="_blank" rel="noopener noreferrer">Monthly_obs_1980_2020.nc</a>) given in a 2.5deg-2.5deg grid. For observations, we used the NOAA Extended Reconstruction SST v5 reanalysis (SST; Huang et al., 2017) and NCEP-NCAR reanalysis (winds; Kalnay et al., 1996) The observed anomalies are calculated as described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5).</p> <p>Given that the stochastic forcing considered is white in time and space (https://doi.org/10.1038/s41612-024-00675-5; Methods, section "Offline simulation of SSH_{12}, PC2, and spatial patterns fron nonlinear inverse model output"), the spatial patterns and lead-lag relationships are better identified using composites. A modification of the methodology that allows for spatially coherent stochastic forcing will be implemented in a future article.</p> <p>When using the data please cite https://doi.org/10.5281/zenodo.10411023 (the data) and Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5; for the methodology). </p> <p>Any question, please contact Cristian Martinez-Villalobos at his email cristian.martinez.v@uai.cl</p> <p>References</p> <p>Martinez-Villalobos, C., Dewitte, B., Garreaud, R.D. <em>et al.</em> Extreme coastal El Niño events are tightly linked to the development of the Pacific Meridional Modes. <em>npj Clim Atmos Sci</em> <strong>7</strong>, 123 (2024). https://doi.org/10.1038/s41612-024-00675-5</p> <p>Huang, B. et al. Extended Reconstructed Sea Surface Temperature, Version 5 (ERSSTv5): Upgrades, Validations, and Intercomparisons. Journal of Climate 30, 8179–8205 (2017).</p> <p>Kalnay, E. et al. The NCEP/NCAR 40-Year Reanalysis Project. Bulletin of the American Meteorological Society 77, 437–471 (1996).</p> <p> </p>
RoHuCAD: Robots and Humans Collaborative Anomaly Detection
<h1>RoHuCAD: Robots and Humans Collaborative Anomaly Detection</h1> <p>RoHuCAD is a dataset of human-robot collaboration in a robotic workshop (check <code>workshop_layout.png</code>). Two robots (collaborative manipulator - cobot, autonomous mobile robot - AMR) assist three human operators in assembly of electronic devices.</p> <p>There are two 8-min long recordings in the dataset. They mostly follow the same scenario, with slightly different anomalies. The data is in ROS Noetic rosbag format.</p> <h2>Included data </h2> <ul> <li>RGBD camera data (color + depth) <ul> <li>3 cameras: <a href="https://www.intelrealsense.com/depth-camera-d435i/">Intel Realsense D435i</a></li> <li>color and depth data at 6 frames per second</li> <li>Intrinsic calibration data</li> <li>Extrinsic calibration data (positions and orientations)</li> </ul> </li> <li>Information about positions of robots <ul> <li>AMR: <a href="https://www.ez-wheel.com/en/development-kit-for-agv-and-amr">Ez-Wheel SWD® Starter Kit</a></li> <li>Cobot: <a href="https://www.universal-robots.com/products/ur10-robot/">Universal Robots UR10e</a></li> </ul> </li> </ul> <h2>Annotations</h2> <p>Annotations of specific anomalies are included (CSV file with columns: event_id, tstart, tend, event_type, person_id, camera_id)</p> <ul> <li>Gestures / poses <ul> <li>BENT</li> <li>T-POSE (hands horizontally to the sides)</li> <li>L+R-UP (both hands up)</li> <li>RH-UP (right hand up)</li> <li>LH-UP (left hand up)</li> <li>SQUAT</li> <li>HI-POSE (waving)</li> </ul> </li> <li>Unsafe behaviour <ul> <li>Human in robot working area</li> <li>Standing back to (moving) robot</li> <li>Looking at phone</li> <li>Human in the way of AMR</li> </ul> </li> <li>Normal activities <ul> <li>Assembling/Working</li> <li>Loading/unloading AMR</li> </ul> </li> </ul> <h2>ROS topics</h2> <ul> <li><code>/tf </code></li> <li><code>/tf_static</code></li> <li><code>/joint_states</code></li> <li>cam_ws2_box <ul> <li><code>/cam_ws2_box/color/camera_info</code></li> <li><code>/cam_ws2_box/color/image_raw/compressed</code></li> <li><code>/cam_ws2_box/depth_registered/camera_info</code></li> <li><code>/cam_ws2_box/depth_registered/image_rect_raw</code></li> </ul> </li> <li>cam_ta2_ws2 <ul> <li><code>/cam_ta2_ws2/color/camera_info</code></li> <li><code>/cam_ta2_ws2/color/image_raw/compressed</code></li> <li><code>/cam_ta2_ws2/depth_registered/camera_info</code></li> <li><code>/cam_ta2_ws2/depth_registered/image_rect_raw</code></li> </ul> </li> <li>cam_ta1_ws2 <ul> <li><code>/cam_ta1_ws2/color/camera_info</code></li> <li><code>/cam_ta1_ws2/color/image_raw/compressed</code></li> <li><code>/cam_ta1_ws2/aligned_depth_to_color/camera_info</code></li> <li><code>/cam_ta1_ws2/aligned_depth_to_color/image_raw</code></li> </ul> </li> </ul> <h2>Acknowledgement</h2> <p>The work leading to these results has received funding from the European Union’s Horizon Europe research and innovation programme within the ULTIMATE project under the Grant Agreement no 101070162.</p>
Tropical multi-datasource and multi-frequency temperature anomalies (TROPTEMP)
<p>Multi-frequency temperature anomalies in the tropical region computed from the datasets: CPC, ECCO2_JPL cube92, J-OFURO, MODIS-Aqua, University of Delaware and Windsat.</p>
Dataset of "Anomaly Detection in Industrial Networks: Current State, Classification, and Key Challenges"
<p>Industrial networks are adapted to their specific requirements, especially in terms of industrial processes. To ensure sufficient security in these networks, it is necessary to set and use security policies that complement government regulations, recommendations, and relevant security standards. This paper aims to provide an in-depth analysis of the anomalies occurring within the networks and propose a structure for collecting valuable data from the experimental site based on dividing anomalies into three main categories:<br>security, operational, and service anomalies (and regular traffic recognition). We present a proof-of-concept solution/design aggregating data in industrial networks for advanced anomaly classification. Multiple data sources such as industrial communication, sensor data (additional sensors controlling device behavior), and HW status data are used as data sources. A total of three scenarios (using a physical testbed) were implemented, where we achieved an accuracy of 0.8540/0.9972 in advanced anomaly classification.</p>
Marine magnetic anomaly data from high resolution surveys off the SW Portuguese coast
<p>This dataset contains <strong>magnetic anomaly grids</strong> that result from the full processing of marine magnetic data collected off the SW Portuguese coast between 2014 and 2019. A total area of ~4400 km<sup>2</sup> was surveyed with average line spacing of 1 nautic mile. Surveys covered the continental shelf and in some regions reaching up to 2500 m bathymetric levels. Total magnetic field data were acquired with a G882 Cesium vapor marine magnetometer towed, towed at sea surface.</p> <p><strong>Full processing</strong> of magnetic data included: layback correction; noise removal; IGRF subtraction; base station correction; line leveling; minimum curvature gridding. The resulting sea level magnetic anomaly grid was further processed for upward continuation and reduction to the pole, providing additional outputs. </p> <p>The following grids are provided in <strong>georeferenced geotiff format</strong>:</p> <ul> <li>Magnetic anomaly (sealevel)</li> <li>Magnetic anomaly reduced to the pole (sealevel)</li> <li>Magnetic anomaly upward continued to 200 m height </li> <li>Magnetic anomaly upward continued to 200 m height, reduced to the pole</li> <li>Magnetic anomaly upward continued to 3000 m height </li> <li>Magnetic anomaly upward continued to 3000 m height, reduced to the pole</li> </ul> <p><strong>Published in</strong>: Neres, M., P. Terrinha, J. Noiva, P. Brito, M. Rosa, L. Batista, C. Ribeiro (2023). <em>New Late Cretaceous and CAMP magmatic sources off West Iberia, from high-resolution magnetic surveys on the continental shelf.</em> <strong>Tectonics</strong>. doi: 10.1029/2022TC007637</p> <p> </p>
Dataset for "Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections"
<p>This dataset is used to reproduce the results presented in the following publication:</p> <p>Goelzer, H., Noel, B. P. Y., Edwards, T. L., Fettweis, X., Gregory, J. M., Lipscomb, W. H., van de Wal, R. S. W., and van den Broeke, M. R.: Remapping of Greenland ice sheet surface mass balance anomalies for large ensemble sea-level change projections, The Cryosphere Discuss., https://doi.org/10.5194/tc-2019-188, in review, 2019.</p> <p> </p>
Anomaly detection in the Zwicky Transient Facility DR3
<p>The feature data set extracted from <a href="https://www.ztf.caltech.edu/page/dr3">ZTF DR3</a> light curves. It was used in <a href="https://arxiv.org/abs/2012.01419">Malanchev et al. 2020</a> to detect anomalous astrophysical sources in ZTF data. </p> <p>"feature_XXX.dat" files contain object-ordered light curve feature data, every object is built on 42 feature values, which are encoded as little endian single precision IEEE-754 float (32bit float) numbers. Feature code-names are the same for all three data sets and are listed in plain text files "feature_XXX.name", one code-name per line. "oid_XXX.dat" files contain ZTF DR object identifiers encoded as little endian 64-bit unsigned integer numbers. "oid_XXX.dat" and "feature_XXX.dat" have same object order, for example the first 8 bytes of "oid_m31.dat" files contain the OID of the ZTF DR3 light curve which feature are presented in the first 168 bytes of "feature_m31.dat" file. "m31", "deep" and "disk" denote different ZTF fields and contain 57 546, 406 611, 1 790 565 objects. Note that observations between 58194 ≤ MJD ≤ 58483 are used, see <a href="https://doi.org/10.1093/mnras/stab316">the paper</a> for field and features details.</p> <p>The sample Python code to access the data as Numpy arrays:</p> <pre><code class="language-python">import numpy as np oid = np.memmap('oid_m31.dat', mode='r', dtype=np.uint64) with open('feature_m31.name') as f: names = f.read().split() dtype = [(name, np.float32) for name in names] feature = np.memmap('feature_m31.dat', mode='r', dtype=dtype, shape=oid.shape) idx = np.argmax(feature['amplitude']) print('Object {} has maximum amplitude {:.3f}'.format(oid[idx], feature['amplitude'][idx]))</code></pre> <p> </p>
Dataset for the paper "Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset"
<p>We present a large-scale anomaly detection dataset collected from IBM Cloud's Console over approximately 4.5 months. This high-dimensional dataset captures telemetry data from multiple data centers, specifically designed to aid researchers in developing and benchmarking anomaly detection methods in large-scale cloud environments. It contains 39,365 entries, each representing a 5-minute interval, with 117,448 features/attributes, as interval_start is used as the index. The dataset includes detailed information on request counts, HTTP response codes, and various aggregated statistics. The dataset also includes labeled anomaly events identified through IBM's internal monitoring tools, providing a comprehensive resource for real-world anomaly detection research and evaluation.</p> <p><strong>File Descriptions</strong></p> <ul> <li><code>location_downtime.csv</code> - Details planned and unplanned downtimes for IBM Cloud data centers, including start and end times in ISO 8601 format.</li> <li><code>unpivoted_data.parquet</code> - Contains raw telemetry data with 413 million+ rows, covering details like location, HTTP status codes, request types, and aggregated statistics (min, max, median response times).</li> <li><code>anomaly_windows.csv</code> - Ground truth for anomalies, listing start and end times of recorded anomalies, categorized by source (Issue Tracker, Instant Messenger, Test Log).</li> <li><code>pivoted_data_all.parquet</code> - Pivoted version of the telemetry dataset with 39,365 rows and 117,449 columns, including aggregated statistics across multiple metrics and intervals.</li> <li><code>demo/demo.[ipynb|html]</code>: This demo file provides examples of how to access data in the Parquet files, available in Jupyter Notebook (<code>.ipynb</code>) and HTML (<code>.html</code>) formats, respectively.</li> </ul> <p>Further details of the dataset can be found in <strong>Appendix B: Dataset Characteristics</strong> of the <a href="https://arxiv.org/abs/2411.09047">paper</a> titled <strong><em>"Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset."</em></strong> Sample code for training anomaly detectors using this data is provided in <a href="https://doi.org/10.5281/zenodo.14598119" target="_blank" rel="noopener">this package</a>.</p> <p> </p> <p>When using the dataset, please cite it as follows:</p> <pre><code>@misc{islam2024anomaly,</code><br><code> title={Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset}, </code><br><code> author={Mohammad Saiful Islam and Mohamed Sami Rakha and William Pourmajidi and Janakan Sivaloganathan and John Steinbacher and Andriy Miranskyy},</code><br><code> year={2024},</code><br><code> eprint={2411.09047},</code><br><code> archivePrefix={arXiv},</code><br><code> url={https://arxiv.org/abs/2411.09047}</code><br><code>}</code></pre> <p> </p>
Global Ocean Heat Content Anomalies and Ocean Heat Uptake based on mapping Argo data using local Gaussian processes
<p>Monthly Ocean Heat Content Anomalies (OHCA) in the top 2000 dbar of the ocean are calculated (during 2004-2024, equatorward of 65 degree latitude) subtracting the mean over the period 2004-2024 from the monthly time series of OHC. Yearly OHCA time series are then calculated that include 1. one point per year, i.e., from averaging Jan to Dec (see files ending in “yearly.nc”), and 2. two points per year, i.e., from averaging Jan to Dec and Jul to Jun, respectively (see files ending in “yearly2.nc”). OHC fields are mapped using locally stationary Gaussian processes (defined over space and time) with data-driven decorrelation scales (Kuusela and Stein, 2018). A linear time trend was included in the estimate of the mean field (along with spatial terms and harmonics for the annual cycle). Mapping is done separately for different vertical sections: 15-20 dbar, 15-300 dbar, 300-700 dbar, 700-1850 dbar, 1800-1850 dbar. The 15-20 dbar (1800-1850 dbar) section is used to estimate OHCA for 0-15 dbar (1850-2000 dbar), where observations are sparser. Different vertical sections are combined to estimate global OHCA time series for 0-2000 dbar, 0-700 dbar, 700-2000 dbar (as indicated in the file names). The attribute "area" is included in the netcdf files and it tells the corresponding surface area for the estimates. Regions of the ocean that are shallower than 300 m or are not sufficiently well sampled by the Argo array are not included. Maps of the ocean masks used for the different vertical sections can be found in the .png files (blue shading indicates the area used for the horizontal integral); the bathymetry mask by Roemmich and Gilson (included in the file RG_ArgoClim_Temperature_2019.nc at https://sio-argo.ucsd.edu/RG_Climatology.html) is also used to define the ocean mask. Ocean Heat Uptake is calculated from the monthly OHCA and then averaged as described above to produce yearly time series included in the files for the different layers.</p> <p>For the uncertainty at each time point, the standard deviation of each OHCA/OHU value in the time series is included. When plotting a time series, the user may consider, e.g., shading plus/minus 1* or 1.96*standard deviation (corresponding to a confidence level of 68% or 95% respectively). These standard deviations in the files are estimated using spatially and temporally dependent conditional simulations of monthly gridded anomalies. When combining different layers, the standard deviation of the sum is conservatively estimated as the sum of the standard deviations. </p> <p>Finally, OHCA/OHU trends are estimated via a least-squares fit and reported in the variable metadata with uncertainties (confidence level of 68%). Trend uncertainties are estimated by repeating the fit for each member of the conditional simulation ensemble described above.</p> <p> </p>
MAMAP2D-Light methane column anomalies NetCDF4 and KMZ (28.09.2024)
<div>MAMAP2D-Light methane column anomalies (28.09.2024) NetCDF4 files consist of methane column anomalies retrieved with the WFM-DOAS retrieval for aircraft instruments in the version "PyWFMD v0.2502_jb", alongside additional metadata and geolocation information. All necessary information is stored in the NetCDF4 attributes and variables.</div> <div> </div> <div>The KMZ files are quicklook data of the CH4 column anomalies.</div> <div> </div> <div>Additional information on the data set can be found in the attached Readme.md</div>
OPSSAT-AD - anomaly detection dataset for satellite telemetry
<p>This is the AI-ready benchmark dataset (OPSSAT-AD) containing the telemetry data acquired on board OPS-SAT---a CubeSat mission that has been operated by the European Space Agency.</p> <p>It is accompanied by the paper with baseline results obtained using 30 supervised and unsupervised classic and deep machine learning algorithms for anomaly detection. They were trained and validated using the training-test dataset split introduced in this work, and we present a suggested set of quality metrics that should always be calculated to confront the new algorithms for anomaly detection while exploiting OPSSAT-AD. We believe that this work may become an important step toward building a fair, reproducible, and objective validation procedure that can be used to quantify the capabilities of the emerging anomaly detection techniques in an unbiased and fully transparent way.</p> <p>The included files are:</p> <ul> <li><code>segments.csv</code> with the acquired telemetry signals from ESA OPS-SAT aircraft,</li> <li><code>dataset.csv</code> with the extracted, synthetic features are computed for each manually split and labeled telemetry segment.</li> <li>code files for data processing and example modeliing (<code>dataset_generator.ipynb</code> for data processing, <code>modeling_examples.ipynb</code> with simple examples, <code>requirements.txt</code>- with details on Python configuration, and the <code>LICENSE</code> file)</li> </ul> <p> </p> <p>Please have a look at our two papers commenting on this dataset:</p> <ul> <li>The benchmark paper with results of 30 supervised and unsupervised anomaly detection models for this collection:<br>Ruszczak, B., Kotowski. K., Nalepa, J., Evans, D.:<strong> The OPS-SAT benchmark for detecting anomalies in satellite telemetry, 2024</strong>, <a href="https://arxiv.org/abs/2407.04730" target="_blank" rel="noopener">preprint arxiv: 2407.04730</a>,</li> <li>the conference paper in which we presented some preliminary results for this dataset:<br>Ruszczak, B., Kotowski. K., Andrzejewski, J., et al.: (2023). Machine Learning Detects Anomalies in OPS-SAT Telemetry. Computational Science – ICCS 2023. LNCS, vol 14073. Springer, Cham, <a href="https://doi.org/10.1007/978-3-031-35995-8_21">DOI:10.1007/978-3-031-35995-8_21</a>.</li> </ul>
TROPRAIN: Tropical Precipitation anomalies
<p>The TROPRAIN dataset is a product containing the daily rainfall anomalies across the entire tropical region (180°W - 180°E/30°S - 30°N). These are calculated from the GOES, GPCP, CPC and TRMM data sets. The data is in NetCDF format and is 3D grids. The objective of this data set is to facilitate researchers to study rainfall anomalies during the occurrence of short and medium duration physical processes.</p>
Daily Rainfall Anomalies in the Brazilian Northeast (DRA-BNE)
<p>This dataset contains the daily precipitation anomalies in the Brazilian northeast calculated from the GPCP dataset (combined sources of precipitation measurement) in the period from 1996-10-01 to 2021-07-31, in the geographic boundaries 42°W - 30°W/20.5°S - 0°. The GPCP dataset has an original resolution of 1 degree, but this dataset was interpolated by the bilinear method, obtaining a resolution of 0.25 degrees.</p>
Reconstructed SST-NINO3.4 anomalies for the years 850 to 1981
<p>This data set gives values of reconstructed SST-NINO3.4 anomalies for the years 850 to 1981.</p> <p>SST-NINO3.4 is a key ENSO (El Niño Southern Oscillation) index, defined as the spatial average of Sea Surface Temperature (SST) over the NINO3.4 spatial box (170°W-120°W, 5°S-5°N).</p> <p>The reconstructed SST-NINO3.4 anomalies are obtained from a new multiproxy reconstruction based on 45 proxy records obtained from the Past Global Changes 2k database (PAGES 2k Consortium, 2017, 2019) and a so-called Random Forest method. The reconstructed anomalies are relative to the 1870-2014 averaged SST-NINO3.4 obtained from the HadISST product (Rayner et al. 2003). Units are degrees Celsius</p>
Packaging Industry Anomaly DEtection (PIADE) Dataset
<p>PIADE dataset contains data from five industrial packaging machines:</p> <ul> <li>Machine s_1: from 2020-01-01 14:00:00 to 2021-12-31 13:00:00</li> <li>Machine s_2: from 2020-06-17 08:00:00 to 2021-12-31 07:00:00</li> <li>Machine s_3: from 2020-10-07 12:00:00 to 2022-01-01 23:00:00</li> <li>Machine s_4: from 2020-01-01 01:00:00 to 2022-01-01 23:00:00</li> <li>Machine s_5: from 2020-01-20 08:00:00 to 2022-01-01 12:00:00</li> </ul> <p>## Raw Data</p> <p>Each row represents a production interval, with the following schema:</p> <ul> <li>interval_start: start of the production interval </li> <li>equipment_ID: equipment identifier </li> <li>alarm: alarm code of the active stop reason, if it occurred </li> <li>type: idle, production, downtime, performance_loss or scheduled_downtime </li> <li>start: start of the production interval </li> <li>end: end of the production interval </li> <li>elapsed: duration of the production interval </li> <li>pi: input packages </li> <li>po: output packages </li> <li>speed: speed (packages per hour)</li> </ul> <p>There are 133 different types of alerts, and 429394 rows.<br> </p> <p>## Sequences (1h) data</p> <p>For each piece of equipment, we define sequences of length = 1 hour and we aggregate raw interval data as follows:</p> <ul> <li>'equipment_ID': machine identifier</li> <li>'#changes': changes in machine state</li> <li>'%downtime': time spent in 'downtime' state</li> <li>'%idle': time spent in 'idle' state</li> <li>'%performance_loss': time spent in 'performance loss' state</li> <li>'%production': time spent in production</li> <li>'%scheduled_downtime': time spent in scheduled downtime</li> <li>'count_sum': sum of all alarm occurrences</li> <li>'A_<XXX>': counter of alarm <XXX> occurrences</li> <li>'<state1>/<state2>': number of transitions from <state1> to <state2></li> </ul> <p> </p>
Magnetic Anomaly Map of Paraná State - Final gridded data
<p>This gridded data is part of the article entitled: "THE MAGNETIC ANOMALY MAP OF PARANÁ STATE: AN<br>INTEGRATION OF AIRBORNE SURVEYS PERFORMED OVER THE YEARS", which was submitted in December 2023 to the Brazilian Journal of Geophysics. The article is still under review. </p> <p>These files include airborne magnetic data integrated at 1800m altitude, and the upwarded data to 2700m. Details of the integration and general interpretations are described in the related article.</p>
Forest condition anomaly index values covering Germany for 2016-2023
<p><strong>General description:</strong><br>In <a href="https://doi.org/10.1016/j.rse.2024.114323" target="_blank" rel="noopener">Lange et. al (2024)</a> we utilised <em>Sentinel-2</em> tree species-specific reflectance time series for extracting forest condition across Germany from 2016 to 2022. These time series' seasonal evolution - computed separately for seven natural regions - serves as reference when calculating a similarity metric – further called <em>forest condition anomaly index</em> (FCA). The FCA is computed between each single reflectance observation and the respective date within the reference time series, also considering the natural temporal deviations caused by phenology. FCA temporal aggregation allowed generating spatially comprehensive forest condition anomaly maps. FCA patterns in space and time are in line with dominant drivers like fires, storms and insect infestations and in agreement with state-of-the-art forest disturbance products using a threshold of FCA = −0.15 for forest loss. More information can be found in the <a href="https://doi.org/10.1016/j.rse.2024.114323" target="_blank" rel="noopener">related publication</a> and in the <a title="UFZ Forest condition monitor" href="https://web.app.ufz.de/forestconditionmonitor" target="_blank" rel="noopener">UFZ Forest condition monitor web-application</a>.</p> <p><br><strong>Data description:<br></strong>Data is provided in GeoTiff format (projection <a href="https://epsg.io/32632" target="_blank" rel="noopener">EPSG:32632</a>). Forest condition anomaly maps are available in a spatial resolution of 20 <em>m</em> for the years 2016 to 2023 as monthly (May to October), seasonal (spring, summer and fall) and yearly maps. Values are scaled by 10 000 to reduce the file size. Final FCA values are obtained by dividing the raw values by 10 000 and range from -1 to 1. A negative value generally indicates a poorer forest condition, for example, due to negative changes in chlorophyll or water content or due to crown defoliation. Through validation using forest surveys, data from the <em>Copernicus Emergency Management System</em> and other current maps of forest cover loss, it can be relatively accurate determined that a value below -0.15 indicates a heavily damaged or dead forest stand. Stronger damage (such as significant needle/leaf loss or tree mortality) is generally captured more precise than light damage (such as slight needle/leaf loss). Moderate forest condition values correspondingly show no anomaly and represent the expected normal condition for the respective tree species at the given time within the year. Positive forest condition values indicate a positive deviation from the expected state, which might stem from from positive chlorophyll or water content changes or from denser foliage or needle cover.</p> <p> </p> <p><strong>File descriptions</strong>: <br>Data is provided in zip archives containing maps in GeoTiff format (projection <a href="https://epsg.io/32632" target="_blank" rel="noopener">EPSG:32632</a>). 4 zip files are provided:</p> <ul> <li><em>FCA_v0007-0005_Germany_2016-2023_yearly_R20m.zip</em> contains 8 yearly FCA maps </li> <li><em>FCA_v0007-0005_Germany_2016-2023_seasonal_R20m.zip </em>contains 24 seasonal FCA maps (spring, summer and fall for 2016 to 2023)</li> <li><em>FCA_v0007-0005_Germany_2016-2019_monthly_R20m.zip</em> contains 24 monthly maps (May to October for 2016 to 2019)</li> <li><em>FCA_v0007-0005_Germany_2020-2023_monthly_R20m.zip</em> contains 24 monthly maps (May to October for 2020 to 2023)</li> </ul> <p> </p> <p><strong>Please note:</strong><br>Forest pixels were selected according to the tree species map from <a href="https://doi.org/10.1016/j.rse.2024.114069" target="_blank" rel="noopener">Blickensdörfer et al. (2024)</a>. </p>
Data set for anomaly detection on a HPC system
<p>This data set contains the data collected on the DAVIDE HPC system (CINECA & E4 & University of Bologna, Bologna, Italy) in the period March-May 2018.</p> <p>The data set has been used to train a autoencoder-based model to automatically detect anomalies in a semi-supervised fashion, on a real HPC system.</p> <p>This work is described in:</p> <p>1) "Anomaly Detection using Autoencoders in High Performance Computing Systems", <a href="https://arxiv.org/search/cs?searchtype=author&query=Borghesi%2C+A">Andrea Borghesi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Bartolini%2C+A">Andrea Bartolini</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Lombardi%2C+M">Michele Lombardi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Milano%2C+M">Michela Milano</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Benini%2C+L">Luca Benini,</a> IAAI19 (proceedings in process) -- https://arxiv.org/abs/1902.08447</p> <p>2) "Online Anomaly Detection in HPC Systems", <a href="https://arxiv.org/search/cs?searchtype=author&query=Borghesi%2C+A">Andrea Borghesi</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Libri%2C+A">Antonio Libri</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Benini%2C+L">Luca Benini</a>, <a href="https://arxiv.org/search/cs?searchtype=author&query=Bartolini%2C+A">Andrea Bartolini, </a>AICAS19 (proceedings in process) -- https://arxiv.org/abs/1811.05269</p> <p>See the git repository for usage examples & details --> https://github.com/AndreaBorghesi/anomaly_detection_HPC</p>
Three Annotated Anomaly Detection Datasets for Line-Scan Algorithms
<h1>Summary</h1> <p>This dataset contains two hyperspectral and one multispectral anomaly detection images, and their corresponding binary pixel masks. They were initially used for real-time anomaly detection in line-scanning, but they can be used for any anomaly detection task.</p> <p>They are in .npy file format (will add tiff or geotiff variants in the future), with the image datasets being in the order of (height, width, channels). The SNP dataset was collected using sentinelhub, and the Synthetic dataset was collected from AVIRIS. The Python code used to analyse these datasets can be found at: https://github.com/WiseGamgee/HyperAD</p> <h1>How to Get Started</h1> <p>All that is needed to load these datasets is Python (preferably 3.8+) and the NumPy package. Example code for loading the Beach Dataset if you put it in a folder called "data" with the python script is:</p> <pre><code>import numpy as np # Load image file hsi_array = np.load("data/beach_hsi.npy") n_pixels, n_lines, n_bands = hsi_array.shape print(f"This dataset has {n_pixels} pixels, {n_lines} lines, and {n_bands}.") # Load image mask mask_array = np.load("data/beach_mask.npy") m_pixels, m_lines = mask_array.shape print(f"The corresponding anomaly mask is {m_pixels} pixels by {m_lines} lines.")</code></pre> <h1>Citing the Datasets</h1> <p>If you use any of these datasets, please cite the following paper:</p> <pre><code>@article{garske2024erx,</code><br><code> title={ERX - a Fast Real-Time Anomaly Detection Algorithm for Hyperspectral Line-Scanning},</code><br><code> author={Garske, Samuel and Evans, Bradley and Artlett, Christopher and Wong, KC},</code><br><code> journal={arXiv preprint arXiv:2408.14947},</code><br><code> year={2024},</code><br><code>}</code></pre> <div> <pre>If you use the beach dataset please cite the following paper as well (original source):</pre> </div> <pre><code>@article{mao2022openhsi, title={OpenHSI: A complete open-source hyperspectral imaging solution for everyone}, author={Mao, Yiwei and Betters, Christopher H and Evans, Bradley and Artlett, Christopher P and Leon-Saval, Sergio G and Garske, Samuel and Cairns, Iver H and Cocks, Terry and Winter, Robert and Dell, Timothy}, journal={Remote Sensing}, volume={14}, number={9}, pages={2244}, year={2022}, publisher={MDPI} }</code></pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.