Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
625
datasets available to search
ShareScore release 0.9.0
Dataset results
625 results for “Anomaly”
Dataset for Anomaly Detection Using Inter-Arrival Curves for Real-time Systems
<p>The dataset shows the input files and detailed results for the experiments discussed in the paper. A README file provide more details on the data.</p>
Walls, Pillars and Beams: A 3D Decomposition of Quality Anomalies
<p>This is an artifact package accompanying the paper "Walls, Pillars and Beams: A 3D Decomposition of Quality Anomalies".</p> <p>The package contains a dataset that can be used to reproduce the experiment or compare it with other approaches that may be proposed in the future. We also provide a video which demonstrates our approach.</p>
EIRSAT-1 Test Campaign and Flight Dataset for Anomaly Detection
<p><span>We have curated a unique dataset derived from EIRSAT-1, Ireland's inaugural domestically produced satellite, as a testing and validation resource for these ML models and the future development cycle of AI-enabled small satellites. This dataset consists of a training set developed during ground testing and containing artificial anomalies induced to train satellite operators, a validation dataset containing real anomalies encountered during the qualification campaign, and an early flight test dataset collected since the satellite was launched on December 1<sup>st</sup>, 2023. This paper presents an in-depth analysis of the efficacy of these ML techniques when applied to the EIRSAT-1 dataset, offering insights into their potential to revolutionize the domain of satellite operations through enhanced autonomy and responsiveness. This study not only showcases the capabilities of these ML techniques in an operational environment but also sets the stage for future research and development in autonomous satellite systems.</span></p>
ComplexVAD Video Anomaly Detection Dataset
<p><strong>Introduction</strong></p> <p>The ComplexVAD dataset consists of 104 training and 113 testing video sequences taken from a static camera looking at a scene of a two-lane street with sidewalks on either side of the street and another sidewalk going across the street at a crosswalk. The videos were collected over a period of a few months on the campus of the University of South Florida using a camcorder with 1920 x 1080 pixel resolution. Videos were collected at various times during the day and on each day of the week. Videos vary in duration with most being about 12 minutes long. The total duration of all training and testing videos is a little over 34 hours. The scene includes cars, buses and golf carts driving in two directions on the street, pedestrians walking and jogging on the sidewalks and crossing the street, people on scooters, skateboards and bicycles on the street and sidewalks, and cars moving in the parking lot in the background. Branches of a tree also move at the top of many frames.</p> <p>The 113 testing videos have a total of 118 anomalous events consisting of 40 different anomaly types.</p> <p>Ground truth annotations are provided for each testing video in the form of bounding boxes around each anomalous event in each frame. Each bounding box is also labeled with a track number, meaning each anomalous event is labeled as a track of bounding boxes. A single frame can have more than one anomaly labeled.</p> <p><strong>At a Glance</strong></p> <ul> <li>The size of the unzipped dataset is ~39GB</li> <li>The dataset consists of Train sequences (containing only videos with normal activity), Test sequences (containing some anomalous activity), a ground truth annotation file for each Test sequence, and a README.md file describing the data organization and ground truth annotation format.</li> <li>The zip files contain a Train directory, a Test directory, an annotations directory, and a README.md file.</li> </ul> <p><strong>License</strong></p> <p>The ComplexVAD dataset is released under <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>.</p> <p>All data:</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2024 SPDX-License-Identifier: CC-BY-SA-4.0</code></pre>
Crustal thicknesses, Moho depths and 3-D density anomaly model for GJI paper: Crustal structure of onshore-offshore Atlantic Canada and environs from constrained 3-D gravity inversion using variable mesh depths by J. Kim Welford
<p>The files are provided as ascii text files in terms of both latitudes/longitudes and eastings/northings. For the 3-D density anomaly model, it is provided with columns of x, y, z, and absolute density. The conversions from latitudes/longitudes to eastings/northings for all of the models and maps in this work are computed with ellipsoid WGS-84 and UTM zone 19 using Generic Mapping Tools.</p>
Ocean Heat Content Anomalies in the North Atlantic based on mapping Argo data using local Gaussian processes defined over space
<p>Monthly Ocean Heat Content Anomalies (OHCA) in the top 2000 dbar of the ocean are calculated (during 2005-2022, in the North Atlantic, north of 20N) subtracting the time mean over the period 2005-2021 from the monthly time series of OHC. OHC fields are mapped using a locally stationary Gaussian process (defined over space) with data-driven decorrelation scales (Kuusela and Stein, 2018). A linear time trend was included in the estimate of the mean field (along with spatial terms and harmonics for the annual cycle). In this product, mapping is done in latitude and longitude with monthly subsets of data. Mapping is done separately for different vertical sections: 15-20 dbar, 15-300 dbar, 300-700 dbar, 700-1850 dbar, 1800-1850 dbar. The 15-20 dbar (1800-1850 dbar) section is used to estimate OHCA for 0-15 dbar (1850-2000 dbar), where observations are sparser. Different vertical sections are combined to estimate global OHCA for 0-2000 dbar. Regions of the ocean that are shallower than 300 m or are not sufficiently well sampled by the Argo array are not included. </p>
ICSE2024_anomaly_data
<p>The data set of ICSE2024 draft.</p>
Model Output for: "MESSENGER observations of Mercury's planetary ion escape rates and their dependence on true anomaly angle"
<p>In the manuscript titled ‘MESSENGER Observations of Mercury’s Planetary Ion Escape Rates and Their Dependence on True Anomaly Angle’, submitted to Geophysical Research Letters, we provide an analysis of test sodium ion (Na+) particles. These test particles, integral to our research, are illustrated in Figure 4. This dataset contains an ASCII file with detailed information and descriptions of these test particles.</p>
Experimental data to manuscript "Resonance-Induced Anomalies in Temperature-Dependent Raman Scattering of PdSe2"
<p>Experimental data to manuscript "Resonance-Induced Anomalies in Temperature-Dependent Raman Scattering of PdSe2"</p>
Dataset related to 'Full-waveform inversion reveals diverse origins of lower mantle positive wave speed anomalies'
<p>This repository contains the global distribution of sources and receivers, tomographic models (netCDF4 format), stacked waveforms from the wavefield modelling (.h5 format), 2D grids of the time-depth correlations (.csv format), and Python scripts required for the full analysis and figures presented in the manuscript.</p>
Long-Tailed Anomaly Detection (LTAD) Dataset
<p><strong>Introduction</strong></p> <p>Anomaly detection (AD) aims to identify defective images and localize their defects (if any). Ideally, AD models should be able to: detect defects over many image classes; not rely on hard-coded class names that can be uninformative or inconsistent across datasets; learn without anomaly supervision; and be robust to the long-tailed distributions of real-world applications. To address these challenges, we formulate the problem of long-tailed AD by introducing several datasets with different levels of class imbalance for performance evaluation.</p> <p>To encourage more follow up works on long-tailed AD, we are publicly releasing the dataset split used in our paper (“Long-Tailed Anomaly Detection with Learnable Class Names” by Chih-Hui Ho, Kuan-Chuan Peng, and Nuno Vasconcelos, CVPR 2024).</p> <p>Files in the unzipped folder:</p> <p>1. ./README.md: This Markdown file</p> <p>2. ./dataset_split: Folder contains long-tail splits from three datasets. See below for details.</p> <p><strong> </strong></p> <p><strong>At a Glance</strong></p> <ul> <li>The size of the unzipped dataset is ~16MB</li> <li>Three datasets are used in this project, including [MVTec](https://www.mvtec.com/company/research/datasets/mvtec-ad), [VisA](https://github.com/amazon-science/spot-diff) and [DAGM](https://www.kaggle.com/datasets/mhskjelvareid/dagm-2007-competition-dataset-optical-inspection). Please download the datasets from their original repositories.</li> <li>The dataset split provided in this folder is organized as follows:<br>```<br>dataset_split<br>|---dagm_lt<br>|---mvtec_lt<br>|---visa_lt<br>|-----|-- exp<br>|-----|-----|----- 100<br>|-----|-----|----- |-----test.json<br>|-----|-----|----- |-----train.json<br>|-----|-----|----- 200<br>|-----|-- step<br>|-----|-- ...<br>```</li> <li>Each long-tailed dataset split contains a subfolder ``imbalance_type/imbalance_factor", where imbalance type can be [exponential (exp), step, reverse exponential (exp_reverse), reverse step (step_reverse)]. The definition of imbalance type and imbalance factor can be found in our paper. Each subfolder contains two json files, one for training and the other for testing.</li> <li>Each entry in the json file contains the meta information of an image and is similar to<br>```<br>{"filename": "candle/test/bad/000.JPG", "label": 1, "label_name": "defective", "clsname": "candle", "maskname": "candle/ground_truth/bad/000.png"}<br>```<br>- filename: location of the input image in the dataset<br>- label: indicates whether the input image is normal (labeled as 0) or defective (labeled as 1)<br>- label name: can be "good" or "defective"<br>- clsname: class name of the input image<br>- maskname (optional): location of the binary image that indicates the defect region. This is only available for test.json, because there is no defect image during training.</li> </ul> <p><strong>Citation</strong></p> <p>If you use the LTAD dataset in your research, please cite our contribution:</p> <pre><code>@InProceedings{Ho_2024_CVPR, author = {Ho, Chih-Hui and Peng, Kuan-Chuan and Vasconcelos, Nuno}, title = {Long-Tailed Anomaly Detection with Learnable Class Names}, booktitle = {The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, month = {June}, year = {2024} } </code></pre> <p><strong>License</strong></p> <p>The LTAD dataset is released under <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>. For the images in the MVTec, VisA, and DAGM datasets, please refer to their websites for their copyright and license terms.</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2023-2024 SPDX-License-Identifier: CC-BY-SA-4.0</code></pre>
Street Scene Video Anomaly Detection Dataset
<p><strong><span>Introduction</span></strong></p> <p><span>The Street Scene dataset consists of 46 training video sequences and 35 testing video sequences taken from a static USB camera looking down on a scene of a two-lane street with bike lanes and pedestrian sidewalks.<span> </span>Videos were collected from the camera at various times during two consecutive summers.<span> </span>All of the videos were taken during the daytime.<span> </span>The dataset is challenging because of the variety of activities taking place such as cars driving, turning, stopping and parking; pedestrians walking, jogging and pushing strollers; and bikers riding in bike lanes. In addition, the videos contain changing shadows, and moving background such as a flag and trees blowing in the wind.</span></p> <p><span>There are a total of 202,545 color video frames (56,135 for training and 146,410 for testing) each of size 1280 x 720 pixels. The frames were extracted from the original videos at 15 frames per second.</span></p> <p><span>The 35 testing sequences have a total of 205 anomalous events consisting of 17 different anomaly types. A complete list of anomaly types and the number of each in the test set can be found in our paper.</span></p> <p><span>Ground truth annotations are provided for each testing video in the form of bounding boxes around each anomalous event in each frame. Each bounding box is also labeled with a track number, meaning each anomalous event is labeled as a track of bounding boxes. Track lengths vary from tens of frames to 5200 which is the length of the longest testing sequence. A single frame can have more than one anomaly labeled.</span></p> <p><span>NOTE: This version of the dataset differs slightly with the original made available in 2020.<span> </span>Some anomalies were found in a few of the normal training sequences.<span> </span>These training frames were deleted from the dataset.<span> </span>Specifically, the following frames were removed:</span></p> <p><span>Train026: frames 1-184 (car taking a u-turn)</span></p> <p><span>Train027: frames 1-229 (jay walkers)</span></p> <p><span>Train031: frames 1-299 (jay walkers, illegally parked car)</span></p> <p><strong><span>At a Glance</span></strong></p> <ul> <li><span>The size of the unzipped dataset is ~46GB</span></li> <li><span>The dataset consists of Train sequences (containing only videos with normal activity), Test sequences (containing some anomalous activity) along with ground truth annotations, and a README.md file describing the data organization and ground truth annotation format.</span></li> <li><span>The zip file contains a Train directory, a Test directory and a README.md file.</span></li> </ul> <p><strong><span>Other Resources</span></strong></p> <p><span>None</span></p> <p><strong><span>Citation</span></strong></p> <p><span>If you use the Street Scene dataset in your research, please cite our contribution:</span></p> <pre><code>@inproceedings{ramachandra2020street, title={Street Scene: A new dataset and evaluation protocol for video anomaly detection}, author={Ramachandra, Bharathkumar and Jones, Michael}, booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision}, pages={2569--2578}, year={2020} } </code></pre> <p><strong><span>License</span></strong></p> <p><span>The Street Scene dataset is released under </span><a href="https://creativecommons.org/licenses/by-sa/4.0/"><span>CC-BY-SA-4.0 license</span></a><span>.</span></p> <p><span>All data:</span></p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2023 SPDX-License-Identifier: CC-BY-SA-4.0 </code></pre>
Love thy neighbour? Tropical tree growth and its response to climate anomalies is mediated by neighbourhood hierarchy and dissimilarity in carbon and water related traits
<div> <div>Data and R code to reproduce all analyses, figures and tables for Krebber et al. 2024, Ecology Letters:<br>Love thy neighbour? Tropical tree growth and its response to climate anomalies is mediated by neighbourhood hierarchy and dissimilarity in carbon and water related traits. </div> <div> </div> <div>All analyses have been conducted and produced in the R environment (R version 4.1.2; R Core Team, 2021; RStudio Team, 2020). Bayesian hierarchical models have been run using the R package brms (Version 2.19.0; Bürkner, 2017). The data and R files containing the code to reproduce the results, figures and tables are described in the README file and code in more detailed is described in <em>RCode_xxx</em> files. Please note that the analyses are highly computational intensive and the scripts should be run on a high performance cluster (models and scripts running models and handeling model outputs presented here have been run with 16 cpus and with 50 - 100 GB of RAM). Packages needed are loaded at the beginning of each script. Please ensure that these and all their dependencies have been previously installed. The R code in the files has been carefully commented.</div> </div>
Data for: Theoretical and practical considerations when using retroelement insertions to estimate species trees in the anomaly zone
<p>A potential shortcoming of concatenation methods for species tree estimation is their failure to account for incomplete lineage sorting. Coalescent methods address this problem but make various assumptions that, if violated, can result in worse performance than concatenation. Given the challenges of analyzing DNA sequences with both concatenation and coalescent methods, retroelement insertions (RIs) have emerged as powerful phylogenomic markers for species tree estimation. Here, we show that two recently proposed quartet-based methods, SDPquartets and ASTRAL_BP, are statistically consistent estimators of the unrooted species tree topology under the coalescent when RIs follow a neutral infinite-sites model of mutation and the expected number of new RIs per generation is constant across the species tree. The accuracy of these (and other) methods for inferring species trees from RIs has yet to be assessed on simulated data sets, where the true species tree topology is known. Therefore, we evaluated eight methods given RIs simulated from four model species trees, all of which have short branches and at least three of which are in the anomaly zone. In our simulation study, ASTRAL_BP and SDPquartets always recovered the correct species tree topology when given a sufficiently large number of RIs, as predicted. A distance-based method (ASTRID_BP) and Dollo parsimony also performed well in recovering the species tree topology. In contrast, unordered, polymorphism, and Camin-Sokal parsimony (as well as an approach based on MDC) typically fail to recover the correct species tree topology in anomaly zone situations with more than four ingroup taxa. Of the methods studied, only ASTRAL_BP automatically estimates internal branch lengths (in coalescent units) and support values (i.e., local posterior probabilities). We examined the accuracy of branch length estimation, finding that estimated lengths were accurate for short branches but upwardly biased otherwise. This led us to derive the maximum likelihood (branch length) estimate for when RIs are given as input instead of binary gene trees; this corrected formula produced accurate estimates of branch lengths in our simulation study, provided that a sufficiently large number of RIs were given as input. Lastly, we evaluated the impact of data quantity on species tree estimation by repeating the above experiments with input sizes varying from 100 to 100,000 parsimony-informative RIs. We found that, when given just 1,000 parsimony-informative RIs as input, ASTRAL_BP successfully reconstructed major clades (i.e clades separated by branches >0.3 CUs) with high support and identified rapid radiations (i.e., shorter connected branches), although not their precise branching order. The local posterior probability was effective for controlling false positive branches in these scenarios.</p>
Dataset used in the study "Residential buildings real estate values linked to summer surface thermal anomaly patterns and urban features: the Florence (Italy) case study."
<p>This dataset repository includes eight raster layers (Reference System EPSG:3035 - ETRS89-extended / LAEA Europe), used in the study "Residential buildings real estate values linked to summer surface thermal anomaly patterns and urban features: the Florence (Italy) case study", and obtained by the adaptation of analyses carried out by previous studies (Morabito et al., 2021; Guerri et al., 2021; 2022).</p> <p>Further information regarding the source, study period, and horizontal resolution is available in the attached text file. </p> <p> </p> <p><strong><em>References</em></strong></p> <p>Guerri, G., Crisci, A., Congedo, L., Munafò, M., Morabito, M., <strong>2022</strong>. A functional seasonal thermal hot-spot classification: Focus on industrial sites. Science of The Total Environment 806, 151383.<a href="http://https://doi.org/10.1016/j.scitotenv.2021.151383"> https://doi.org/10.1016/j.scitotenv.2021.151383</a>.</p> <p>Guerri, G., Crisci, A., Messeri, A., Congedo, L., Munafò, M., Morabito, M., <strong>2021</strong>. Thermal Summer Diurnal Hot-Spot Analysis: The Role of Local Urban Features Layers. Remote Sensing 13, 538. <a href="https://doi.org/10.3390/rs13030538">https://doi.org/10.3390/rs13030538</a>.</p> <p>Morabito, M., Crisci, A., Guerri, G., Messeri, A., Congedo, L., Munafò, M., <strong>2021</strong>. Surface Urban Heat Islands in Italian Metropolitan Cities: Tree Cover and Impervious Surface Influences. Science of The Total Environment 751, 142334. <a href="https://doi.org/10.1016/j.scitotenv.2020.142334">https://doi.org/10.1016/j.scitotenv.2020.142334</a>.</p>
Datasets for Simulation-based Anomaly Detection for Multileptons at the LHC
<p>The simulated background and signal data used for a signal model agnostic machine learning search. This search examined the decay of the Higgs boson to leptons working off of LHC data from the Atlas experiment. Details are provided in the paper entitled "Simulation-based Anomaly Detection for Multileptons at the LHC". </p>
InSAR data for "Transcrustal compressible fluid flow explains the Altiplano-Puna deformation anomaly"
<p>InSAR data presented in the paper "Transcrustal compressible fluid flow explains the Altiplano-Puna deformation anomaly". See file "README" for detailed descriptions of each item.</p>
Magnetic Anomaly Characteristics and Magnetic Basement Structure in Changning Area of Southern Sichuan Basin
<p>The data are used in the article "Magnetic Anomaly Characteristics and Magnetic Basement Structure in Changning Area of Southern Sichuan Basin".</p>
SDUST2021GRA: Global marine gravity anomaly model recovered from Ka-band and Ku-band satellite altimeter data
<p>SDUST2021GRA is the global marine gravity anomaly model on a grid of 1′×1′, which is established from the altimeter data of <strong> </strong>Ka-band and Ku-band altimetry satellite including HY-2A. Its spatial coverage is 80°S-80°N. Assessed by the shipborne gravity data, the accuracy of SDUST2021GRA in the global is 2.37 mGal, and that in the open ocean is about 1.5 mGal.</p>
Monthly water storage anomalies, precipitation, and groundwater recharge in two karstic basins, southwest China (2003-2014)
<p>This dataset includes the regionally-averaged monthly hydrological data in two karstic basins, southwest China which are processed or estimated for the manuscript entitled "A novel approach for assessing groundwater recharge by combining GRACE and baseflow with case studies in karst areas of southwest China" by Huang et al., 2022 (submitted to Water Resources Research, Major Revision).</p> <p><strong>Basin description:</strong></p> <p>The Wujiang River Basin (WRB, ~87,900 km<sup>2</sup>, ~70% karstification) and Xijiang River Basin (XRB, ~360,000 km<sup>2</sup>, ~44% karstification) are two typical karstic basins in southwest China. The Wujiang River is the largest tributary in the southern part of the upper Yangtze River. It is originated from the Wumeng Mountain in the Yunnan-Guizhou Plateau and flows from Guizhou to Chongqing. The Xijiang River is the largest tributary of the Pearl River (the largest river in southern China). It is originated from the mountains in eastern Yunnan and flows to Guangxi, Guangdong, and finally into the South China Sea.</p> <p><strong>Data description:</strong></p> <p>1. TWSA (unit: mm in equivalent water thickness) is the average terrestrial water storage anomaly data obtained from three release 6 GRACE (Gravity Recovery and Climate Experiment) mascon solutions, i.e., the Center for Space Research (CSR) at the University of Texas (<a href="http://www2.csr.utexas.edu/grace/RL06_mascons.html">http://www2.csr.utexas.edu/grace/RL06_mascons.html</a>), Jet Propulsion Laboratory (JPL, <a href="https://podaac.jpl.nasa.gov/dataset/TELLUS_GRAC-GRFO_MASCON_CRI_GRID_RL06_V2">https://podaac.jpl.nasa.gov/dataset/TELLUS_GRAC-GRFO_MASCON_CRI_GRID_RL06_V2</a>), and Goddard Space Flight Center (GSFC, <a href="https://earth.gsfc.nasa.gov/geo/data/grace-mascons">https://earth.gsfc.nasa.gov/geo/data/grace-mascons</a>). The anomalies were estimated by removing a mean background value during 2006-2012.</p> <p>2. SMSA and SWSA (unit: mm in equivalent water thickness) are the soil moisture storage anomaly and surface water storage anomaly data based on the model simulations by WGHM (v2.2d) provided by Dr. Hannes Müller Schmied (email: hannes.mueller.schmied@em.uni-frankfurt.de) at Institute of Physical Geography, Goethe-University Frankfurt. The anomalies were estimated by removing a mean background value during 2006-2012.</p> <p>3. Precipitation data (unit: mm/month) were based on the monthly gridded (0.5×0.5 degree) data product obtained from China Meteorological Administration (CMA, https://data.cma.cn/) which was interpolated from ground-based data measured by meteorological stations.</p> <p>4. Groundwater recharge (unit: mm/month) was estimated based on the groundwater budget method, i.e., the summation of groundwater storage change (GWSC) and baseflow. GRACE-based recharge was estimated using the GWSC derived from GRACE TWSA, WGHM-simulated SMSA and SWSA and in situ-based reservoir storage data. Observation-based recharge was estimated using the GWSC based on in situ groundwater-level data and specific yield (or storage coefficient). Both GRACE- and observation-based recharge were estimated using the baseflow separated by a multiple linear regression using the in situ streamflow as predictand, and precipitation and water table depth data as predictors.</p> <p>Time span: 2003-2014. The time lable like "200301" means January, 2003. "200312" means December, 2003.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.