Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,300
datasets available to search
ShareScore release 0.7.1
Dataset results
9,300 results for “Detection”
Raw Particle Number Size-Distribution Data of twin-DMPS equipped with two CPCs for nanoparticle detection for SMEAR II station, Hyytiälä, Finland, Spring 2017
<p>Raw size-Distribution data from twin-DMPS system (Aalto et al., 2001), where the nano-DMA (measuring up to 40 nm, short Hauke type DMA) is quipped with two detectors:<br> a TSI 3776 and a modified Airmodus A20 (Kangasluoma et al., 2015)</p> <p>Data acquired during in March-May 2017 at the SMEAR II station in Hyytiälä, Finland.<br> Data associated with the publication Stolzenburg, Laurila et al. (2023), Atmos. Meas. Techn., "Improved counting statistics of an ultrafine differential mobility particle size spectrometer system"</p> <p>Files DMYYDDMM_A20.Dat contain the raw DMPS data, with YYMMDD indicating the day of the measurement.<br> Data are provided alternating between data acquired with the nano-DMA and with the long-DMA, on a scan by scan basis.<br> First line of each scan cycle (for both DMAs) always indicates the start and end times of the voltage scan.<br> Second line gives the parameters related to the DMPS as given below:<br> (sheath flow in [l per min], aerosol flow in [l per min], DMA inner electrode diameter in [m], DMA outer electrode diameter in [m], DMA classification length in [m], other parameters)<br> Following lines give<br> (for long-DMA): set voltage at DMA [in V], concentration measured by TSI3772 in [per cm3]<br> (for nano_DMA): et voltage at DMA [in V], concentration measured by TSI 3776 in [per cm3], concentration measured by mod. Airmodus A20 in [per cm3]</p> <p>File dmps_data_format_specifier.text gives a conversion from voltage to diameter and indicates the measurement time at each voltage during the stepping of the DMPS.<br> Needs to be used to convert measured concentrations in counts per set-interval.</p> <p>Files GR_J_overview.xlsx gives size-distribution derived quantities during that campaign.<br> Header defines Date, Growth Rate and Formation Rate measured at different sizes [in nm] and by the two different CPCs connected to the nano-DMA.<br> Growth rates in [nm per h], formation rate in [per cm3 per s].</p> <p>Other data related to the campaign can be obtained from the corresponding author upon reasonable request.<br> juha.kangasluoma@helsinki.fi</p> <p>References:</p> <p>Stolzenburg, Laurila et al. "Improved counting statistics of an ultrafine differential mobility particle size spectrometer system",<br> Atmos. Meas. Techn., in press, 2023</p> <p>Aalto et al., "Physical characterization of aerosol particles during nucleation events",<br> Tellus B, vol. 53, pp. 344-358, 2001</p> <p>Kangasluoma et al., "Sub-3 nm Particle Detection with Commercial TSI 3772 and Airmodus A20 Fine Condensation Particle Counters",<br> Aerosol Sci. Techn., vol. 49, pp. 674-681, 2015</p>
Detecting cosmic voids via maps of geometric optics parameters
<ul> <li>lensing-ddbb4ac.pdf - research data in pdf format</li> <li>void_matches*.dat - plain text results files corresponding to Table 3 and Figures 2, 4, 6, 8.</li> <li>lensing-ddbb4ac-journal.tar.gz - source package for producing the article pdf, together with the reproducibility package, but without the git history; appropriate for ArXiv</li> <li>lensing-ddbb4ac-git.bundle - git source package that can be unbundled with 'git clone lensing-e4f7af0-git.bundle' and used for reproducibility: to download data, do calculations, analyse them, plot them and produce the research data pdf</li> <li>software-ddbb4ac.tar.gz - this should contain all the software, apart from a minimal POSIX-compatible system and LaTeX packages, needed for compiling and installing the software used in producing this work</li> <li>lensing-ddbb4ac-snapshot.tar.gz - source files of the project; these should be enough, provided that external software packages can be downloaded, to reproduce the full project</li> </ul> <p>The authors grant a perpetual, non-exclusive licence to distribute this pdf preprint.</p> <p>All the other materials here are free-licensed, as stated in the individual files and packages.</p>
Fish tag data remotely detected using whole stream antennas or hand held tag readers in the Kuparuk, Itkilik, and Sagavanirktok drainages near Toolik Field Station, Alaska, from 2010 to 2017
From 2009 to 2017, the FISHSCAPE Project (grant numbers 1719267, 1417754, and 0902153), based at Toolik Field Station, has monitored physical, chemical, and biological parameters within three watersheds: The Kuparuk (including Toolik Lake and Toolik outlet stream); The Sagavanirktok (primarily Oksrukuyik Creek, but also including sections of the Ailish and Atigun Rivers and the Galbraith Lakes); and The Itkillik (primarily the I-Minus outlet stream, a tributary that that feeds into the Itkilik River). Target species were primarily Arctic grayling and Lake trout, although Arctic char, Burbot, Dolly varden, round whitefish, and slimey sculpin were also captured. This file contains the detectioned fish tags using whole stream or hand-held antennas in the three watersheds. We had no field season in 2014 and thus did not deploy antennaes. Fish were tagged with Passive Integrated Transponder (PIT) tags which can be read with a whole stream antenna to track the migration of the fish, predominately Arctic grayling, throughout the systems. Fish tags detected with a handheld readers are designated in Site ID as "XXX_capture". For "capture" fish time is arbitraily set at '7:00:00'' of the day of capture and tagging because actual time was not recorded. The individual fish data (date, tag number, length, weight, species) associated with the tag can be found in the 2009-2017_FISHSCAPE_fish_tagging file.
Go-nogo categorization and detection task
Open the record for dataset details and reuse information.
Wind tunnel distributed temperature sensing with actively heated fibers and microstructures for detecting wind direction
<p>Wind tunnel tests were performed using distributed temperature sensing with actively heated fibers that had microstructures attached in opposing directions on neighboring fibers. These microstructures created a temperature difference between fibers that depended on wind speed, providing a prototype for distributed sensing of wind direction. These data are connected to a publication detailing this work and method, <a href="https://www.atmos-meas-tech-discuss.net/amt-2019-188/">"Distributed observations of wind direction using microstructures attached to actively heated fiber-optic cables"</a>.</p> <p>Data are stored in a netcdf format and includes the instrument reported temperature ('instr_temp') and calibrated temperature ('cal_temp') with the various parameters tested in the linked paper available as coordinates, labeled along an 'expname' dimension.</p> <p>The included ipython notebooks provide examples and explanations for using these laboratory data.</p>
Anomaly detection in the Zwicky Transient Facility DR3
<p>The feature data set extracted from <a href="https://www.ztf.caltech.edu/page/dr3">ZTF DR3</a> light curves. It was used in <a href="https://arxiv.org/abs/2012.01419">Malanchev et al. 2020</a> to detect anomalous astrophysical sources in ZTF data. </p> <p>"feature_XXX.dat" files contain object-ordered light curve feature data, every object is built on 42 feature values, which are encoded as little endian single precision IEEE-754 float (32bit float) numbers. Feature code-names are the same for all three data sets and are listed in plain text files "feature_XXX.name", one code-name per line. "oid_XXX.dat" files contain ZTF DR object identifiers encoded as little endian 64-bit unsigned integer numbers. "oid_XXX.dat" and "feature_XXX.dat" have same object order, for example the first 8 bytes of "oid_m31.dat" files contain the OID of the ZTF DR3 light curve which feature are presented in the first 168 bytes of "feature_m31.dat" file. "m31", "deep" and "disk" denote different ZTF fields and contain 57 546, 406 611, 1 790 565 objects. Note that observations between 58194 ≤ MJD ≤ 58483 are used, see <a href="https://doi.org/10.1093/mnras/stab316">the paper</a> for field and features details.</p> <p>The sample Python code to access the data as Numpy arrays:</p> <pre><code class="language-python">import numpy as np oid = np.memmap('oid_m31.dat', mode='r', dtype=np.uint64) with open('feature_m31.name') as f: names = f.read().split() dtype = [(name, np.float32) for name in names] feature = np.memmap('feature_m31.dat', mode='r', dtype=dtype, shape=oid.shape) idx = np.argmax(feature['amplitude']) print('Object {} has maximum amplitude {:.3f}'.format(oid[idx], feature['amplitude'][idx]))</code></pre> <p> </p>
Smartbay Marine Species Object Detection Training dataset
<h1>Training dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Species Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using species names.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 training dataset format. </p>
Smartbay Marine Types Object Detection Training dataset
<h1>Training Dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Types Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using broad "Marine Type" classes.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 trainign dataset format. </p>
MaDroid: A Maliciousness-aware Multifeatured Dataset for Detecting Android Malware
<p>MaDroid is a maliciousness-aware multifeatured dataset of system calls focused on APK anomaly detection. The dataset includes 50,429 well-marked normal and abnormal system call sequences, with 24,789 and 25,640 sets of normal and abnormal sequences, respectively, for a total of 1.1 billion system call feature information. Each APK is labeled with the latest VT checksum information, and the sequence data includes 81 groups of system calls, system call parameters, and return values. The size of the whole dataset is 457 GB (19 GB after compression), including 236 GB of malicious system call sequence data. The APKs from which the system call feature sequences are derived cover mobile apps of different types released at different times in the past 14 years (2010-2023), covering 10 mainstream app markets, including Google Play, PlayDrone, Anzhi, etc. The APKs are also used as the source of the system call feature sequences, and the system call sequence data is used as the source of the APKs. anzhi, etc. We store the source code and dataset in two open platforms, GitHub and Zenodo, respectively.</p><h2>DataSet</h2><p>Release address: <a href="http://doi.org/10.5281/zenodo.7997398">http://doi.org/10.5281/zenodo.7997398</a></p><ul><li>The dataset consists of two classifications, Normal and Malware, with a total of 21 zip files. The installation files of each sequence come from 10 application markets such as Google Play, PlayDrone, Anzhi, etc. The Malware classification contains information about the running system call sequences of some APKs in the Drebin dataset.</li><li>RF, MLP and GBDT models were used to establish benchmarks for the dataset, the use of the models can be found in the source code.</li><li>The file `merge_all_csv_count_online_check_replenish.csv` is the dataset APK information. We provide APK name (SHA256 name for APK only), classification, APK capacity, number of sequences, log capacity, CVT, OVT value, check time, etc.</li></ul><h2>Source Code</h2><p>Release address: <a href="https://github.com/HNUSystemsLab/MaDroid">https://github.com/HNUSystemsLab/MaDroid</a></p><ul><li>The released source code contains two folders, `Source_Code` and `ml_metadata`. Where `Source_Code` is the automated framework for data collection, the tool chain and some notes on the structure of the source files. `ml_metadata` contains the metadata used for machine learning, the partitioned data on which the article builds its benchmark.</li><li>The automation framework is described in detail in the `Readme.md` document in the `Source_Code` directory. It consists of four main parts: environment requirements, program structure, quick start (working steps), model training and evaluation (including training and evaluation). It describes in detail the preparation of the environment, the data import method, the functional description of each file in the source code directory, the working principle of model training and evaluation, and other related contents.</li></ul><h2>Tips: </h2><ul><li>MAS is another name of MaDroid, the content shown here is the final version of "Readme.md".</li><li>A Large-scale Multi-feature Dataset for Anomaly Detection of Mobile Applications, which is the name of the document during our experiment.</li></ul>
Detecting local variations across metazoan communities in backreef depressions of Reunion Island (Mascarene Archipelago) through environmental DNA survey
<p>The back-reef depressions, or lagoons, of Reunion Island (western Indian Ocean) host a high abundance of organisms living amongst the coral reefs and are critical sites for artisanal fishing, tourism, and shoreline stability for the island. Over time, increasing degradation of Reunionese reefs has been observed due to overexploitation, beach erosion and eutrophication. Efforts to mitigate the impact of these pressures on aquatic organisms include biodiversity surveys primarily performed through visual censuses that can be logistically complex and may unintentionally overlook organisms. Surveys integrating environmental DNA (eDNA) collections have provided rapid biodiversity assessments, while helping to circumvent some limitations of visual surveys. The present study describes the results of an exploratory eDNA survey, which aims to characterize metazoan communities of four Reunionese lagoons located along the west coast of the island. As eDNA surveys first require deliberate study design and optimization for each new context, we sought to establish a modernized workflow implementing specialized equipment to collect and preserve samples to facilitate future studies in these lagoons. During the austral summer of 2023, samples were pumped directly from surface and bottom depths at each site through self-preserving filters which were then processed for DNA metabarcoding using regions of the 12S ribosomal RNA (12S), small ribosomal subunit 18S (18S) and Cytochrome Oxidase I (COI) genes. The survey detected high species richness that varied by site, and in a single collection period, recovered the presence of 60 teleost families and numerous invertebrate taxa, including members of the coral faunal community that are less studied in Reunion. Distinct biological communities were observed at each site, and within a single lagoon, suggesting that these differences are due to site-specific factors (e.g., environmental variables, geographic distance, etc.). Although continued protocol optimization is needed, the present findings demonstrate the successful application of an eDNA-based survey for biodiversity assessment within Reunionese lagoons.</p>
Real-time black ice detection using YOLOX on drone
<p><strong>Detailed Info:</strong> https://github.com/hsh060824/blackice-drone-dataset</p> <p> </p> <p><strong>Dataset Type</strong>: Object Detection Dataset (with bounding boxes)</p> <p> </p> <p><strong>Overview</strong></p> <p>Road safety during winter months remains a critical concern due to the elusive nature of black ice, a thin layer of ice that forms on road surfaces, making it challenging for drivers to identify and navigate safely. In an effort to address this issue, our research team at Cheongshim International Academy (CSIA) has conducted extensive studies on real-time black ice detection utilizing YOLOX, a state-of-the-art object detection algorithm, deployed on drones. As a significant contribution to the research community, we are pleased to share our meticulously curated image dataset, which encapsulates diverse scenarios and conditions representative of real-world black ice occurrences.</p> <p> </p> <p><strong>Background</strong></p> <p>Black ice poses a significant threat to road safety, especially during winter, as it is often challenging for drivers to detect, leading to increased risks of accidents and hazardous road conditions. Our dataset aims to fill the gap in existing resources by providing a comprehensive collection of images showcasing various instances of black ice under different environmental conditions. The dataset covers diverse scenarios, including different lighting conditions, road surfaces, and black ice formations, making it a valuable resource for developing and testing robust black ice detection models.</p> <p> </p> <p><strong>Significances of the Dataset</strong></p> <p>The significance of this dataset lies in its potential to advance the development of effective black ice detection algorithms. By sharing our dataset with the research community, we aim to facilitate the creation of more accurate and reliable models for real-time detection of black ice using drone technology. The dataset includes annotations in COCO format, providing detailed information about the location and characteristics of black ice instances in each image.</p> <p> </p> <p><strong>Categorization</strong></p> <p>In our pursuit of advancing the field of computer vision and contributing to ongoing research endeavors, we proudly introduce three distinct image datasets meticulously curated by our research team. These datasets, categorized as "White," "Black," and "Outdoors (OD)," cater to unique scenarios and are designed to fuel the development of specialized models addressing specific challenges in visual recognition.</p> <p> </p> <p><strong>White Dataset: </strong></p> <ul> <li><strong>Composition:</strong> This dataset comprises 413 images, each meticulously annotated with an average of 1.1 annotations per image, depicting the unique optical characteristics of black ice.</li> <li><strong>Properties:</strong> The average proportion of instance pixel area is 3.16%, emphasizing the subtlety of the black ice formations. The average image brightness is measured at 149.358.</li> <li><strong>Capture Environment:</strong> The images were taken in controlled indoor laboratory conditions, ensuring consistency and repeatability.</li> <li><strong>Creation Method:</strong> The dataset was generated by cooling asphalt samples in a freezer to temperatures ranging from -4°C to -20°C. Subsequently, 4°C water was sprayed onto the sample surfaces, creating black ice. The dataset captures the optical properties of black ice, showcasing its interaction with light.</li> <li><strong>Significance: </strong>Valuable for highlighting the optical characteristics of black ice, enhancing model accuracy in well-lit scenarios.</li> </ul> <p> </p> <p><strong>Black Dataset: </strong></p> <ul> <li><strong>Composition:</strong> This dataset comprises 814 images, with a detailed annotation structure averaging 3.5 annotations per image, showcasing the challenges of recognition in low-light conditions.</li> <li><strong>Properties:</strong> The average proportion of instance pixel area is notably higher at 12.37%, reflecting the complex and varied formations of black ice. The average image brightness is measured at 123.028.</li> <li><strong>Capture Environment:</strong> Similar to the White Dataset, images were captured in a controlled indoor laboratory environment. Asphalt pelt was placed under the black iced asphalt pieces to replicate realistic scenarios.</li> <li><strong>Creation Method:</strong> The dataset creation involved the same process of cooling asphalt samples, followed by spraying water to create black ice. To simulate real-world conditions, asphalt pelt was used as a background, and various shapes of black ice were randomly placed in each image.</li> <li><strong>Significance:</strong> Realistic emulation of black ice using backgrounds made up of asphalt pelts, providing essential drark images for robust model training.</li> </ul> <p> </p> <p><strong>Outdoor (OD) Dataset</strong></p> <ul> <li><strong>Composition:</strong> This dataset is the most extensive, consisting of 1624 images, with an average of 1.5 annotations per image, capturing the challenges of recognizing black ice in outdoor winter conditions.</li> <li><strong>Properties:</strong> The average proportion of instance pixel area is 12.34%, mirroring the complexity of real-world outdoor scenarios. The average image brightness is significantly lower at 56.575.</li> <li><strong>Capture Environment:</strong> Unlike the indoor datasets, the OD dataset was captured outdoors in winter conditions where black ice naturally forms.</li> <li><strong>Creation Method:</strong> Black ice was created on the asphalt road of Cheongsim International High School by spraying +4°C water onto the surface. DJI Tello's built-in camera was used for capturing images from various angles, simulating drone-like perspectives. This dataset is designed to closely replicate real-world scenarios, providing a valuable resource for training models for outdoor applications.</li> <li><strong>Significance: </strong>Represents real-world outdoor scenarios, offering a unique perspective for developing models capable of handling diverse and challenging conditions.</li> </ul> <p> </p> <p><strong>Cameras: </strong></p> <ul> <li>iPhone SE2 (Apple, California)</li> <li>iPhone SE3 (Apple, California)</li> <li>iPhone 12 (Apple, California)</li> <li>iPhone 14 Pro (Apple, California)</li> <li>Q9 (LG Electronics, Seoul, Korea)</li> <li>V30 (LG Electronics, Seoul, Korea)</li> <li>Tello (DJI, ShenZhen, China)</li> </ul>
CloneCorp: Cross-language clone detection dataset
<p>Data set of mobile apps and code examples to evaluate clone detection algorithms across languages (Kotlin, Swift, and Dart)</p>
Cherri - Accurate detection of functional RNA-RNA interactions sites
<p><strong>CheRRI</strong> - Pipeline for the Identification of putative RNA-RNA interaction sites.</p> <p> </p> <p>This repository contains all CheRRI's models computed and mentioned in the content.txt, listing data and their descriptions. All models can be used to classify interaction sites in CheRRI's eval mode.</p> <p> </p> <p>The source code for CheRRI is avalbile on <a href="https://github.com/BackofenLab/Cherri#install-cherri-conda-package">GitHub</a> and can be cited using this Software Heritage citation:</p> <ul> <li><span>Müller T, Mautner S, Videm P, Eggenhofer F, Raden M, Backofen R (2024) CheRRI - Accurate classification of the biological relevance of putative RNA-RNA interaction sites (Version 0.8). [Computer software]. Software Heritage, <a href="https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri">https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri</a></span></li> </ul> <div> <div> <div> <p>The pipeline contains Machine Learning segments which were annotated using DOME:</p> </div> </div> </div> <ul> <li><span><a href="https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25">https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25</a></span></li> </ul>
Supporting Data - Sentinel-1 Detection of Ice Slabs on the Greenland Ice Sheet
<p>This dataset contains supporting data accompanying Culberg, R., Michaelides, R. J., and Miller, J. Z.: Sentinel-1 Detection of Ice Slabs on the Greenland Ice Sheet, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2023-2652">https://doi.org/10.5194/egusphere-2023-2652</a>, 2023. The final accepted manuscript will be linked via the same preprint server at the time of publication. The dataset contains the following files:</p> <ul> <li>Sentinel-1 HV and HV/HH backscatter mosaics of the Greenland Ice Sheet formed using data from 1 Oct 2016 - 30 April 2017.</li> <li>Estimated average annual summer melt extent between 1 Nov 2014 and 31 Aug 2020, detected using seasonal variations in Sentinel-1 HH backscatter.</li> <li>The firn aquifer extent over Greenland derived from Sentinel-1 in Brangers et al. (2020), reprojected to EPSG:3413.</li> <li>The ice mask used in the study, derived from the BedMachine Greenland ice mask.</li> <li>The training and validation datasets derived from the Jullien et al. (2023) ice slabs detections from ice penetrating radar data that were used to optimize ice slab detection thresholds for the Sentinel-1 backscatter mosaics. </li> </ul>
Dataset for the paper "Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset"
<p>We present a large-scale anomaly detection dataset collected from IBM Cloud's Console over approximately 4.5 months. This high-dimensional dataset captures telemetry data from multiple data centers, specifically designed to aid researchers in developing and benchmarking anomaly detection methods in large-scale cloud environments. It contains 39,365 entries, each representing a 5-minute interval, with 117,448 features/attributes, as interval_start is used as the index. The dataset includes detailed information on request counts, HTTP response codes, and various aggregated statistics. The dataset also includes labeled anomaly events identified through IBM's internal monitoring tools, providing a comprehensive resource for real-world anomaly detection research and evaluation.</p> <p><strong>File Descriptions</strong></p> <ul> <li><code>location_downtime.csv</code> - Details planned and unplanned downtimes for IBM Cloud data centers, including start and end times in ISO 8601 format.</li> <li><code>unpivoted_data.parquet</code> - Contains raw telemetry data with 413 million+ rows, covering details like location, HTTP status codes, request types, and aggregated statistics (min, max, median response times).</li> <li><code>anomaly_windows.csv</code> - Ground truth for anomalies, listing start and end times of recorded anomalies, categorized by source (Issue Tracker, Instant Messenger, Test Log).</li> <li><code>pivoted_data_all.parquet</code> - Pivoted version of the telemetry dataset with 39,365 rows and 117,449 columns, including aggregated statistics across multiple metrics and intervals.</li> <li><code>demo/demo.[ipynb|html]</code>: This demo file provides examples of how to access data in the Parquet files, available in Jupyter Notebook (<code>.ipynb</code>) and HTML (<code>.html</code>) formats, respectively.</li> </ul> <p>Further details of the dataset can be found in <strong>Appendix B: Dataset Characteristics</strong> of the <a href="https://arxiv.org/abs/2411.09047">paper</a> titled <strong><em>"Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset."</em></strong> Sample code for training anomaly detectors using this data is provided in <a href="https://doi.org/10.5281/zenodo.14598119" target="_blank" rel="noopener">this package</a>.</p> <p> </p> <p>When using the dataset, please cite it as follows:</p> <pre><code>@misc{islam2024anomaly,</code><br><code> title={Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset}, </code><br><code> author={Mohammad Saiful Islam and Mohamed Sami Rakha and William Pourmajidi and Janakan Sivaloganathan and John Steinbacher and Andriy Miranskyy},</code><br><code> year={2024},</code><br><code> eprint={2411.09047},</code><br><code> archivePrefix={arXiv},</code><br><code> url={https://arxiv.org/abs/2411.09047}</code><br><code>}</code></pre> <p> </p>
Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows
<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, “On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,” IEEE Communications Magazine, pp. 2–8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>
Labeled Images at OBSEA for Object Detection Algorithms
<p>Images from OBSEA underwater cameras labeled with marine species to train AI-based Object Detection algorithms.</p>
Dataset: Label-free detection of methicillin resistance in Staphylococcus aureus using different Raman-spectroscopy approaches
<p>This is the dataset accompanying the submission of the manuscript: Label-free detection of methicillin resistance in Staphylococcus aureus using different Raman-spectroscopy approaches in the journal Microbiology Spectrum.</p> <p>The data description is the following:</p> <p>Strains<br> 16859MRSA= Strain AUSTR-07-16859 MRSA<br> 16859MSSA= Strain AUSTR-07-16859 MSSA<br> CC8MRSA= Strain 08V15773<br> CC8MSSA= Strain MRSA2010-174<br> AUSTR05MRSA= Strain AUSTR-05-15441 MRSA<br> AUSTR05MSSA= Strain AUSTR-05-15441 MSSA<br> CC361MRSA= Strain UAE-Abu Dhabi-020<br> CC361MSSA= Strain UAE-Dubai-80-MS 1368.9/09</p> <p>Datasets<br> UVRR: UV-Resonance Raman with 244 nm excitation on bulk samples, calibration standard Polystyrene, measurements were time series of 10 consecutive spectra, for each strain and batch 25 time series were collected from 3 different slides<br> 532nm: Single cell analysis with 532nm excitation, calibration standard 4AAP, one spectrum per bacterial cell was collected<br> 785nm: Bulk analysis of bacterial colonies using 785 nm excitation and a Raman fibre probe, calibration standard 4AAP, bulk analysis, individual spectra of colonies were collected</p> <p>Data structure is in the metadata files.<br> Individual spectra are in the folders sorted by the date they were measured.</p>
OPSSAT-AD - anomaly detection dataset for satellite telemetry
<p>This is the AI-ready benchmark dataset (OPSSAT-AD) containing the telemetry data acquired on board OPS-SAT---a CubeSat mission that has been operated by the European Space Agency.</p> <p>It is accompanied by the paper with baseline results obtained using 30 supervised and unsupervised classic and deep machine learning algorithms for anomaly detection. They were trained and validated using the training-test dataset split introduced in this work, and we present a suggested set of quality metrics that should always be calculated to confront the new algorithms for anomaly detection while exploiting OPSSAT-AD. We believe that this work may become an important step toward building a fair, reproducible, and objective validation procedure that can be used to quantify the capabilities of the emerging anomaly detection techniques in an unbiased and fully transparent way.</p> <p>The included files are:</p> <ul> <li><code>segments.csv</code> with the acquired telemetry signals from ESA OPS-SAT aircraft,</li> <li><code>dataset.csv</code> with the extracted, synthetic features are computed for each manually split and labeled telemetry segment.</li> <li>code files for data processing and example modeliing (<code>dataset_generator.ipynb</code> for data processing, <code>modeling_examples.ipynb</code> with simple examples, <code>requirements.txt</code>- with details on Python configuration, and the <code>LICENSE</code> file)</li> </ul> <p> </p> <p>Please have a look at our two papers commenting on this dataset:</p> <ul> <li>The benchmark paper with results of 30 supervised and unsupervised anomaly detection models for this collection:<br>Ruszczak, B., Kotowski. K., Nalepa, J., Evans, D.:<strong> The OPS-SAT benchmark for detecting anomalies in satellite telemetry, 2024</strong>, <a href="https://arxiv.org/abs/2407.04730" target="_blank" rel="noopener">preprint arxiv: 2407.04730</a>,</li> <li>the conference paper in which we presented some preliminary results for this dataset:<br>Ruszczak, B., Kotowski. K., Andrzejewski, J., et al.: (2023). Machine Learning Detects Anomalies in OPS-SAT Telemetry. Computational Science – ICCS 2023. LNCS, vol 14073. Springer, Cham, <a href="https://doi.org/10.1007/978-3-031-35995-8_21">DOI:10.1007/978-3-031-35995-8_21</a>.</li> </ul>
Chemical sensors for fire detection and nuisance rejection under EN-5420 standard conditions and reduced-scale chamber
<p>The dataset was acquired using a gas sensor array placed in the celling of a validated standard fire room (240 m3) located in Minimax Company. The dataset includes measurements of three different campaigns that were performed over 15 months. The dataset includes standard EN-54 smoldering fires and non-standard smoldering fires (such as plastic fires; PVC, cables Fire). In order to generate scenarios that may result in false-positive alarms when gas sensors are used, different nuisance experiments were also performed (such as cleaners, and air fresheners). Additionally, an additional measurement campaign was performed in a small chamber. The small-scale experiments dataset includes scale-down replicates of the fire and nuisances experiments performed in the standard fire room (EN-54 smoldering fire experiments, non-standard fires, and nuisance experiments).</p> <p>Citation request: Ana Solórzano et al, Early fire detection based on gas sensor arrays: Multivariate calibration and validation, Sensors and Actuators B: Chemical, 2021, <a href="https://doi.org/10.1016/j.snb.2021.130961">https://doi.org/10.1016/j.snb.2021.130961</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.