Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Dataset for "The influence of the amount of recycled material on the microstructure and properties of the second generation of single-domain YBCO bulks"
<p>The development of a recycling process for various REBCO materials is crucial considering both environmental sustainability and economic efficiency, particularly in light of the upcoming large-scale applications. In this paper, a novel general recycling process based on chemical dissolution was employed to grow REBCO bulks; recycled material obtained by recycling defective YBCO single-domain bulks was added (15 wt. %, 30 wt. % and 45 wt. %) to raw materials to prepare recycled YBCO precursor powder. Subsequently, recycled single-domain YBCO bulks were produced using Top-Seeded Melt Growth. The waste recycling related to of single-domain bulks growth was chosen, as it represents the most challenging form of waste in the context of REBCO superconductor production. The properties and microstructure of recycled bulks were further analyzed to determine the influence of the amount of recycled material used and compared to commercially produced bulks. Single-domain YBCO bulks were grown successfully from the recycled precursor powder. Furthermore, it was found that their properties could be tuned by varying the amount of the added recycled powder, allowing the use of vast amounts of REBCO waste for the preparation of bulks, when achieving the best possible properties is not essential for a given application. Given that the underlying recycling process is designed to work for all REBCO systems and any form of waste, it has significant implications for the sustainability and cost-effectiveness of REBCO superconductor production. </p>
Dataset - Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks
<p>This data is complementary to the paper by Leijnse et al. 2022 "Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks" <br> https://doi.org/10.5194/nhess-2021-181</p> <p>This data is made available in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE</p> <p>For questions about the data ask: tim.leijnse@deltares.nl</p> <p>For more information about the tool to generate the used synthetic tracks TCWiSE see: <a href="https://www.deltares.nl/en/software/tcwise/">https://www.deltares.nl/en/software/tcwise/</a></p> <p> </p>
Dataset of "Photoelectrochemical generation of H2O2 using hematite (α-Fe2O3) and gas diffusion electrode (GDE)"
<p>In contrast to the industrial-scale production of H2O2 the electrochemical or photoelectrochemical synthesis is environmentally friendly. In the present work, <br>the photoelectrochemical generation of H2O2 was studied by combining the hematite (α-Fe2O3/FTO/glass) photoanode and gas diffusion electrode (GDE) modified by <br>incorporation of tin (II) phthalocyanine (SnPc) in its hydrophilic layer. The experiments were carried out in a photoelectrochemical cell with two compartments <br>separated by a proton exchange membrane under applied bias and AM1.5 irradiation (100 mW/cm2). The generated amount of H2O2 was determined by chemical analysis <br>(visible light spectrophotometry) of the electrolyte. As a tool to determine the efficiency of such a process, the Faradaic efficiency (FE) was calculated. The <br>best configuration used air as an inlet gas for GDE and phosphate buffer (pH 6.4) as an electrolyte in the cathodic compartment. The combination of hematite and <br>GDE (with SnPc) was the most effective in H2O2 photoelectrochemical generation. The highest value of FE was 52.4 % for GDE (O2 reduction to H2O2) and 0.4 % for <br>hematite photoanode (H2O oxidation to H2O2).</p>
Four lipidomics datasets (mouse liver, mouse pancreatic islets, mouse soleus muscle and mouse visceral adipose tissue), generated for the publication Mehl et al., "A multiorgan map of metabolic, signalling, and inflammatory pathways that coordinately control fasting glycemia in mice"
<p>Mehl, Thorens et al present a multiomics study aimiing to<span> identify the pathways that are coordinately regulated in pancreatic </span><span>b</span><span>-cells, muscle, liver, and fat to control fasting glycemia we fed C57Bl/6, DBA/2 and Balb/c mice a regular chow or a high fat diet for 3, 10 and 30 days. We measured fasted glycemia, insulinemia and whole-body insulin resistance. Transcriptomic and lipidomic analysis were used in a data fusion approach to identify organ-specific pathways related to the glycemic levels across all conditions investigated. In pancreatic islets, constant insulinemia despite higher glycemic levels were associated with reduced expression of mRNAs encoding hormone and neurotransmitter receptors as well as OXPHOS, cadherins, integrins and gap junction proteins. Higher glycemia and whole-body insulin resistance were associated, in muscle, with reduced expression of mRNAs encoding insulin signaling proteins and enzymes of the glycolysis, Krebs’ cycle and OXPHOS pathways, as well as endocytosis and exocytosis proteins; in hepatocytes, with lower expression of mRNAs of the insulin signaling pathway, of branched chain amino acid catabolism and of OXPHOS; in adipose tissue, with increased expression of mRNAs of innate immunity and lipid catabolism. These data provide a map of the pathways that are coordinately recruited in the investigated tissues to control fasting glycemia and a resource for further studies of interorgan communication in glucose homeostasis. </span></p>
Molecular datasets from "SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design"
<p>Herein find the molecular datasets from "<a href="https://chemrxiv.org/articles/SMILES-Based_Deep_Generative_Scaffold_Decorator_for_De-Novo_Drug_Design/11638383">SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design</a>". These were generated with SMILES-based scaffold decorator generative models trained with two training sets (DRD2 and ChEMBL). These generative models require a partially-built molecule (scaffold) as input and output several possible completions for each scaffold. Each dataset corresponds to a model trained with the ChEMBL or DRD2 sets, wither multi-step (ms) or single-step (ss) and the provenance of the scaffolds (validation set, or non-dataset).</p> <p>The molecules generated are annotated with a set of descriptors. The DRD2 datasets have the predicted probability of each molecule to be active on DRD2 (p) obtained from a Random Forest model. The ChEMBL model's descriptors are related to the synthesizability of the molecules (see manuscript). Also, the datasets decorated from validation set scaffolds are annotated whether they are part of the validation set (in_validation).</p>
Dataset for the publication entitled "An exact system of generation for face-milled hypoid gears with uniform depth taper: application to hypoid gear drives with high gear ratio"
Open the record for dataset details and reuse information.
Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows
<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, “On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,” IEEE Communications Magazine, pp. 2–8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>
Dataset for the publication "Implementation of an exact completing method of generation for face-milled spiral bevel gears with uniform depth taper"
<p>This dataset contains geometric and graphics data associated with the referenced paper, enabling the reproduction of the conducted research. </p>
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
Dataset generated to evaluate in situ sampling strategies to reconstruct fine-scale ocean currents in the context of SWOT satellite mission (H2020 EuroSea project)
<p><strong>Dataset generated in Subtask 2.3.1 of the H2020 EuroSea project.</strong></p> <ul> <li> <p><em>H2020 EuroSea project:</em><br> The H2020 EuroSea project aims at improving and integrating the European Ocean Observing and Forecasting System (see official website: <a href="https://eurosea.eu/">https://eurosea.eu/</a>). It has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 862626).</p> </li> <li> <p><em>Task 2.3:</em><br> Task 2.3 has the objective to improve the design of multi-platform experiments aimed to validate the Surface Water and Ocean Topography (SWOT) satellite observations with the goal to optimize the utility of these observing platforms. Observing System Simulation Experiments (OSSEs) have been conducted to evaluate different configurations of the in situ observing system, including rosette and underway CTD, gliders, conventional satellite nadir altimetry and velocities from drifters. High-resolution models have been used to simulate the observations and to represent the “ocean truth”. Several methods of reconstruction have been tested: spatio-temporal optimal interpolation, machine-learning techniques, model data assimilation and the MIOST tool. The planned OSSEs are detailed in this public report <a href="https://doi.org/10.3289/eurosea_d2.1">Barceló-Llull et al. (2020)</a> and the complete analysis is available here <a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al. (2022)</a>. Contributors to Task 2.3 are CSIC (Spain), CLS (France), SOCIB (Spain), IMT-Atlantique (France) and Ocean-Next (France).</p> </li> <li> <p><em>Subtask 2.3.1:</em><br> Subtask 2.3.1 aims to evaluate different in situ sampling strategies to reconstruct fine-scale ocean currents (~20 km) in the context of SWOT. An advanced version of the classic optimal interpolation used in field experiments, which considers the spatial and temporal variability of the observations, has been applied to reconstruct different configurations with the objective to evaluate the best sampling strategy to validate SWOT.</p> </li> <li> <p><em>Where?</em><br> The analysis focuses on two regions of interest: (i) the western Mediterranean Sea and (ii) the Subpolar North West Atlantic. In the western Mediterranean Sea, the target area is located within a swath of SWOT, while in the North West Atlantic the region of study includes a crossover of SWOT during the fast-sampling phase.</p> </li> </ul> <p><strong>Report with the full analysis</strong></p> <p>The complete analysis can be found in this report: <a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al. (2022)</a>.</p> <p><strong>Codes for the analysis</strong></p> <p>The codes generated to develop Subtask 2.3.1 can be found on GitHub: <a href="https://github.com/bbarcelollull/EuroSea_subTask_2.3.1">https://github.com/bbarcelollull/EuroSea_subTask_2.3.1</a></p> <p><strong>The dataset</strong></p> <p>The dataset includes:</p> <p>1) Model outputs used to simulate the observations in different configurations in both regions of study. The folder "2D_model_outputs" contains 2D data used to simulate SSH observations for the analysis of the temporal correlation scale (<a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al., 2022</a>, p. 28-42). The folder "3D_model_outputs" contains 3D model outputs used to simulate observations of temperature and salinity. Note that eNATL60 outputs have been interpolated onto a new regular grid. </p> <p>2) Simulated configurations (or sampling strategies) in each region (PKL file format).</p> <p>3) Observations simulated in each configuration in both regions of study. The observations simulated are temperature and salinity. ADCP horizontal velocities are also simulated, however for eNATL60 they will be corrected in the future to account for the rotated original axes. File format: region_configuration_period_model.nc. The folder "SSH" includes the simulated SSH observations for the analysis of the temporal correlation scale (<a href="https://doi.org/10.3289/eurosea_d2.3">Barceló-Llull et al., 2022</a>, p. 28-42).</p> <p>4) Reconstructed fields with the spatio-temporal optimal interpolation. File format: region_configuration_period_model_stOI_Lx_Lt_cd_YYYYMMDDhhmm_var.nc (stOI = spatio-temporal optimal interpolation, Lx = spatial correlation scale, Lt = temporal correlation scale, cd = map on the central date of the sampling, YYYYMMDDhhmm = date and time of the map, var = variable interpolated (temperature and salinity) or the derived variables (dynamic height, geostrophic velocities and the Rossby number)).</p> <p>5) Compared fields (ocean truth from model outputs vs. reconstructed fields) for each region and model (PKL file format).</p> <p> </p>
Dataset for Accessing Cosmic Radiation as an Entropy Source for a Non-Deterministic Random Number Generator
<p>The dataset contains all gathered data from the experiment from Wednesday, March 16, 2022 11:58:41.929 AM UTC+0 (1647431921929) until Sunday, April 3, 2022 1:08:35.353 PM UTC+0 (1648991315353). The experiment was executed during physical presence within the Arctic Circle in Tromsø, Norway 69° 40' 53.117'' N 18° 58' 36.027'' E at 35m elevation above sea level. The dataset was gathered with a prototype [1] based on the CREDO android application [2]. The main research is to use Ultra High Energy Cosmic Rays (UHECR) as an entropy source for a Random Bit Generator (RBG). </p> <p>The associated publication will probably have the title "Accessing Cosmic Radiation as an Entropy Source for a Non-Deterministic Random Number Generator"</p> <p>In order to reproduce the results the SQLite3 database "mrng_arctic_experiment_2022.db" is needed. To get the visual representations of the detections use "image_decoding_and_codesnippets.py" to generate the cleaned (414 detections / ~15MB) or the uncleaned (5567 detections / ~195 MB) dataset. The compressed folder "raw_data_incl_space_weather.7z" contains all raw data as gathered with the MRNG prototype, unprocessed, uncleaned, and unmerged. </p> <p> </p> <p>[1] https://github.com/StefanKutschera/mrng-prototype, visited on 27.03.2023</p> <p>[2] https://github.com/credo-science/credo-detector-android, visited on 27.03.2023</p>
Randomly generated dataset
<p>This dataset is randomly generated using the built-in function from python random.randint(). This csv file contains 2 columns, index and value. Index represents the unique row id and value represents the randomly generated value at each row.</p>
Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.
<p>This repository contains input and ouput files used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives: </p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex -> DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i><index>_<pdbcode></i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. </p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -> FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -> DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 -> DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -> DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>
Artificial Neural Networks-generated Dataset: pH, Total Alkalinity, and Hydrogen Ion Concentration in Ría de Vigo (NW Spain), 1995–2020
<p>This dataset comprises input data from INTECMAR and the predicted outcomes. The variables and their units are as follows:</p> <p>station: 'Station ID [1-6]'</p> <p>year: 'Year [1995-2020]'</p> <p>month: 'Month [1-12]'</p> <p>day: 'Day'</p> <p>latitude: 'Latitude (decimal degrees)'</p> <p>longitude: 'Longitude (decimal degrees)'</p> <p>depth: 'Depth (meters)'</p> <p>temperature: 'Temperature (degrees Celsius)'</p> <p>salinity: 'Salinity (psu)'</p> <p>phosphate: 'Phosphate (umol/kg)'</p> <p>nitrate: 'Nitrate (umol/kg)'</p> <p>silicate: 'Silicate (umol/kg)'</p> <p>cweek: 'Cosine week'</p> <p>sweek: 'Sine week'</p> <p>TA: 'Total Alkalinity predicted (umol/kg)'</p> <p>NTA: 'Normalized Total Alkalinity (umol/kg)'</p> <p>NAT_st: 'Normalized per station Total Alkalinity (umol/kg)'</p> <p>NTA_gl: 'Normalized globally Total Alkalinity (umol/kg)'</p> <p>pHTS_insitu: 'pH insitu (pH units)'</p> <p>HT: 'Hydrogen ion concentration predicted (nmol/kg)'</p> <p> </p> <p>The authors gratefully acknowledge the financial support by the Programa de axudas á etapa predoutoral da Xunta de Galicia (Axencia Galega de Innovación) (Grant nº IN606A-2022/025). F.F.P. and A.V. were supported by REDEIRA (TED2021-132188B-I00) project, funded by MCIN/AEI/10.13039/501100011033. The authors also express their gratitude to the Instituto Tecnolóxico para o Control do Medio Mariño de Galicia (INTECMAR), for the analyses and production of the database used to make predictions.</p>
Dataset and plot generation script for article "Probabilistic short-range forecasts of high precipitation events : optimal decision thresholds and predictability limits" by Francois Bouttier and Hugo Marchal, submitted in Dec 2023.
<p>Dataset and plot generation script for article "Probabilistic short-range forecasts of high precipitation events : optimal decision thresholds and predictability limits" by Francois Bouttier and Hugo Marchal, submitted in NHESS journal in Dec 2023.</p> <p>For further technical details read the file READMEdata in the zipfile. The script MAKEFIG remakes all the figures from the data.</p> <p>For scientific details read the associated article preprint on the NHESS egusphere website.</p>
Dataset and neural network weights to the paper: "Generative diffusion for regional surrogate models from sea-ice simulations"
<p>All the needed code and data to reproduce the results from the paper: "Generative diffusion for regional surrogate models from sea-ice simulations".<br>While most of the code is a frozen clone of the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, this capsule also includes the dataset and neural network weights to train and apply the surrogate models.</p> <p>The <strong>dataset</strong> for training and evaluation can be found at <em>data/nextsim</em>, which includes three different Zarr folders for training/validation/testing. The dataset is based on neXtSIM simulation data and ERA5 forcing data and extracted from the <a href="https://ige-meom-opendap.univ-grenoble-alpes.fr/thredds/catalog/meomopendap/extract/catalog.html">SASIP shared data OpenDAP server</a>:</p> <ul> <li>The neXtSIM simulations were performed by Gauillaume Boutin and published in the paper "<a href="https://doi.org/10.5194/tc-17-617-2023">Arctic sea ice mass balance in a new coupled ice–ocean model using a brittle rheology framework</a>" (Boutin et al., 2023) and available as Zenodo <a href="../records/7277523">dataset</a> (Boutin et al., 2022).</li> <li>The forcing data is based on the ERA5 reanalysis dataset published in the paper: "<a href="https://doi.org/10.1002/qj.3803">The ERA5 global reanalysis</a>" (Hersbach et al., 2020) and available as dataset from the Copernicus Climate Change Service (C3S, Copernicus Climate Change Service, 2023). The here used forcing data is based on the <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels">hourly reanalysis data on single levels</a> and interpolated with nearest neighbors to the curvilinear grid as used in the output from the neXtSIM simulations. <strong>Disclaimer:</strong> The results contain modified Copernicus Climate Change Service information, 2023. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains.</li> </ul> <p>The <strong>neural network weights</strong> are included under <em>data/models </em>and split into weights for the deterministic models and the diffusion models.<br>These neural network weights have been used to generate the results presented in the paper.</p> <p>In this capsule, the <em>notebooks</em> folder includes also the figures used within the paper and additional trajectory data used in the qualitative analysis of the paper.</p> <p>Generally, we recommend to just download the <em>data.tar.gz </em>file and use otherwise the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, since the here included code can be outdated. We further refer to the repository for additional information.</p> <p> </p> <p>Contained in this capsule:</p> <ul> <li>configs.tar.gz: The configuration files for the experiments.</li> <li>data.tar.gz: The dataset and neural network weights.</li> <li>diffusion_nextsim.tar.gz: The main code for the neural network etc.</li> <li>environment.yaml: The anaconda environment file, can be used to install the needed packages.</li> <li>notebooks.tar.gz: The notebooks that were used to create the figures in the paper. The figures from the paper and the data from the qualitative analysis are included as well.</li> <li>readme.md: The readme file from the repository.</li> <li>scripts.tar.gz: The scripts used for the experiments.</li> <li>setup.py: the file to install the <em>diffusion_nextsim</em> package in a python environment.</li> </ul> <p>References:</p> <p>Guillaume Boutin, Heather Regan, Einar Ólason, Laurent Brodeau, Claude Talandier, Camille Lique, & Pierre Rampal. (2022). Data accompanying the article "Arctic sea ice mass balance in a new coupled ice-ocean model using a brittle rheology framework" (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7277523</p> <p>Boutin, G., Ólason, E., Rampal, P., Regan, H., Lique, C., Talandier, C., Brodeau, L., and Ricker, R.: Arctic sea ice mass balance in a new coupled ice–ocean model using a brittle rheology framework, The Cryosphere, 17, 617–638, https://doi.org/10.5194/tc-17-617-2023, 2023.</p> <p>Copernicus Climate Change Service (2023): ERA5 hourly data on single levels from 1940 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI: <a href="https://doi.org/10.24381/cds.adbb2d47">10.24381/cds.adbb2d47</a>.</p> <p>Hersbach H, Bell B, Berrisford P, et al. The ERA5 global reanalysis. <em>Q J R Meteorol Soc</em>. 2020; 146: 1999–2049. <a href="https://doi.org/10.1002/qj.3803">https://doi.org/10.1002/qj.3803</a></p> <p> </p>
Source Data and ambient ozone dataset generated in "Substantially underestimated global health risks of current ozone pollution"
<p>Existing assessments might have underappreciated ozone-related health impacts worldwide. Here our study assesses current global ozone pollution using the high-resolution (0.05°) estimation from a geo-ensemble learning model, with key focuses on population exposure and all-cause mortality burden. Our model demonstrates strong performance, achieving a mean bias of less than -1.5 parts per billion against in-situ measurements. We estimate that 66.2% of the global population is exposed to excess ozone for short term (> 30 days per year), and 94.2% suffers from long-term exposure. Furthermore, severe ozone exposure levels are observed in Cropland areas, particularly over Asia. Importantly, the all-cause ozone-attributable deaths significantly surpass previous recognition from specific diseases worldwide. Notably, mid-latitude Asia (30°N) and the western United States show high mortality burden, contributing substantially to global ozone-attributable deaths. Our study highlights current significant global ozone-related health risks and may benefit the ozone-exposed population in the future.</p>
Global Surface Ozone Concentration Dataset 1990-2017 Generated by Bayesian Maximum Entropy Data Fusion With RAMP Bias Correction
<p>This dataset reports estimates of surface ozone concentration at fine spatial resolution for 1990 to 2017, at 0.5 degree horizontal resolution. Also reported is the variance. Estimates correspond to this paper:</p> <p><span>Becker, J. S.</span><span>, DeLang, M. N., K.-L. Chang, M. L. Serre, O. R. Cooper, <u>H. Wang</u>, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, M. Brauer, and <span>J. J. West</span> (2023) Using Regionalized Air Quality Model Performance and Bayesian Maximum Entropy data fusion to map global surface ozone concentration, <em>Elementa Science of the Anthropocene</em>, 11: 1, doi: 10.1525/elementa.2022.00025.</span></p> <p>The dataset reports estimates of surface ozone for the OSDMA8 metric (the 6-month ozone-season average of the daily maximum 8-hr concentration), estimated through a data fusion of ozone observations from the Tropospheric Ozone Assessment Report (TOAR) database, and output from multiple global atmospheric models. Estimates are created in each year by a combination of M3Fusion to create a multi-model composite, Regional Air Quality Model Performance (RAMP) regional and nonlinear bias correction, and Bayesian Maximum Entropy (BME) data fusion in space and time. The estimates here are the final results using a weighted RAMP bias correction. </p>
HIKARI-2021: Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic
<p>Available datasets from the paper Generating Encrypted Network Traffic for Intrusion Detection Datasets.</p> <p>To produce the dataset follow the technical detail in <a href="https://github.com/andreysfc/generating-encrypted-network">github</a></p>
Dataset of 30 energy customers with flexibility data, and distributed generation, considering residential, small commerce, large commerce, and industrial customers
<p>The dataset has 30 customers: ten residential, ten small commerce, five large commerce, and five industrial customers. The combination of several energy customer types allows the creation of a dataset with different types of consumption profiles, generation, and flexibility, and, therefore, different values of participation in demand response events.</p> <p>The residential profiles of the considered customers use the data available in the Working Group on Intelligent Data Mining and Analysis (IDMA): https://site.ieee.org/pes-iss/data-sets/</p> <p>The values represent a week period using 15 minutes reading periods. All the values are expressed in kWh and the matrixes were created as [customer x time_period].</p> <p> </p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.