Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,943 results for “Machine learning”

Learn how ShareScore rates datasets ↗
zenodo56/100

Sample data for "Machine learning for large-scale forecasting"

<p>This dataset includes sample data for the Netherlands to run the machine learning baseline as described in the paper titled <em>Machine learning for large-scale crop yield forecasting</em>, accessible at&nbsp;<a href="https://doi.org/10.1016/j.agsy.2020.103016">https://doi.org/10.1016/j.agsy.2020.103016</a>.&nbsp;The software implementation of the machine learning baseline is available at:&nbsp;<a href="https://github.com/BigDataWUR/MLforCropYieldForecasting">https://github.com/BigDataWUR/MLforCropYieldForecasting</a>.</p> <p><strong>Notes:</strong></p> <p>The NUTS classification (Nomenclature of territorial units for statistics) is a hierarchical system for dividing up the economic territory of the EU and the UK (see Eurostat, 2016) for more details).</p> <p>Data</p> <p>The dataset consists of 11 CSV files. They are formatted to work as sample inputs to the machine learning baseline.</p> <ol> <li><strong>Crop Area Fractions </strong>(NUTS2, NUTS1):&nbsp;We aggregated the predictions of the machine learning baseline from NUTS2&nbsp;to national (NUTS0) level&nbsp;by weighting them on the modeled crop area. Cerrani and L&oacute;pez Lozano (2017) have described in detail the algorithm used to model crop areas for different NUTS levels. The data comes from the MARS Crop Yield Forecasting System (MCYFS) of European Commission&#39;s Joint Research Centre (JRC) (see Lecerf et al., 2019).</li> <li><strong>Centroids (NUTS2)</strong>: Data includes latitude, longitude and distance to coast of the centroids of NUTS2 regions.</li> <li><strong>Meteo Daily Data and Meteo Dekadal Data </strong>(NUTS2):&nbsp;The data comes from MCYFS&nbsp;(see EC-JRC, 2020). By default, the implementation uses daily data.</li> <li><strong>Remote Sensing Data</strong> (NUTS2, see Copernicus Global Land Service, 2020): Data includes fraction of absorbed photosynthetically active radiation (FAPAR) aggregated to NUTS2.</li> <li><strong>Soil Data</strong>: Data includes soil moisture information that can be used to calculate soil water holding capacity. The data comes from MCYFS (see Lecerf et al., 2019).</li> <li><strong>WOFOST data </strong>(NUTS2): The World Food Studies (WOFOST) crop model (van Diepen et al., 1989; Supit et al., 1994; de Wit et al.&nbsp; 2019) is a simulation model for the quantitative analysis of the growth and production of annual field crops. It is a mechanistic, dynamic model that explains daily crop growth on the basis of the underlying processes, such as photosynthesis, respiration and how these processes are influenced by environmental conditions.&nbsp;The crop simulation is fed by weather, soil and crop data. Observed meteorological data is interpolated on a regular 25 km grid using a method based on the distance, altitude and climatic region similarity between the center of grid cells and weather stations (see Van der Goot, 1998). WOFOST runs on the intersection between the 25 km meteorological grid and soil units based on the European soil map (http://esdac.jrc.ec.europa.eu/). In order to have the output data aggregated to administrative regions such as countries or provinces, simulation units are further intersected with the boundaries of these regions. The outputs at soil unit (STU) level are aggregated to grid level in an area weighted manner. Gridded simulations are aggregated to lowest NUTS level 3 considering the arable land area of each grid, derived from GLOBCOVER and CORINE Land Cover (Cerrani and Lopez Lozano, 2017). From NUTS3 to higher levels, crop area fractions for the current year, retrieved from Eurostat, are used to weight and aggregate the output (Cerrani and Lopez Lozano, 2017).</li> <li><strong>GAES data</strong>: GAES data includes agro-climatic features of regions, such as&nbsp;elevation and slope (from USGS-EROS, 2021), field size (from&nbsp;Lesiv et al., 2019), irrigated (crop) areas (from&nbsp;EC-JRC, 2020) and crop areas&nbsp;(from&nbsp;EC-JRC, 2020).</li> <li><strong>National yield statistics&nbsp;</strong>(NUTS0): These are the official Eurostat national yield statistics (Eurostat, 2020a).&nbsp;We used these yield statistics as reference to compare&nbsp;the machine learning predictions aggregated to NUTS0 and the actual MCYFS forecasts (see van der Velde and Nisini, 2019).</li> <li><strong>Regional yield statistics&nbsp;</strong>(NUTS2): We used NUTS2 yield statistics&nbsp;as labels to train and evaluate machine learning algorithms. We got NUTS2&nbsp;yield statistics from&nbsp;The Central Bureau of Statistics (CBS) of the Netherlands&nbsp;(NL-CBS, 2020).</li> <li><strong>Past MCYFS Yield Forecasts&nbsp;</strong>(NUTS0): These are actual forecasts made by MCYFS in the past (see van der Velde and Nisini, 2019). We used the official Eurostat national yield statistics (see point 7 above) as the reference to compare the machine learning predictions aggregated to NUTS0 and&nbsp;MCYFS forecasts.</li> </ol> <p><strong>Crop ID and name mapping</strong></p> <p>2 : grain maize</p> <p>6 : sugar beets</p> <p>7 : potatoes</p> <p>90 : soft wheat</p> <p>93 : sunflower</p> <p>95 : spring barley</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We would like to thank S. Niemeyer from the European Commission&rsquo;s Joint Research Centre (JRC) for the permission to provide open&nbsp;access to the Netherlands data. Similarly, we would like to thank M. van der Velde, L. Nisini and I. Cerrani from JRC for sharing with us past MCYFS forecasts&nbsp;and Eurostat national yield statistics.</p>

opencc-by-4.0Dec 2020View details →
zenodo52/100

Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms (supplementary material)

<p>Supplementary material for Sawi et al., 2023, <i>Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms </i>(The Seismic Record). Catalog of repeating earthquakes in sequences on a 10-km long segment of the San Andreas Fault in California from 1984-2019.&nbsp;</p><p>&nbsp;</p><p><strong>Catalog Header</strong></p><p>YR/MO/DY...........Date of event</p><p>HR/MN/SC...........Time of event</p><p>LAT/LON/DEP........Location of event</p><p>EX/EY/EZ...........Relative location uncertainty (in m)</p><p>MAG................NCSN magnitude</p><p>evID.................NCSN event ID</p><p>seqID................Repeating earthquake sequence ID</p><p>isRESp............Is quasi-periodic RES (bool)</p><p>&nbsp;</p><p><strong>References:&nbsp;</strong></p><p>Sawi T., Waldhauser F., Holtzman B. K., Groebner, N. (2023) Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms. The Seismic Record.&nbsp;</p><p>Waldhauser, F., and Schaff, D. P. (2021). A Comprehensive Search for Repeating Earthquakes in Northern California: Implications for Fault Creep, Slip Rates, Slip Partitioning, and Transient Stress. J Geophys Res B Solid Earth, 126(11), 1–22.&nbsp;<a href="https://doi.org/10.1029/2021JB022495">https://doi.org/10.1029/2021JB022495</a></p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Gravity Spy Machine Learning Classifications of LIGO Glitches from Observing Runs O1, O2, O3a, and O3b

<p>This data set contains all classifications that the Gravity Spy Machine Learning model for LIGO glitches from the first three observing runs (<a href="https://doi.org/10.7935/K57P8W9D">O1</a>, <a href="https://doi.org/10.7935/CA75-FM95">O2</a> and O3, where O3 is split into <a href="https://doi.org/10.7935/nfnt-hm34">O3a</a> and <a href="https://doi.org/10.7935/pr1e-j706">O3b</a>). Gravity Spy classified all noise events identified by the <a href="https://doi.org/10.1016/j.softx.2020.100620">Omicron trigger pipeline</a> in which Omicron identified that the signal-to-noise ratio was above 7.5 and the peak frequency of the noise event was between 10 Hz and 2048 Hz. To classify noise events, Gravity Spy made <a href="https://en.wikipedia.org/wiki/Constant-Q_transform">Omega scans</a> of every glitch consisting of 4 different durations, which helps capture the morphology of noise events that are both short and long in duration.</p> <p>There are <a href="https://doi.org/10.1088/1361-6382/aa5cea">22 classes</a> used for O1 and O2 data (including No_Glitch and None_of_the_Above), while there are <a href="https://doi.org/10.1088/1361-6382/ac1ccb">two additional classes</a> used to classify O3 data (while None_of_the_Above was removed).</p> <p>For O1 and O2, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, Chirp, Extremely_Loud, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>For O3, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, <strong>Blip_Low_Frequency</strong>, Chirp, Extremely_Loud, <strong>Fast_Scattering</strong>, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>The data set is described in <a href="https://doi.org/10.1088/1361-6382/acb633"><strong>Glanzer </strong><em>et al</em><strong>. (2023)</strong></a>, which we ask to be cited in any publications using this data release. Example code using the data can be found in this <a href="https://colab.research.google.com/drive/19q_lItODPk7qw_sohlHyWPnAbY0FZyt8?usp=sharing"><strong>Colab notebook</strong></a>.</p> <p>If you would like to download the Omega scans associated with each glitch, then you can use the gravitational-wave data-analysis tool <a href="https://gwpy.github.io/docs/stable/">GWpy</a>. If you would like to use this tool, please install anaconda if you have not already and create a virtual environment using the following command</p> <pre><code class="language-bash">conda create --name gravityspy-py38 -c conda-forge python=3.8 gwpy pandas psycopg2 sqlalchemy</code></pre> <p>After downloading one of the CSV files for a specific era and interferometer, please run the following Python script if you would like to download the data associated with the metadata in the CSV file. We recommend not trying to download too many images at one time. For example, the script below will read data on Hanford glitches from O2 that were classified by Gravity Spy and filter for only glitches that were labelled as Blips with 90% confidence or higher, and then download the first 4 rows of the filtered table.</p> <pre><code class="language-python">from gwpy.table import GravitySpyTable H1_O2 = GravitySpyTable.read('H1_O2.csv') H1_O2[(H1_O2["ml_label"] == "Blip") &amp; (H1_O2["ml_confidence"] &gt; 0.9)] H1_O2[0:4].download(nproc=1)</code></pre> <p>Each of the columns in the CSV files are taken from various different inputs:&nbsp;</p> <p>[&lsquo;event_time&rsquo;, &lsquo;ifo&rsquo;, &lsquo;peak_time&rsquo;, &lsquo;peak_time_ns&rsquo;, &lsquo;start_time&rsquo;, &lsquo;start_time_ns&rsquo;, &lsquo;duration&rsquo;, &lsquo;peak_frequency&rsquo;, &lsquo;central_freq&rsquo;, &lsquo;bandwidth&rsquo;, &lsquo;channel&rsquo;, &lsquo;amplitude&rsquo;, &lsquo;snr&rsquo;, &lsquo;q_value&rsquo;] contain metadata about the signal from the <a href="https://virgo.docs.ligo.org/virgoapp/Omicron/">Omicron pipeline</a>.&nbsp;</p> <p>[&lsquo;gravityspy_id&rsquo;] is the unique identifier for each glitch in the dataset.&nbsp;</p> <p>[&lsquo;1400Ripples&rsquo;, &lsquo;1080Lines&rsquo;, &lsquo;Air_Compressor&rsquo;, &lsquo;Blip&rsquo;, &lsquo;Chirp&rsquo;, &lsquo;Extremely_Loud&rsquo;, &lsquo;Helix&rsquo;, &lsquo;Koi_Fish&rsquo;, &lsquo;Light_Modulation&rsquo;, &lsquo;Low_Frequency_Burst&rsquo;, &lsquo;Low_Frequency_Lines&rsquo;, &lsquo;No_Glitch&rsquo;, &lsquo;None_of_the_Above&rsquo;, &lsquo;Paired_Doves&rsquo;, &lsquo;Power_Line&rsquo;, &lsquo;Repeating_Blips&rsquo;, &lsquo;Scattered_Light&rsquo;, &lsquo;Scratchy&rsquo;, &lsquo;Tomte&rsquo;, &lsquo;Violin_Mode&rsquo;, &lsquo;Wandering_Line&rsquo;, &lsquo;Whistle&rsquo;] contain the machine learning confidence for a glitch being in a particular Gravity Spy class (the confidence in all these columns should sum to unity). These use the original 22 classes in all cases.</p> <p>[&lsquo;ml_label&rsquo;, &lsquo;ml_confidence&rsquo;] provide the machine-learning predicted label for each glitch, and the machine learning confidence in its classification.&nbsp;</p> <p>[&lsquo;url1&rsquo;, &lsquo;url2&rsquo;, &lsquo;url3&rsquo;, &lsquo;url4&rsquo;] are the links to the publicly-available <a href="https://gwdetchar.readthedocs.io/en/stable/omega/">Omega scans</a> for each glitch. &lsquo;url1&rsquo; shows the glitch for a duration of 0.5 seconds, &lsquo;url2&rsquo; for 1 seconds, &lsquo;url3&rsquo; for 2 seconds, and &lsquo;url4&rsquo; for 4 seconds.</p> <p>For the most recently uploaded training set used in Gravity Spy machine learning algorithms, please see <a href="https://zenodo.org/record/1486046#.YZfcar3MJqs">Gravity Spy Training Set</a> on Zenodo.&nbsp;</p> <p><br> For detailed information on the training set used for the original Gravity Spy machine learning paper, please see <a href="https://zenodo.org/record/1476156#.YZfchL3MJqs">Machine learning for Gravity Spy: Glitch classification and dataset</a> on Zenodo.</p>

opencc-by-4.0Nov 2021View details →
zenodo52/100

Unravelling the physiological and psychosocial signatures of pain by machine learning

<p>These datasets include information from 118 subjects, with 81 chronic pain patients across three different cohorts, Complex Regional Pain Syndrome (CRPS), Low Back Pain (LBP) and Spinal Cord Injury with Neuropathic Pain (SCI NP) and 37 healthy subjects. Each participant underwent 40 repetitions of experimentally induced pain, resulting in a total of 4,697 pain trials. Physiological signals (EDA and EEG) and psychosocial information have been recorded and collected. Age, gender, height, weight, BMI, medications, fatigue, sleep quality, perceived health, quality of life, sleep quality, and sick leave and validated questionnaires: Hospital Anxiety and Depression Scale (HADS), Pain Catastrophizing Score (PCS), and Pain Self-Efficacy Questionnaire (PSEQ) and Multidimensional Assessment of Interoceptive Awareness (MAIA).</p> <p>This information has been used for the publication "Unravelling the physiological and psychosocial signatures of pain by machine learning".</p> <p><span><em><strong>Cite this dataset as:</strong></em></span><br>N. Gozzi, G. Preatoni, F. Ciotti, M. Hubli, P. Schweinhardt, A. Curt, S. Raspopovic, Unraveling the physiological and psychosocial signatures of pain by machine learning. Med 0 (2024). &nbsp;<a href="https://doi.org/10.1016/j.medj.2024.07.016" target="_blank" rel="noopener">10.1016/j.medj.2024.07.016</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Exploring AdaBoost and Random Forests machine learning approaches for infrared pathology on unbalanced data sets

<p>The use of infrared spectroscopy to augment decision-making in histopathology is a promising direction for the diagnosis of many disease types. Hyperspectral images of healthy and diseased tissue, generated by infrared spectroscopy, are used to build chemometric models that can provide objective metrics of disease state. It is important to build robust and stable models to provide confidence to the end user. The data used to develop such models can have a variety of characteristics which can pose problems to many model-building approaches. Here we have compared the performance of two machine learning algorithms &ndash; AdaBoost and Random Forests &ndash; on a variety of non-uniform data sets. Using samples of breast cancer tissue, we devised a range of training data capable of describing the problem space. Models were constructed from these training sets and their characteristics compared. In terms of separating infrared spectra of cancerous epithelium tissue from normal-associated tissue on the tissue microarray, both AdaBoost and Random Forests algorithms were shown to give excellent classification performance (over 95% accuracy) in this study. AdaBoost models were more robust when datasets with large imbalance were provided. The outcomes of this work are a measure of classification accuracy as a function of training data available, and a clear recommendation for choice of machine learning approach.</p>

opencc-by-4.0May 2021View details →
zenodo52/100

Dataset for a machine learning tool to improve lymph node staging with FDG-PET/CT

<p>This upload provides Open Data associated with the publication&nbsp;&quot;A machine learning tool to improve prediction of mediastinal lymph node metastases in non-small cell lung cancer using routinely obtainable [<sup>18</sup>F]FDG-PET/CT parameters&quot; by Rogasch JMM <em>et al.</em> (2022).</p> <p>The upload contains the&nbsp;anonymized dataset&nbsp;with 10 features necessary for the final GBM model that was presented in the publication. However, the original full dataset with&nbsp;40 features was excluded from this Open Data repository because it may not comply with strict rules of data anonymization. The full dataset can be obtained from the corresponding author (julian.rogasch@charite.de) upon reasonable request.</p> <p>Besides the dataset, this upload provides the original python and R scripts that were used as well as&nbsp;their output.</p> <p>A description of all&nbsp;files&nbsp;can be found in &quot;content_description_2022_11_19.txt&quot;.</p> <p>A user-friendly web tool that implements the final machine learning model can be found here:&nbsp;<a href="https://baumgagl.github.io/PET_LN_calculator/">PET_LN_calculator</a>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
edi52/100

Bioacoustic Dataset of African and Florida Manatee Vocalizations for Machine Learning Applications, 2020-2022

This data package presents a comprehensive acoustic library of manatee vocalizations for machine learning (ML) and classifier development. It includes recordings from two species, African and Florida manatees, sampled across four locations. The species are combined due to the acoustic similarity of their vocalizations, providing a diverse and representative training set for ML algorithms. The dataset consists of 0.5-second WAV clips categorized as either containing manatee vocalizations (MV, n=18,129 clips) or not (Noise, n=23,444 clips). MV clips may include multiple vocalizations or truncated calls. All clips were manually verified by two researchers with expertise in manatee acoustics. Recordings were collected using stationary hydrophones deployed in natural habitats, with variable signal-to-noise ratios (SNR) resulting from changes in distance between the vocalizing manatees and the recorders. Background noise across sites is relatively low, with minimal anthropogenic noise; caution is advised when applying models to noisier environments. No dolphin species are believed to be present at the recording sites, and models trained on this dataset should be used cautiously in dolphin-inhabited regions to avoid false positives. If you use this dataset, please reach out to the listed contacts, we are interested in learning how it supports your work.

openCC (other)Sep 2025View details →
zenodo48/100

Dataset Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models

<p>This dataset contains the raw data (01_raw_data), processed data (02_processed_data), and plotting scripts (03_figures) related to the paper:</p> <p>"Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models"<br>Atmika Bhardwaj, Jens-Uwe Sommer, Marco Werner</p> <p>Macromolecules <strong>2024</strong>; DOI: <a href="10.1021/acs.macromol.4c00920">10.1021/acs.macromol.4c00920</a></p> <p>Please refer to the README.md files in their respective folders.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

Dataset of "Advanced machine learning techniques for State-of-Health estimation in lithium-ion batteries: A comparative study"

This research focuses on State-of-Health (SOH) estimation of lithium-ion (Li-ion) batteries to enhance lifespan and reliability. Using Samsung INR18650-35E cells, 600 cycles were analyzed with machine learning (ML) techniques, including Gaussian Process Regression (GPR), Support Vector Regression (SVR), Feed-Forward Neural Network (FFNN) and Adaptive Neuro-Fuzzy Inference System (ANFIS). Input features from charging and discharging cycles were selected with Pearson Correlation Analysis (PCA) and Exhaustive Search (ES) to optimize inputs for each ML method. Models were tested on datasets of varying sizes to evaluate performance and overfitting, including an experiment where SOH estimation of one battery was performed using training data from another. The findings highlight each model's strengths and limitations, guiding their application in battery health prediction.

opencc-by-4.0Nov 2024View details →
zenodo48/100

RADIT: A Machine Learning-Reconstructed Dataset of River Discharge, Temperature, and Heat Flux into the Arctic Ocean

<p>The Reconstructed Arctic-draining river DIscharge and Temperature (RADIT) dataset provides daily records of river discharge, temperature, and heat flux for 25 major Arctic-draining rivers from 1950 to 2023. Using machine learning methods and ERA5-Land reanalysis data, we reconstructed these key hydrological variables with high accuracy (most NSEs &gt; 0.8).</p> <p>Due to licensing restrictions and to encourage adherence to the stated licenses of the original input data, this dataset only provides the reconstructed (filled) values. Users can obtain the complete historical observational data from their original publicly available sources as detailed in our documentation. By combining these original observations with our reconstructed data, a comprehensive and continuous daily dataset from 1950 to 2023 can be assembled. Clear instructions and links for downloading the original observational data used in this study can be found at: <a href="https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data" target="_blank" rel="noopener">https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data</a>. Should you encounter any issues or have questions, please feel free to contact the first author, Zihan Wang (zhwang2018@163.com).</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

ML-TOMCAT V2.0: Machine-Learning-Based Satellite-Corrected Global Stratospheric Ozone Profile Dataset

<p>MLTOMCAT V2 is 46 years (1979-2024) of gap free ozone profile data sets that is created by correcting biases in a TOMCAT Chemical Transport Model (CTM) simulated ozone profiles. We use Random Forest regression model to correct model biases.&nbsp;</p> <p>Each file contain monthly mean zonal mean ozone profiles. There are 6 data files.</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_vmr_V2.nc</a>&nbsp;contains ozone profiles on&nbsp;geometric height levels (1 to 60 km) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Similarly,&nbsp;</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_vmr_V2.nc</a>&nbsp;contains ozone profiles on 43 MLS pressure levels&nbsp;&nbsp;(1000 to 0.1&nbsp;hPa) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Please note that data below 300 hPa (~8km) and 1 hPa (~50 km) should be used with caution.</p> <p>There are two straospheric column files</p> <p>ML-TOMCAT-SCO_120ppb_boundary_V2_197901-202412.nc and</p> <p>ML-TOMCAT-SCO_150ppb_boundary_V2_197901-202412.nc</p> <p>Stratospheric column files calculated using 120 ppb and 150 ppb as a chemical ozone boundaries.</p> <p>A manuscript describing MLTOMCAT would be published in EESD (Dhomse et al., 2021).</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Potential forest conservation value rasters for Denmark from Assmann et al. "LiDAR data fusion and machine learning identify temperate forests of high conservation value"

<p>Potential forest conservation value (high / low) rasters for Denmark based on a remote sensing data fusion approach. Please see manuscript (below) for a detailed description of the methods and data products.&nbsp;</p> <p><br>Jakob J. Assmann, Pil B. M. Pedersen, Jesper E. Moeslund, Cornelius Senf, Urs A. Treier, Derek Corcoran, Zs&oacute;fia Koma, Thomas Nord-Larsen, Signe Normand. In prep. LiDAR data fusion and machine learning identify temperate forests of high conservation value.</p> <p><br>When using the data, please cite the above manuscript.&nbsp;</p> <p><br>Files description:</p> <ul> <li>Compressed and cloud optimised rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:3857 <ul> <li>forest_quality_ranger_biowide_10m_cog_epsg3857.tif &nbsp; &nbsp; RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m_cog_epsg3857.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m_cog_epsg3857.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m_cog_epsg3857.tif &nbsp; &nbsp; GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li>Aggregated rasters of potential forest conservation value projections for Denmark (100 m res.) in EPSG:25832 <ul> <li>forest_quality_ranger_biowide_100m.tif RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_100m.tif RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_100m.tif GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_100m.tif GBM model projections based on SustainScapes stratification&nbsp;</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li>Uncompressed and tiled rasters of potential forest conservation value projections for Denmark (10 m res.) in EPSG:25832<br>Please note: the archives contain approx. 42k tiles, each 10 x 10 km, as well as a VRT file for covenient loading.&nbsp; <ul> <li>forest_quality_ranger_biowide_10m.zip RandomForest model projections based on BIOWIDE stratification (!! best performing model !!)</li> <li>forest_quality_ranger_sustainscapes_10m.zip RandomForest model projections based on SustainScapes stratification</li> <li>forest_quality_gbm_biowide_10m.zip GBM model projections based on BIOWIDE stratification</li> <li>forest_quality_gbm_sustainscapes_10m.zip GBM model projections based on SustainScapes stratification</li> </ul> </li> </ul>

opencc-by-4.0Dec 2023View details →
zenodo48/100

A Bayesian Machine Learning Framework for Animal Telemetry Data

<p>The data and tutorial in this repository are intended to be used in conjunction with the tutorial with our manuscript titled "A Bayesian Machine Learning Framework for Animal Telemetry Data." Telemetry data for three lesser prairie-chickens are provided here as .csv files. For more information about the data, please refer to our manuscript or contact Andrew Whetten or David Haukos for more information.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Dataset for "Machine learning predictions on an extensive geotechnical dataset of laboratory tests in Austria"

<p>This dataset comprises over 20 years of geotechnical laboratory testing data collected primarily from Vienna, Lower Austria, and Burgenland. It includes 24 features documenting critical soil properties derived from particle size distributions, Atterberg limits, Proctor tests, permeability tests, and direct shear tests. Locations for a subset of samples are provided, enabling spatial analysis.</p> <p>The dataset is a valuable resource for geotechnical research and education, allowing users to explore correlations among soil parameters and develop predictive models. Examples of such correlations include liquidity index with undrained shear strength, particle size distribution with friction angle, and liquid limit and plasticity index with residual friction angle.</p> <p>Python-based exploratory data analysis and machine learning applications have demonstrated the dataset's potential for predictive modeling, achieving moderate accuracy for parameters such as cohesion and friction angle. Its temporal and spatial breadth, combined with repeated testing, enhances its reliability and applicability for benchmarking and validating analytical and computational geotechnical methods.</p> <p>This dataset is intended for researchers, educators, and practitioners in geotechnical engineering. Potential use cases include refining empirical correlations, training machine learning models, and advancing soil mechanics understanding. Users should note that preprocessing steps, such as imputation for missing values and outlier detection, may be necessary for specific applications.</p> <p><strong>Key Features</strong>:</p> <ul> <li><strong>Temporal Coverage</strong>: Over 20 years of data.</li> <li><strong>Geographical Coverage</strong>: Vienna, Lower Austria, and Burgenland.</li> <li><strong>Tests Included</strong>: <ul> <li>Particle Size Distribution</li> <li>Atterberg Limits</li> <li>Proctor Tests</li> <li>Permeability Tests</li> <li>Direct Shear Tests</li> </ul> </li> <li><strong>Number of Variables</strong>: 24</li> <li><strong>Potential Applications</strong>: Correlation analysis, predictive modeling, and geotechnical design.</li> </ul> <p><strong>Technical Details</strong>:</p> <ul> <li>Missing values have been addressed using K-Nearest Neighbors (KNN) imputation, and anomalies identified using Local Outlier Factor (LOF) methods in previous studies.</li> <li>Data normalization and standardization steps are recommended for specific analyses.</li> </ul> <p><strong>Acknowledgments</strong>:<br>The dataset was compiled with support from the European Union's MSCA Staff Exchanges project 101182689 Geotechnical Resilience through Intelligent Design (GRID).</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows

<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, &ldquo;On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,&rdquo; IEEE Communications Magazine, pp. 2&ndash;8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Elevating Cybersecurity for Smart Grid Systems—A Container-Based Approach Enhanced by Machine Learning

<p>README<br>Title<br>Elevating Cybersecurity for Smart Grid Systems&mdash;A Container-Based Approach Enhanced by Machine Learning</p> <p>Authors<br>Mays Abukeshek, School of Computer Science, Faculty of Technology, University of Sunderland, University of Huddersfield, UK<br>Email: mays.abukeshek@sunderland.ac.uk, Mays.abukeshek@hud.ac.uk<br>Basel Barakat, School of Computer Science, Faculty of Technology, University of Sunderland, UK<br>Email: basel.barakat@sunderland.ac.uk<br>Bamidele Ajayi, School of Computer Science, Faculty of Technology, University of Sunderland, UK<br>Email: bamidele.ajayi@research.sunderland.ac.uk<br>Abstract<br>This dataset supports the paper "Elevating Cybersecurity for Smart Grid Systems&mdash;A Container-Based Approach Enhanced by Machine Learning," which presents a comprehensive implementation of a cybersecurity solution for smart grid network containers. The methodology utilizes:</p> <p>Qualys API-based vulnerability scanning and reporting system for vulnerability identification<br>Docker deployment for security and isolation<br>Advanced load balancing techniques for resource optimization<br>Machine learning-powered anomaly detection for threat identification and vulnerability prioritization.<br>The dataset contains details of several simulated attacks enabling effective training and evaluation of a robust machine-learning model.</p> <p>Data Description<br>The dataset includes logs from conducted attacks on containerized nodes, generated to reflect real-world scenarios. The simulated attacks include:</p> <p>Denial of Service (DoS)<br>Remote-to-Local (R2L)<br>User-to-Root (U2R)<br>Probes<br>Contents<br>Csv_file.csv: This file contains the dataset used for training and evaluating the machine learning models. The columns in the dataset represent various features and results of the simulated attacks.<br>Data Columns and Rows<br>Timestamp:</p> <p>Description: The exact date and time when the data was recorded.<br>time: 2023-06-01 12:00:00</p> <p>Attack_Type:</p> <p>Description: The type of cyber-attack conducted.<br>Possible Values: DoS, R2L, U2R, Probe<br>Example: DoS<br>Notes: Categorizes the type of attack, crucial for training classification models.<br>CPU_Utilization (%):</p> <p>Description: The percentage of CPU resources used during the attack.<br>Example: 52.3<br>Notes: Indicates the load on the CPU during the attack, useful for assessing the impact of attacks on system performance.<br>Memory_Utilization (%):</p> <p>Description: The percentage of memory resources used during the attack.<br>Example: 63.4<br>Notes: Shows memory usage which can be a critical factor in understanding system performance under attack conditions.<br>Network_Bandwidth (Mbps):</p> <p>Description: The bandwidth of the network in Megabits per second.<br>Example: 100<br>Notes: Reflects the network load and is essential for analyzing the impact on network performance.<br>Vulnerabilities_Detected:</p> <p>Description: The number of vulnerabilities detected during the attack.<br>Example: 289<br>Notes: Indicates the effectiveness of the vulnerability scanning process and the system's exposure to threats.<br>Mean_Response_Time (ms):</p> <p>Description: The average response time in milliseconds during the attack.<br>Example: 87<br>Notes: Important for evaluating the responsiveness of the system under attack conditions.<br>Throughput (requests/second):</p> <p>Description: The number of requests the system can handle per second during the attack.<br>Example: 1068<br>Notes: Measures the capacity and efficiency of the system under load.<br>Example Row<br>Timestamp &nbsp; &nbsp;Attack_Type &nbsp; &nbsp;CPU_Utilization (%) &nbsp; &nbsp;Memory_Utilization (%) &nbsp; &nbsp;Network_Bandwidth (Mbps) &nbsp; &nbsp;Vulnerabilities_Detected &nbsp; &nbsp;Mean_Response_Time (ms) &nbsp; &nbsp;Throughput (requests/second)<br>2023-06-01 12:00:00 &nbsp; &nbsp;DoS &nbsp; &nbsp;52.3 &nbsp; &nbsp;63.4 &nbsp; &nbsp;100 &nbsp; &nbsp;289 &nbsp; &nbsp;87 &nbsp; &nbsp;1068<br>Usage<br>This dataset can be used to:</p> <p>Train and evaluate machine learning models for cybersecurity applications in smart grid systems.<br>Analyze the performance of different machine learning models in detecting and prioritizing vulnerabilities.<br>Understand the impact of various types of cyber-attacks on containerized environments.<br>Methodology<br>The dataset was created using a combination of Qualys API-based vulnerability scanning and Docker containerization. Multiple container clusters were subjected to various simulated attacks, and the performance of machine learning models was evaluated based on accuracy, precision, recall, and F1-scores.</p> <p>Acknowledgments<br>This research was supported by the University of Sunderland and the University of Huddersfield.</p> <p>References<br>Please refer to the full paper for detailed methodology, implementation, and analysis:<br>IEEE</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Using machine learning to integrate genetic and environmental data to model genotype-by-environment interactions

<p>Files generated from the study described in&nbsp;<a href="https://doi.org/10.1101/2024.02.08.579534">Fernandes et. al (2024)</a> .</p> <p>The file "cvs_h2s.csv" comprises the coefficient of variation and the Cullis heritability for each environment.</p> <p>The file "all_predictions.csv" contains the predictions from all the models evaluated, in different cross-validation (CV) scenarios.</p> <p>The file "coincidence_index.csv" has the Coincidence Index (CI) for each CV and models evaluated in our study.</p> <p>Our study used the multi-environment maize yield trials data from the Genomes to Fields 2022 initiative (<a href="https://doi.org/10.1186/s13104-023-06421-z">Lima et. al 2024</a>).</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Viral Pneumonia Classification Using Machine and Transfer Learning Techniques

<p>Pneumonia is considered a deadly and harmful disease throughout the world. Pneumonia can be lethal if not treated promptly with antibiotics. As a result, early detection of pneumonia increases the likelihood of recovery and lowers mortality. X-rays are one of the most important diagnostic tools for pneumonia. Because of its lower diagnostic costs, the chest X-ray is routinely used to diagnose various lung illnesses. Indeed, diagnosis can be subjective for various reasons, including disease presentation, which might be confusing in chest X-ray images or misdiagnosed as another condition. As a result, the employment of chest X-rays for the diagnoses of pneumonia disease is considered a way forward to fight the challenges being faced with during the examination process and expert readings of results. The dataset comprises 1,067 Pneumonia Chest X-ray images that were curated from the Hopskin Diagnostic Center Nigeria for Research Purposes. This was used to classify Pneumonia disease for pneumonia class encoding. The result yield Pneumonia Disease with High Accuracy, precision and Recall.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

A large data-set of CASP protein refinement simulations for machine-learning

<p>The uploaded trajectory data originates from our own laboratory&#39;s refinement method in CASP11 and CASP12 for which the reference crystal structure is available in the PDB. In total the trajectory data consists of&nbsp; 904 trajectories with 3419 ns cumulative simulation time and 1,709,704 snapshots with a delta t =2 ps from 42 different protein systems.</p> <p><strong>File Overview</strong></p> <ul> <li><strong>trajectory_data_pdbs.tar.gz :</strong> contains the PDB files of the different trajectories as well as the starting model and reference crystal structure for each target</li> <li><strong>casp_normalized_all_data_final.csv.gz :&nbsp; </strong>contains the trajectory features calculated for each snapshot from the trajectory PDBs</li> <li><strong>cv_folds.csv : </strong>contains the 7 fold cross-validation assignment used to assess the performance of the model<br> &nbsp;</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Machine learning for Gravity Spy: Glitch classification and dataset

<p>We present the first version of the training set used in the Gravity Spy citizen science project. This training set, discussed in detail <a href="https://www.sciencedirect.com/science/article/pii/S0020025518301634">here</a>, was utilized to train the convolutional neural network employed in the Gravity Spy project. We anticipate moving forward to release more labelled Gravity Spy data sets, including a refined version of this training set which can be found here&nbsp;<a href="https://doi.org/10.5281/zenodo.1476551">10.5281/zenodo.1476551</a>, and data sets containing the annotations provided by our citizen science volunteers.</p> <p><strong>Data Set Information</strong></p> <p>There are three files provided in this data set</p> <ul> <li><strong>trainingset_v1d0_metadata.csv</strong> <ul> <li>This file has three columns, <em>gravityspy_id, label, </em>and <em>sample_type.</em><em> gravityspy_id </em>is the unique 10 character hash given to every Gravity Spy sample. <em>label</em> is the string label of the sample. <em>sample_type </em>indicates whether this sample was used in the paper for testing training or validating the models. This is provided for those who would like to do direct comparisons to the network described in the paper.</li> </ul> </li> <li><strong>trainingsetv1d0.h5</strong> <ul> <li>This file contains the exact arrays used in the paper for every Gravity Spy sample. Each Gravity Spy sample is defined by four different images with varying temporal duration, <em>0.5, 1.0, 2.0, and 4.0</em> second, respectively. This also determines the naming conventions of the PNGs: <em>interferometer_gravityspyid_spectrogram_duration.png (e.g. H1_Fv3p6eROvA_spectrogram_0.5.png, H1_Fv3p6eROvA_spectrogram_1.0.png, H1_Fv3p6eROvA_spectrogram_2.0.png, H1_Fv3p6eROvA_spectrogram_4.0.png</em>).</li> <li>This file contains all the information needed for each sample in the Gravity Spy dataset (i.e. the label, the sample type of the sample, the unique id of the sample, and the image data for that sample. <ul> <li>/1080Lines/validation/xUEyaWr34c Group<br> /1080Lines/validation/xUEyaWr34c/0.5.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/1.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/2.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/4.0.png Dataset {1, 140, 170}</li> </ul> </li> </ul> </li> <li><strong>trainingsetv1d0.tar.gz</strong> <ul> <li>Contains the raw PNGs of the Gravity Spy training set.</li> <li>The structure of the folder is <em>/&quot;label&quot;/&quot;sample_type&quot;/&quot;pngs&quot;</em></li> </ul> </li> </ul> <p><strong>Data Set Parsing Information</strong></p> <p>To read and crop out the plot axis and labels of the provided PNGs, the following small python code using scikit-image should work.</p> <p>from skimage import io</p> <p>image_data = io.imread(&quot;filename_of_image&quot;)</p> <p>x=[66, 532]; y=[105, 671]</p> <p>image_data = image_data[x[0]:x[1], y[0]:y[1], :3]</p>

opencc-by-4.0Oct 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record