Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data and software: Heat flux for semi-local machine-learning potentials

<p><br> This repository contains data, code, and related artefacts supporting the following publication:</p> <p>&quot;Heat flux for semi-local machine-learning potentials&quot;<br> by Marcel F. Langer, Florian Knoop, Christian Carbogno, Matthias Scheffler, and Matthias Rupp<br> arXiv: TBD<br> doi: TBD<br> &nbsp;</p> <p>More details can be found in the main README.md file, and the README.md files in the subfolders.</p> <p><br> For any further questions, feel free to contact mail@marcel.science, @marceldotsci&nbsp;on Twitter, or @marcel@sigmoid.social.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

EO4WildFires: An Earth Observation multi-sensor, time-series machine-learning-ready benchmark dataset for wildfire impact prediction

<p>This paper presents a benchmark dataset called EO4WildFires; a multi-sensor (multi spectral; Sentinel-2, Synthetic-Aperture Radar - SAR; Sentinel-1, meteorological parameters; NASA Power) time-series dataset that spans 45 countries, which can be used for developing machine learning and deep learning methods targeted for the estimation of the area that a forest wildfire might cover.</p> <p>This novel EO4WildFires dataset is annotated using EFFIS (European Forest Fire Information System) as forest fire detection and size estimation data source. A total of 31,742 wildfire events are gathered from 2018 to 2022. For each event, Sentinel-2 (multispectral), Sentinel-1 (SAR) and meteorological data are assembled into a single data cube. The meteorological parameters that are included in the data cube are: ratio of actual partial pressure of water vapor to the partial pressure at saturation, average temperature, bias corrected average total precipitation, average wind speed, fraction of land covered by snowfall, percent of root zone soil wetness, snow depth, snow precipitation, as well as percent of soil moisture.</p> <p>The main problem that this dataset is designed to address, is the severity forecasting before wildfires occur. The dataset is not used to predict wildfire events, but rather to predict the severity (size of area damaged by fire) of a wildfire event, if that happens in a specific place under the current and historical forest status, as recorded from multispectral and SAR images, and meteorological data.</p> <p>Using the data cube for the collected wildfire events, the EO4WildFires dataset is used to realize three (3) different preliminary experiments, in order to evaluate the contributing factors for wildfire severity prediction. The first experiment evaluates wildfire size using only the meteorological parameters, the second one utilizes both the multispectral and SAR parts of the dataset, while the third exploits all dataset parts. In each experiment, machine learning models are developed, and their accuracy is evaluated.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

A catalog of associated, machine-learning-derived phase arrival times for ten days of seismic data in the Yellowstone region

<p>This dataset contains the associated phase picks and event information from applying a deep learning phase picker to continuous data recorded over March 25 &ndash; April 3, 2014, on 20 three-component stations and 14 vertical-component stations in the Yellowstone region. This 10-day period contains an M<sub>w</sub> 4.8 event, the largest earthquake in the Yellowstone region since 1980. The catalog and deep learning phase picker are described in Armstrong et al. (submitted).</p> <p>The arrivals were associated using the method described by Baker et al. (2021) and located using HypoInverse2000 (Klein, 2002). There are 1,053 events in this catalog, including 855 that were previously unidentified. Events that also appear in the University of Utah Seismograph Stations catalog have an event identifier (evid) beginning with &ldquo;6&rdquo;, while new events begin with &ldquo;9&rdquo;.&nbsp;</p> <p>Columns include:</p> <ul> <li>A simple event number</li> <li>the network, station, channel, and location code for the arrival time</li> <li>the arrival time in UTC (arrival_time) and Unix (arrival_time_epoch) format</li> <li>any static correction applied to the arrival time</li> <li>the P-pick first motion polarity as determined by a machine learning model - up (1), down (-1), or unknown (0)</li> <li>the arrival time residual&nbsp;</li> <li>the take off angle in degrees&nbsp;</li> <li>the event latitude and longitude in degrees</li> <li>the event depth in km</li> <li>the event origin time in UTC (origin_time) and Unix (origin_time_epoch) format</li> <li>the azimuthal gap of the event in degrees</li> <li>the root mean square error (RMS) of the event location</li> <li>the event identifier (evid) - begins with a &ldquo;6&rdquo; for events in the UUSS catalog and a &ldquo;9&rdquo; for new events</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

A dataset of Earth Observation Data for Lithological Mapping using Machine Learning

<p><strong>Dataset Information</strong></p> <p>Machine Learning (ML) algorithms had successfully contributed in the creation of automated methods of recognizing patterns in high-dimensional data. Remote sensing data&nbsp; covers&nbsp; wide&nbsp; geographical areas and could be used to solve the problem of the demand of various&nbsp; in-situ data.&nbsp; Lithologicall mapping using remotely sensed data&nbsp; is one of the most challenging&nbsp; applications of ML algorithms. In the framework of the &ldquo;AI for Geoapplications&rdquo; project , ML and especially Deep Learning (DL) methodologies are investigated&nbsp; for&nbsp; the identification and characterization of the lithology based on remote sensing data in various&nbsp; pilot areas&nbsp; in Greece.&nbsp; In order to train and test the various ML algorithms, a dataset consisting of&nbsp; 30 ROIs selected&nbsp; mainly&nbsp; from low -vegetated areas,&nbsp; that cover 2% of the total&nbsp; area of Greece was created</p> <p><strong>Dataset Preprocessing</strong></p> <p>Dataset preprocessing was executed using a combination of SNAP, QGIS and ENVI tools.</p> <p>Preprocessing steps:</p> <p>Defining areas with the following properties:</p> <ul> <li> <p>Zero cloud and snow coverage</p> </li> <li> <p>No water bodies</p> </li> <li> <p>Minimum vegetation</p> </li> </ul> <p>For the Aster Images:</p> <ul> <li> <p>Subset on defined areas</p> </li> <li> <p>Mosaic images when needed</p> </li> <li> <p>Digitising clouds</p> </li> </ul> <p>For the Labels:</p> <ul> <li> <p>We got the Soil map from YPEN (<a href="https://ypen.gov.gr/">https://ypen.gov.gr/</a>)</p> </li> <li> <p>Subset on defined areas</p> </li> <li> <p>All categories are represented with good analogies</p> </li> <li> <p>Clip label files with digitised clouds</p> </li> <li> <p>Rasterize</p> </li> </ul> <p>&nbsp;</p> <p>For the Labels we have eighteen categories for the twenty-eight areas that we collected data.&nbsp;We use the following coding&nbsp;for the&nbsp;Labels of our <strong>Dataset</strong>:</p> <table> <tbody> <tr> <td> <p><strong>Alluvial deposits</strong></p> </td> <td> <p><strong>0</strong></p> </td> </tr> <tr> <td> <p><strong>Limestone colluvial deposits</strong></p> </td> <td> <p><strong>1</strong></p> </td> </tr> <tr> <td> <p><strong>Limestones</strong></p> </td> <td> <p><strong>2</strong></p> </td> </tr> <tr> <td> <p><strong>Schists</strong></p> </td> <td> <p><strong>3</strong></p> </td> </tr> <tr> <td> <p><strong>Quaternary sediments</strong></p> </td> <td> <p><strong>4</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss</strong></p> </td> <td> <p><strong>5</strong></p> </td> </tr> <tr> <td> <p><strong>Slope fan debris</strong></p> </td> <td> <p><strong>6</strong></p> </td> </tr> <tr> <td> <p><strong>Mixed flysch</strong></p> </td> <td> <p><strong>7</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch shale and cherts</strong></p> </td> <td> <p><strong>8</strong></p> </td> </tr> <tr> <td> <p><strong>Dolomites</strong></p> </td> <td> <p><strong>9</strong></p> </td> </tr> <tr> <td> <p><strong>Granite</strong></p> </td> <td> <p><strong>10</strong></p> </td> </tr> <tr> <td> <p><strong>Sandstone flysch</strong></p> </td> <td> <p><strong>11</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch colluvial deposits</strong></p> </td> <td> <p><strong>12</strong></p> </td> </tr> <tr> <td> <p><strong>Peridotite and Gabbro</strong></p> </td> <td> <p><strong>13</strong></p> </td> </tr> <tr> <td> <p><strong>River bed deposits</strong></p> </td> <td> <p><strong>14</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss colluvial deposits</strong></p> </td> <td> <p><strong>15</strong></p> </td> </tr> <tr> <td> <p><strong>Not available</strong></p> </td> <td> <p><strong>-100</strong></p> </td> </tr> <tr> <td> <p><strong>cloud coverage</strong></p> </td> <td> <p><strong>-999</strong></p> </td> </tr> </tbody> </table> <p>The following table lists the available <strong>areas </strong>and the <strong>categories </strong>that each contains<strong>:&nbsp;<a href="https://docs.google.com/spreadsheets/d/17q0L5Ltz7V4uBY9i6DhULJsJCtf7BOY1nbblB-hf3Pw/edit?usp=share_link">Lithology_Dataset</a> </strong></p> <p>&nbsp;</p> <p>For the <strong>Sentinel-2 images</strong>, we made the following process:</p> <ul> <li> <p><strong>Resampling 10m</strong></p> </li> <li> <p><strong>Subset on defined areas</strong></p> </li> </ul> <p>The Sentinel-2 map contains: Sentinel 2 false colour composite 11/8/4 with OSM background</p> <p>The Final step is the collocation of the previous into a datacube i.e a multidimensional array with 25 bands (datacube dimensions differentiate for every area) using the Aster image as base (15m spatial resolution).&nbsp;</p> <ul> <li> <p>Bands 1-14: Aster</p> </li> <li> <p>Bands 15-24: S2</p> </li> <li> <p>Band 25: Label</p> </li> </ul> <p>The code for preprocessing the dataset in order to be used for machine learning algorithms can be found in the following link:&nbsp;&nbsp;</p> <p><a href="https://github.com/georgegiannop/Lithology">https://github.com/georgegiannop/Lithology</a></p> <p><strong>Citation</strong></p> <p>If you use this dataset in your work, please cite our paper:</p> <p>Vernikos, I., Giannopoulos, G., Christopoulou, A., Begaj, A., Stefouli, M., Bratsolis, E., and Charou, E.: A dataset of Earth Observation Data for Lithological Mapping using Machine Learning, EGU General Assembly 2023, Vienna, Austria, 24&ndash;28 Apr 2023, EGU23-17570,&nbsp;<a href="https://doi.org/10.5194/egusphere-egu23-17570">https://doi.org/10.5194/egusphere-egu23-17570</a>, 2023.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2

<p>Supplementary files containing datasets needed to reproduce the results of the manuscript &quot;Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states&quot; by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 &amp; 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 1

<p><strong>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al (https://doi.org/10.1101/2023.02.21.529387).</strong></p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. models.zip - contains thermodynamically curated steady-state and nonlinear kinetic models of <em>E. coli </em>metabolism used in this study. Also contains the samples of steady-state metabolite concentrations and metabolic fluxes used in the study presented in Figure 3 (steady-state samples used for preparing Figures 2 and 4).</p> <p>2. renaissance_incidence_results.zip - self-explanatory (Figure 2a and 2b)</p> <p>3. ODE_solutions.zip - self-explanatory (Figure 2c)</p> <p>4. bioreactor_simulations1-3.zip - self-explanatory (Figure 2d)</p> <p>5. steady_state_analysis.zip - RENAISSANCE results obtained for each of the steady states (Figure 3a)</p> <p>6. subspace_analysis.zip - RENAISSANCE results presented in Figure 3b-g</p> <p><strong>The remaining datasets are published in the following links</strong></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.7930084</em></p> <p><em>&nbsp;- https://doi.org/10.5281/zenodo.10391802</em></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

MISATO - Machine learning dataset for structure-based drug discovery

<p>Developments in Artificial Intelligence (AI) have had an enormous impact on scientific research in recent years. Yet, relatively few robust methods have been reported in the field of structure-based drug discovery. To train AI models to abstract from structural data, highly curated and precise biomolecule-ligand interaction datasets are urgently needed. We present MISATO, a curated dataset of almost 20000 experimental structures of protein-ligand complexes, associated molecular dynamics traces, and electronic properties. Semi-empirical quantum mechanics was used to systematically refine protonation states of proteins and small molecule ligands. Molecular dynamics traces for protein-ligand complexes were obtained in explicit water. The dataset is made readily available to the scientific community via simple python data-loaders. AI baseline models are provided for dynamical and electronic properties. This highly curated dataset is expected to enable the next-generation of AI models for structure-based drug discovery. Our vision is to make MISATO the first step of a vibrant community project for the development of powerful AI-based drug discovery tools.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

CREMP: Conformer-rotamer ensembles of macrocyclic peptides for machine learning

<p>CREMP:&nbsp;A&nbsp;resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides. CREMP contains 36,198 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 31.3 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the &ldquo;pickle&rdquo; folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single &ldquo;pickle.tar.gz&rdquo; archive. In the &ldquo;sdf_and_json&rdquo; folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, &ldquo;sdf_and_json.tar.bz2&rdquo;. A single summary CSV file is also provided containing &rdquo;sequence&rdquo;, &ldquo;smiles&rdquo;, &ldquo;num_monomers&rdquo;, &ldquo;num_atoms&rdquo;, &ldquo;num_heavy_atoms&rdquo;, along with the CREST metadata &ldquo;totalconfs&rdquo;, &ldquo;uniqueconfs&rdquo;, &ldquo;lowestenergy&rdquo;, &ldquo;poplowestpct&rdquo;, &ldquo;temperature&rdquo;, &ldquo;ensembleenergy&rdquo;, &ldquo;ensembleentropy&rdquo;, and &ldquo;ensemblefreeenergy&rdquo;. The number of unique conformers with different 3D structures is given by &ldquo;uniqueconfs&rdquo;, while &ldquo;totalconfs&rdquo; includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 32 GB for "pickle.tar.gz" and 210 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Data and code for "Phase transitions in inorganic halide perovskites from machine learning potentials: The impact of size, rate, and the underlying exchange-correlation functional"

<p>This record contains databases with data from density functional theory calculations used for training a series of neuroevolution potentials (NEPs), which are also included here. Information is also included for how to access the databases and run the NEP models.</p> <p><strong>Databases</strong><br> The <code>*.db</code> files are databases with the results from density functional theory (DFT) calculations. These are sqlite databases in ase format, see <a href="https://wiki.fysik.dtu.dk/ase/tutorials/tut06_database/database.html">here</a> for more information. The <code>demo-database-access.py</code> script illustrates the most basic access.</p> <p><strong>Models</strong><br> The neuroevolution potential (NEP) models described in the publication can be found in the <code>nep-*.txt</code> files. They can be used in conjunction with the <a href="https://gpumd.org">GPUMD package</a>. The <a href="https://calorine.materialsmodeling.org">calorine package</a> provides a Python interface to GPUMD.</p> <p><strong>Primitive structures</strong><br> Several primitive structures in extended xyz format can be found in the <code>*.xyz</code> files. These structures have been relaxed using the NEP models included here. The <code>demo-for-using-structures-and-models.py</code> script illustrates how to access the structures and models.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Dataset - What are the Machine Learning best practices reported by practitioners on Stack Exchange?

<p>The data correspond to the posts (questions and answers) retrieved by querying for posts related to the tag &#39;machine learning&#39; and the phrase &#39;best practice(s).&#39; The data were used as the basis for a study currently under review on discussing machine learning best practices as discussed by practitioners in question-and-answer communities such as Stack Exchange. The information from each type of post (i.e., questions and answers) is presented in multiple formats (i.e., .txt, .csv, and .xlsx).</p> <p>&nbsp;</p> <p><strong>Answers - Variables</strong></p> <ul> <li><strong>AID</strong>:<strong>&nbsp;</strong>&nbsp;Unique identification of the answer in the Q&amp;A website.</li> <li><strong>ParentId</strong>: Unique identification of the question associated with the answer in the Q&amp;A website&nbsp;</li> <li><strong>AcceptedAnswerId</strong>&nbsp;: In the case in which an answer is the most voted question associated with the&nbsp;<em>ParentId</em>, and it is different from the accepted answer, a different identifier from the&nbsp;<em>AID</em>&nbsp;is available. In the case in which the accepted question had a&nbsp;<em>score</em>&nbsp;lower than 1, a -1 is assigned.&nbsp;</li> <li><strong>ABody:</strong>&nbsp;&nbsp;HTML text of the answer.</li> <li><strong>Score:</strong>&nbsp;Upvotes - downvotes of the answer.</li> <li><strong>url_Answer:</strong>&nbsp;&nbsp;URL of the answer. The question URL can be from different websites.&nbsp;&nbsp;</li> <li><strong>type:</strong>&nbsp;best or accepted. Accepted in the case that the information belongs to the accepted answer of the&nbsp;<em>ParentId&nbsp;</em>question and best in the case in which it is the most voted question of the&nbsp;<em>ParentId&nbsp;</em>question.</li> <li><strong>Date:&nbsp;</strong>Creation date of the answer.</li> </ul> <p><strong>Questions - Variables</strong></p> <ul> <li><strong>QID</strong>: Unique identification of the question in the Q&amp;A website.&nbsp;</li> <li><strong>AcceptedAnswerId</strong>: Unique identification of the accepted answer for a specific question in the Q&amp;A website. In the case in which a question had a most-voted answer different from the accepted one, and the accepted one had a negative score, a -1 was assigned to the&nbsp;&nbsp;<em>AcceptedAnswerId</em><strong>.&nbsp;</strong></li> <li><strong>BestAnswerId</strong>: Unique identification of the most voted answer for a specific question in the Q&amp;A website. In the case in which the most voted and accepted questions were the same, then a -1 was assigned to the&nbsp;<em>BestAnswerId</em>.&nbsp;&nbsp;</li> <li><strong>Qtitle</strong>: Title of the question.</li> <li><strong>QBody</strong>: HTML text of the question.</li> <li><strong>Score</strong>: Upvotes - downvotes of the questions.</li> <li><strong>QTags</strong>: Tags that are associated with each question.</li> <li><strong>url_question</strong>: URL of the question. The question URL can be from different websites. &nbsp;</li> <li><strong>Date</strong>: Creation date of the question</li> </ul> <p>This dataset is a subset of the Stack Exchange dump of 03.2021 (<a href="https://archive.org/details/stackexchange_20210301">https://archive.org/details/stackexchange_20210301</a>) in which a series of filters were applied to obtain the data used in the study.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Dataset - Downscaling ERA5 Wind Speed Data: A Machine Learning approach considering Topographic Influences

<p>This dataset provides three products:</p> <p><strong>1. The topographic data.&nbsp;&nbsp;</strong></p> <p>These data are provided as GeoTIFF files for Europe with 1km x 1km spatial resolution. These maps include:</p> <ul> <li>Digital Elevation Model (DEM) map: Europe_DEM.tif</li> <li>Slope map: Europe_slope.tif</li> <li>Aspect map: Europe_aspect.tif</li> <li>Topographic Position Index (TPI) with a 5 km radius map: Europe_TPI_5.tif</li> <li>Topographic Position Index (TPI) with a 75 km radius map: Europe_TPI_75.tif</li> <li>Terrain Diversity Index (TDI) map: Europe_TDI.tif</li> </ul> <p>These data can be used as input maps for the preprocessing step. In addition, the two TPI maps can also be used in the regression process.</p> <p><strong>2. The resulting map of the preprocessing step.</strong>&nbsp;</p> <p>This map offers predictions on the quality of ERA5 data across Europe and is also provided as a GeoTIFF file with 1km x 1km spatial resolution under the name:</p> <ul> <li>&nbsp;Europe_classification.tif</li> </ul> <p>In this map, Class1 represents a good ERA5 quality with an RMSE of less than 1.5 m/s, Class2 represents a moderate ERA5 quality with an RMSE bigger than 1.5 m/s but less than 3 m/s, while Class 3 indicates a poor ERA5 quality with an RMSE greater than 3 m/s.</p> <p><strong>3. The downscaled wind speed time series data. </strong></p> <p>Europe has been divided into 64 equal area blocks to accommodate the large data size. Each downscaled dataset is provided as a NetCDF file, offering hourly wind speed time series for a year (8760 hours) at approximately 1km x 1km spatial resolution. Each NetCDF file has three dimensions: 'lon' representing longitude, 'lat' representing latitude, and 'time' representing the hour. The variable name for wind speed in the NetCDF file is 'WindSpeed'. The 'WindSpeed' variable is stored as an Int32 data type in the NetCDF file, with values multiplied by 10000 in order to significantly reduce the data size. To utilize this variable, please divide it by 10000.</p> <p>For regions identified as Class1 and Class2, the downscaled wind speed is obtained through a simple nearest neighbour spatial interpolation of ERA5 due to the good quality of ERA5 in these regions. However, for the regions identified as Class3, the downscaled wind speed is derived using the machine learning-based regression approach described in the relevant publication. The geographic extent and the visual representation for each block are provided&nbsp;in 'Readme.pdf' document.</p> <p>&nbsp;</p> <p>To cite this dataset, please cite our published paper in Environmental Research Letters (<strong>DOI:</strong> 10.1088/1748-9326/aceb0a)</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Machine-learning based lightning nowcasting data archive

<p>This data archive contains the&nbsp;derived data supporting the findings of article &quot;Lightning nowcasting with aerosol-informed machine learning and satellite-enriched dataset&quot;. The paper is currently in the preprint version:&nbsp; https://doi.org/10.21203/rs.3.rs-2616886/v1</p> <p>The prediction results in this data archive are generated by various models:</p> <p>1. Current model. The model involves data input of aerosol observations together with meteorological variables and auxiliary datasets, as well as data enrichment by Geostationary Lightning Mapper (GLM). In the demo of the dataset, the year of 2020 is trained and predicted on a cross-validation scheme.&nbsp;</p> <p>2. LMA model. The model acts as the baseline model considering only data label obtained from the ground-based Lightning Mapping Array (LMA), which observes accurate lightning occurrence in limited&nbsp;spatial range.</p> <p>3. No-AOD model. The model acts as the baseline model considering no aerosol observation is utilized during the machine learning process.&nbsp;</p> <p>The model results are demonstrated in a continuous value in 0-1. Trade-offs between Probability of Detection (POD)&nbsp;and False Alarm Ratio (FAR) can be optimized by selection of different thresholds.&nbsp;</p> <p>Other datasets:</p> <p>1. Dataset for training. It is for the public use of machine learning training for the current model and no-AOD model (training input features vary).</p> <p>2. PM2.5 dataset.&nbsp;The real-time spatially continuous and hourly-level PM<sub>2.5</sub>&nbsp;dataset is obtained following a published method by Zeng&nbsp;&nbsp;et al..&nbsp;In this method, the fundamental in-situ measurements are obtained from Air Quality System&nbsp;(AQS) monitoring network operated by United States Environmental Protection Agency.</p> <p>Reference:</p> <p>Siwei Li, Ge Song, Jia Xing et al. Lightning nowcasting with aerosol-informed machine learning and satellite-enriched dataset, 14 March 2023, PREPRINT (Version 1) available at Research Square [https://doi.org/10.21203/rs.3.rs-2616886/v1]</p> <p>Zeng, Z.&nbsp;et al.&nbsp;Estimating hourly surface PM2. 5 concentrations across China from high-density meteorological observations by machine learning. Atmospheric Research&nbsp;254, 105516 (2021).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Classifying protein kinase conformations with machine learning: data

<p>This data collection accompanies the manuscript &quot;Classifying protein kinase conformations with machine learning&quot;.</p> <p>It is created using the&nbsp;<a href="https://github.com/edikedik/kinactive">kinactive</a>&nbsp;v0.1&nbsp;tool written in pure Python v3.10. <strong>Note that the data are&nbsp;provided for the reference and reproducibility purposes and will not be compatible with later versions of&nbsp;`kinactive` built upon&nbsp;<a href="https://github.com/edikedik/lXtractor">lXtractor</a> &gt;&nbsp;0.1.1.</strong> Refer to the&nbsp;<a href="https://kinactive.readthedocs.io/en/latest/index.html">kinactive documentation</a>&nbsp;for instructions on how to obtain an actualized version of the structural kinome collection.</p> <p>File descriptions:</p> <ul> <li>db_v3.tar.gz -- a structural kinome collection archive. One can unpack it and inspect the contents or&nbsp;load it into the Python interpreter using `kinactive` or `lXtractor` tools.</li> <li>db_af2.tar.gz -- an AlphaFold2 kinome collection for Swiss-Prot sequences.</li> <li>default_*_vs.tsv -- structure/sequence variables calculated with lXtractor and used in an interpretable ML pipeline.</li> <li>*_features.tsv -- lists of ranked features selected by the <a href="https://github.com/edikedik/eBoruta">eBoruta</a> tool for each classifier.</li> <li>Supplement_labels.tsv -- ML model predictions for each PK domain structure found in db_v3.</li> <li>predictions_af2.csv -- Active/Inactive and DFG labels predicted for domains in db_af2.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Machine Learning Features from Proton Therapy Treatment Simulations with the Bergen DTC Prototype for Range Verification

<p>Extracted features from the simulation data found at DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.8192778">10.5281/zenodo.8192778</a></p> <p>Each simulation constitutes a single data sample. The following features were extracted.</p> <p>Detector features:</p> <ul> <li>Total number of active pixels</li> <li>Total number of clusters (hits)</li> <li>Number of clusters over threshold (5, 20 pixels)</li> <li>Mean and standard deviation of cluster sizes</li> <li>The number of clusters of any given size (1&ndash;72)</li> <li>Mean and standard deviation of x- and y-coordinates over each layer (0&ndash;42), and the entire detector</li> <li>Number of active pixels in each layer (0&ndash;42)</li> <li>Number of clusters (hits) in each layer (0&ndash;42)</li> <li>Total energy deposition of the hits in each layer (0&ndash;42)</li> </ul> <p>Higher-level detector features, i.e., function fits (linear, cubic, exponential) with their mean squared residuals&nbsp;over&nbsp;the following quantities:</p> <ul> <li>Active pixels over layer</li> <li>Number of clusters over layer</li> <li>Total deposited energy over layer</li> </ul> <p>201 RSP features extracted from the beam spot, the phantom rotation, and its 3D RSP image.</p> <p>Two datasets are included in two separate archive files:</p> <ul> <li><strong>features.tar.gz:</strong> 715-HN phantom by CIRS Inc. (Norfolk, VA, United States), digitized by Giacometti et al. (2017).</li> <li><strong>features-vhf.tar.gz:</strong> The Visible Human Female (VHF) Head phantom (Ackermann et al. 1995), courtesy of the U.S. National Library of Medicine, resampled&nbsp;to 1 mm voxels and scaled down to 80% size in the simulation.</li> </ul> <p>After extracting features, some outliers were removed from the datasets: 14&nbsp;samples for 715-HN and 3 samples for VHF. The rest of the samples were split into train (70%), validation (10%), and test (20%) sets, for both phantoms separately, which can be found in separate CSV files: features_train.csv, features_val.csv, features_test.csv (715-HN) and features-vhf_train.csv, features-vhf_val.csv, features-vhf_test.csv (VHF).</p> <p>The last file (features_shifted_test.csv (715-HN) and features-vhf_shifted_test.csv (VHF)) contains 40 additional samples for each data point in the respective test set, representing a simulated lateral shift between 1 mm and 10 mm in 1 mm intervals in all directions along the x- and y-axis of the beam.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Data deposit accompanying Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields

<p>Dataset accompanying the paper: <em>&quot;Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields&quot;</em>. Contains the training sets curated during active learning as well as .xyz files used for creating the Figures.&nbsp;<br> <br> The paper highlights that the computational efficiency of ML force fields not only results in decreased computational costs for routine catalytic investigations but also facilitates more comprehensive exploration of catalytic pathways.</p> <p><strong>Published in NPJ Computational Materials</strong>:&nbsp;<a href="https://www.nature.com/articles/s41524-023-01124-2">https://www.nature.com/articles/s41524-023-01124-2</a><br> Formerly on Arxiv:&nbsp;<a href="https://arxiv.org/abs/2301.09931">https://arxiv.org/abs/2301.09931</a></p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Machine learning methods for gap-filling in greenhouse gas emissions databases

<p>Datasets for use with code related to &quot;Machine learning methods for gap-filling in greenhouse gas emissions databases&quot; manuscript submitted to the Journal of Industrial Ecology. Code for using the datasets can be found at&nbsp;<a href="https://github.com/luke-scot/ml-ghg-databases">https://github.com/luke-scot/ml-ghg-databases</a>.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Field Line Resonances estimated using Machine Learning methods

<p>This data set contains the machine learning input matrix (composed by 1D Fourier cross-spectra) + additional information, for the Classification algorithm implemented in Foldes et al. (Automatic Detection of Field Line Resonance Frequencies in the Earth&rsquo;s Plasmasphere, 2023) for the pair of station Tartu-Birzai (TAR-BRZ).</p> <p>Each file contains the following header at line 1. Columns are:</p> <p>- P(f0)-P(f211): Cross-phase value per frequency bin</p> <p>- YEAR</p> <p>- DOY (Day Of Year)</p> <p>- HOUR</p> <p>- ToD_flag: &quot;Umbra&quot;, &quot;Penumbra&quot;, &#39;Light&#39;</p> <p>- L: McIllwain parameter</p> <p>- stat_tag: &quot;tarbrz&quot;</p> <p>- Kp</p> <p>- Kp_w_05d: Kp index weighted on a 12hrs time window</p> <p>- Kp_w_10d: Kp index weighted on a 24hrs time window</p> <p>- Kp_w_15d: Kp index weighted on a 36hrs time window</p> <p>- Kp_w_20d: Kp index weighted on a 2-day time window</p> <p>- Kp_w_25d: Kp index weighted on a 2.5-day time window</p> <p>- Kp_w_30d: Kp index weighted on a 3-day time window</p> <p>- Kp_m_05d: Kp index max on a 12hrs time window</p> <p>- Kp_m_10d: Kp index max on a 24hrs time window</p> <p>- Kp_m_15d: Kp index max on a 36hrs time window</p> <p>- Kp_m_20d: Kp index max on a 2-day time window</p> <p>- Kp_m_25d: Kp index max on a 2.5-day time window</p> <p>- Kp_m_30d: Kp index max on a 3-day time window</p> <p>- DST</p> <p>- DST_m_05d: DST index min on a 12hrs time window</p> <p>- DST_m_10d: DST index min on a 24hrs time window</p> <p>- DST_m_15d: DST index min on a 36hrs time window</p> <p>- DST_m_20d: DST index min on a 2-day time window</p> <p>- DST_m_25d: DST index min on a 2.5-day time window</p> <p>- DST_m_30d: DST index min on a 3-day time window</p> <p>- F107: F10.7 solar activity proxy</p> <p>- EField: Earth Electric co-rotation field</p> <p>- f(mHz): FLR frequency in mHz</p> <p>- df(mHz): Uncertainty on the validated frequency</p> <p>- class: 0 for &quot;NoFreq&quot;, 1 for &quot;Freq&quot; and 2 for &quot;PBL&quot;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Technical Debt Prioritization Using Machine Learning

<p>Technical debt (TD) identification tools can find thousands of technical debt items (TDIs) in a software project. Remedying all of them would take months or even years, so prioritization and decision-making are needed to make this process efficient. On the other hand, advances in machine learning over the last few decades have allowed researchers to apply methods to cluster behaviors and identify patterns in software engineering data. In this study, we aim to develop machine learning methods to decide whether and when a given TDI should be paid off in \st{real} software projects. We performed a survey to collect data from Java open-source software projects hosted on GitHub. From the 2,616 survey responses, we created a dataset using three different labeling strategies - &quot;pay or not&quot;, 3-classes, and priority. We applied nine well-known machine learning methods over 27 source code metrics to build models to predict if and when a TDI should be paid off. The best methods for determining whether an item should be paid off achieved a mean accuracy of 0.86 and an F1-score of 0.85. For when to make the payment, we applied four approaches. Their performance achieved an accuracy of 0.59 using traditional analysis and 0.83 with tuned analysis for the most flexible method.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Machine Learning Applications in Marketing: Literature Review and Research Agenda

<p>Currently, machine learning applications in marketing allow to optimize strategies, personalize experiences and improve decision making. However, there are still several research gaps, so the objective is to examine the research trends in the use of machine learning in marketing. A bibliometric analysis is proposed to assess the current scientific activity, following the parameters established by PRISMA-2020. Machine learning applications in marketing have experienced steady growth and increased attention in the academic community. Key references, such as Miklosik and Evans, and prominent journals, such as IEEE Access and Journal of Business Research, have been identified. A thematic evolution towards big data and digital marketing is observed, and thematic clusters such as &quot;digital marketing&quot;, &quot;interpretation&quot;, &quot;prediction&quot;, and &quot;healthcare&quot; stand out. These findings demonstrate the continued importance and research potential of this evolving field.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Ensemble Machine Learning Prediction of Potential FAPAR: Monthly time-series 2021 and Long-Term Comparison with Actual FAPAR

<p><strong>General Description</strong></p> <p>The dataset contains composites at 250 m spatial resolution of (1) &nbsp;monthly potential FAPAR for the year 2021 from ensemble ML model predictions, (2) the model deviance for each prediction, (3) the yearly average of potential FAPAR, (4) the yearly average of actual FAPAR and (5) the yearly average of the difference between actual and potential (actual minus potential) FAPAR. The dataset is based on the <a href="https://zenodo.org/record/8392976">95th percentile of the monthly aggregated FAPAR</a>&nbsp;derived from&nbsp;<a href="http://glass.umd.edu/Overview.html">250&thinsp;m 8&thinsp;d GLASS V6 FAPAR</a>. Potential FAPAR was predicted by fitting an ensemble ML model using globally distributed training points (cca 3 Mio) and a set of 52 biophysical covariates including several layers related to human pressure. The code for modeling potential FAPAR is openly available at <a href="http://github.com/Open-Earth-Monitor/Global_FAPAR_250m">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m</a>. The dataset can be used in many applications like land degradation modeling, land productivity mapping, and land potential mapping.&nbsp;</p> <p><strong>Data Details</strong></p> <ul> <li><strong>Time period:</strong> January 2021 - December 2021</li> <li><strong>Type of data: </strong>Fraction of Absorbed Photosynthetically Active Radiation (FAPAR)</li> <li><strong>How the data was collected or derived:</strong> Derived from 250m 8 d GLASS V6 FAPAR</li> <li><strong>Statistical methods used: </strong>Ensemble machine learning</li> <li><strong>Limitations or exclusions in the data: </strong>The dataset does not include data for Antarctica.</li> <li><strong>Coordinate reference system:</strong> EPSG:4326</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.0008094, 179.9999424, 87.37000)</li> <li><strong>Spatial resolution:</strong> 1/480 d.d. = 0.00208333 (250m)</li> <li><strong>Image size: </strong>172,800 x 71,698</li> <li><strong>File format: </strong>Cloud Optimized Geotiff (COG) format.</li> </ul> <p><strong>Support</strong></p> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: <a href="https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues</a></p> <p><strong>Reference</strong></p> <p>Hackl&auml;nder, J., Parente, L., Ho, Y.-F., Hengl, T., Simoes, R., Consoli, D., Şahin, M., Tian, X., Herold, M., Jung, M., Duveiller, G., Weynants, M., Wheeler, I., (2023?) &quot;Land potential assessment and trend-analysis using 2000&ndash;2021 FAPAR monthly time-series at 250 m spatial resolution&quot;, submitted to PeerJ, preprint available at: <a href="https://doi.org/10.21203/rs.3.rs-3415685/v1">https://doi.org/10.21203/rs.3.rs-3415685/v1</a></p> <p>&nbsp;</p> <p><strong>Name convention</strong></p> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p> <ol> <li><strong>generic variable name:</strong> pot.fapar = Potential Fraction of Absorbed Photosynthetically Active Radiation</li> <li><strong>variable procedure combination: </strong>eml = ensemble machine learning</li> <li><strong>Position in the probability distribution / variable type:</strong> m = mean</li> <li><strong>Spatial support:</strong> 250m</li> <li><strong>Depth reference: </strong>s = surface</li> <li><strong>Time reference begin time:</strong> 20210101 = 2021-01-01</li> <li><strong>Time reference end time:</strong> 20211231 = 2021-12-31</li> <li><strong>Bounding box: </strong>go = global (without Antarctica)</li> <li><strong>EPSG code:</strong> epsg.4326 = EPSG:4326</li> <li><strong>Version code:</strong> v20230924 = 2023-09-24 (creation date)</li> </ol>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record