Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

26

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

26 results for “ensemble machine learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods

<p>This is the R-code as well as the underlying data&nbsp;needed to reproduce the results of the springer book chapter: &quot;Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods&quot;</p> <p>For more information contact:&nbsp;dlieske@mta.ca</p> <p>&nbsp;</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

CREMP-CycPeptMPDB: Conformer-rotamer ensembles of macrocyclic peptides for machine learning with permeability annotations

<p>CREMP-CycPeptMPDB: A resource generated for the rapid development and evaluation of machine learning models for permeable macrocyclic peptides. CREMP-CycPeptMPDB contains 3,258 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 8.7 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations and with experimental membrane permeability measurements obtained from the <a href="http://cycpeptmpdb.com/" target="_blank" rel="noopener">CycPeptMPDB</a> database. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>This dataset complements the <a title="CREMP" href="../doi/10.5281/zenodo.7931444" target="_blank" rel="noopener">CREMP dataset</a>, which contains a larger selection of conformer ensembles for homodetic macrocyclic peptides.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the &ldquo;pickle&rdquo; folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single &ldquo;pickle.tar.gz&rdquo; archive. In the &ldquo;sdf_and_json&rdquo; folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, &ldquo;sdf_and_json.tar.bz2&rdquo;. A single summary CSV file is also provided containing &rdquo;sequence&rdquo;, &ldquo;smiles&rdquo;, &ldquo;num_monomers&rdquo;, &ldquo;num_atoms&rdquo;, &ldquo;num_heavy_atoms&rdquo;, along with the CREST metadata &ldquo;totalconfs&rdquo;, &ldquo;uniqueconfs&rdquo;, &ldquo;lowestenergy&rdquo;, &ldquo;poplowestpct&rdquo;, &ldquo;temperature&rdquo;, &ldquo;ensembleenergy&rdquo;, &ldquo;ensembleentropy&rdquo;, and &ldquo;ensemblefreeenergy&rdquo;. The number of unique conformers with different 3D structures is given by &ldquo;uniqueconfs&rdquo;, while &ldquo;totalconfs&rdquo; includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 13 GB for "pickle.tar.gz" and 84 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 1)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 2)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types classification and relative entropy</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CREMP: Conformer-rotamer ensembles of macrocyclic peptides for machine learning

<p>CREMP:&nbsp;A&nbsp;resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides. CREMP contains 36,198 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 31.3 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the &ldquo;pickle&rdquo; folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single &ldquo;pickle.tar.gz&rdquo; archive. In the &ldquo;sdf_and_json&rdquo; folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, &ldquo;sdf_and_json.tar.bz2&rdquo;. A single summary CSV file is also provided containing &rdquo;sequence&rdquo;, &ldquo;smiles&rdquo;, &ldquo;num_monomers&rdquo;, &ldquo;num_atoms&rdquo;, &ldquo;num_heavy_atoms&rdquo;, along with the CREST metadata &ldquo;totalconfs&rdquo;, &ldquo;uniqueconfs&rdquo;, &ldquo;lowestenergy&rdquo;, &ldquo;poplowestpct&rdquo;, &ldquo;temperature&rdquo;, &ldquo;ensembleenergy&rdquo;, &ldquo;ensembleentropy&rdquo;, and &ldquo;ensemblefreeenergy&rdquo;. The number of unique conformers with different 3D structures is given by &ldquo;uniqueconfs&rdquo;, while &ldquo;totalconfs&rdquo; includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 32 GB for "pickle.tar.gz" and 210 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Ensemble Machine Learning Prediction of Potential FAPAR: Monthly time-series 2021 and Long-Term Comparison with Actual FAPAR

<p><strong>General Description</strong></p> <p>The dataset contains composites at 250 m spatial resolution of (1) &nbsp;monthly potential FAPAR for the year 2021 from ensemble ML model predictions, (2) the model deviance for each prediction, (3) the yearly average of potential FAPAR, (4) the yearly average of actual FAPAR and (5) the yearly average of the difference between actual and potential (actual minus potential) FAPAR. The dataset is based on the <a href="https://zenodo.org/record/8392976">95th percentile of the monthly aggregated FAPAR</a>&nbsp;derived from&nbsp;<a href="http://glass.umd.edu/Overview.html">250&thinsp;m 8&thinsp;d GLASS V6 FAPAR</a>. Potential FAPAR was predicted by fitting an ensemble ML model using globally distributed training points (cca 3 Mio) and a set of 52 biophysical covariates including several layers related to human pressure. The code for modeling potential FAPAR is openly available at <a href="http://github.com/Open-Earth-Monitor/Global_FAPAR_250m">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m</a>. The dataset can be used in many applications like land degradation modeling, land productivity mapping, and land potential mapping.&nbsp;</p> <p><strong>Data Details</strong></p> <ul> <li><strong>Time period:</strong> January 2021 - December 2021</li> <li><strong>Type of data: </strong>Fraction of Absorbed Photosynthetically Active Radiation (FAPAR)</li> <li><strong>How the data was collected or derived:</strong> Derived from 250m 8 d GLASS V6 FAPAR</li> <li><strong>Statistical methods used: </strong>Ensemble machine learning</li> <li><strong>Limitations or exclusions in the data: </strong>The dataset does not include data for Antarctica.</li> <li><strong>Coordinate reference system:</strong> EPSG:4326</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.0008094, 179.9999424, 87.37000)</li> <li><strong>Spatial resolution:</strong> 1/480 d.d. = 0.00208333 (250m)</li> <li><strong>Image size: </strong>172,800 x 71,698</li> <li><strong>File format: </strong>Cloud Optimized Geotiff (COG) format.</li> </ul> <p><strong>Support</strong></p> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: <a href="https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues</a></p> <p><strong>Reference</strong></p> <p>Hackl&auml;nder, J., Parente, L., Ho, Y.-F., Hengl, T., Simoes, R., Consoli, D., Şahin, M., Tian, X., Herold, M., Jung, M., Duveiller, G., Weynants, M., Wheeler, I., (2023?) &quot;Land potential assessment and trend-analysis using 2000&ndash;2021 FAPAR monthly time-series at 250 m spatial resolution&quot;, submitted to PeerJ, preprint available at: <a href="https://doi.org/10.21203/rs.3.rs-3415685/v1">https://doi.org/10.21203/rs.3.rs-3415685/v1</a></p> <p>&nbsp;</p> <p><strong>Name convention</strong></p> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p> <ol> <li><strong>generic variable name:</strong> pot.fapar = Potential Fraction of Absorbed Photosynthetically Active Radiation</li> <li><strong>variable procedure combination: </strong>eml = ensemble machine learning</li> <li><strong>Position in the probability distribution / variable type:</strong> m = mean</li> <li><strong>Spatial support:</strong> 250m</li> <li><strong>Depth reference: </strong>s = surface</li> <li><strong>Time reference begin time:</strong> 20210101 = 2021-01-01</li> <li><strong>Time reference end time:</strong> 20211231 = 2021-12-31</li> <li><strong>Bounding box: </strong>go = global (without Antarctica)</li> <li><strong>EPSG code:</strong> epsg.4326 = EPSG:4326</li> <li><strong>Version code:</strong> v20230924 = 2023-09-24 (creation date)</li> </ol>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Artifact for "MAAT: A Novel Ensemble Approach to Addressing Fairness and Performance Bugs for Machine Learning Software"

<p>This artifact is for the paper entitled &ldquo;MAAT: A Novel Ensemble Approach to Addressing Fairness and Performance Bugs for Machine Learning Software&rdquo;, which is accepted by ESEC/FSE 2022. MAAT is a novel ensemble approach to improving the fairness-performance trade-off for ML software. It outperforms state-of-the-art bias mitigation methods. The artifact has also been placed on GitHub (https://github.com/chenzhenpeng18/FSE22-MAAT) under the Apache License, publicly accessible to other researchers. In this artifact, we provide the source code of MAAT and other existing bias mitigation methods that we use in our study, as well as the intermediate results, the installation instructions, and a replication guideline (included in the README). The replication guideline provides detailed steps to replicate all the results for all the research questions.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Ensemble of optimised machine learning algorithms for predicting surface soil moisture content at global scale (v1.0)

<p>This study investigates the estimation of daily SSM using eight optimised ML algorithms and ten ensemble models (constructed via model bootstrap aggregating techniques and five-fold cross-validation). The algorithmic implementations were trained and tested using the international soil moisture network (ISMN) data collected from 1722 stations distributed across the World.&nbsp;</p>

openother-openJun 2023View details →
zenodo36/100

Ensemble BLUP, Machine Learning, and Deep Learning Models Predict Maize Yield Better Than Each Model Alone.

<p>Data and scripts exploring ensembling strategies using the models developed in <a href="https://academic.oup.com/g3journal/advance-article/doi/10.1093/g3journal/jkad006/6982634">Kick et al., 2023</a> (see also <a href="https://zenodo.org/record/7401113">1</a>, <a href="https://zenodo.org/record/6916775">2</a>). Download all files to a single directory then run setup.sh or manually unzip using tar.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Filename</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>setup.sh</td> <td>Simple script that unzips zipped directories</td> </tr> <tr> <td>ext_data</td> <td>Reduced data from Kick et al. 2023</td> </tr> <tr> <td>ext_data_notebooks</td> <td>Contains python notebooks containing analysis and R markdown file containing visualization of results. Python and R data objects are written to allow results to be read in instead of re-generated.</td> </tr> <tr> <td>output</td> <td>Folder containing a placeholder file.</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This research used resources provided by the United States Department of Agriculture&rsquo;s Agricultural Research Service (project number 5070-21000-041-000-D). The SCINet project of the USDA Agricultural Research Service (project number 0500-00093-001-00-D) was instrumental in the training of the models used in this work. In addition, we would like to acknowledge those presently and historically involved in generating data for the Genomes to Fields Initiative.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-3.0-usMar 2023View details →
dryad36/100

An ensemble machine learning bioavailable strontium isoscape for Eastern Canada

Open the record for dataset details and reuse information.

publicMay 2025View details →
zenodo32/100

Input data and some models (all except multi-model ensembles) for JAMES paper "Machine-learned uncertainty quantification is not magic"

<p>The tar file contains two directories: data and models. &nbsp;Within "data," there are 4 subdirectories: "training" (the clean training data -- without perturbations), "training_all_perturbed_for_uq" (the lightly perturbed training data), "validation_all_perturbed_for_uq" (the moderately perturbed validation data), and "testing_all_perturbed_for_uq" (the heavily perturbed validation data). &nbsp;The data in these directories are unnormalized. &nbsp;The subdirectories "training" and "training_all_perturbed_for_uq" each contain a normalization file. &nbsp;These normalization files contain parameters used to normalize the data (from physical units to z-scores) for Experiment 1 and Experiment 2, respectively. &nbsp;To do the normalization, you can use the script normalize_examples.py in the code library (ml4rt) with the argument input_normalization_file_name set to one of these two file paths. &nbsp;The other arguments should be as follows:</p><p>--uniformize=1</p><p>--predictor_norm_type_string="z_score"</p><p>--vector_target_norm_type_string=""</p><p>--scalar_target_norm_type_string=""</p><p>&nbsp;</p><p>Within the directory "models," there are 6 subdirectories: for the BNN-only models trained with clean and lightly perturbed data, for the CRPS-only models trained with clean and lightly perturbed data, and for the BNN/CRPS models trained with clean and lightly perturbed data. &nbsp;To read the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Landslide susceptibility maps using ensemble machine learning models on basin and regional level in Lombardy, Italy

<p>A selection of landslide susceptibility maps computed through ensemble machine learning models for the basins of Val Tartano, Upper Valtellina and Valchiavenna, and on a regional level for the Lombardy region in Italy.</p> <p>A list of the used base machine learning methods:</p> <ul> <li>Random Forest,</li> <li>AdaBoost,</li> <li>Neural Networks.</li> </ul> <p>A list of the used ensemble models:</p> <ul> <li>Stacking,</li> <li>Blending,</li> <li>Soft Voting.</li> </ul> <p>A full list of the model combinations can be found in the "Case Studies" document.</p> <p>The maps are in WGS 84/ UTM zone 32N (EPSG:32632).</p> <p>The map production process details are discussed in Xu et al. 2024. If you use the dataset, please, cite also the paper:</p> <p><em>Qiongjie Xu, Vasil Yordanov, Lorenzo Amici &amp; Maria Antonia Brovelli (2024) Landslide susceptibility mapping using ensemble machine learning methods: a case</em><br><em>study in Lombardy, Northern Italy, International Journal of Digital Earth, 17:1, 2346263, DOI:10.1080/17538947.2024.2346263</em></p> <p>The maps are produced as part of the "Geoinformatics and Earth Observation for Landslide Monitoring" Italy-Vietnam.</p> <p>The work is partially funded by the Italian Ministry of Foreign Affairs and International Cooperation within the project &ldquo;Geoinformatics and Earth Observation for Landslide Monitoring&rdquo; CUP D19C21000480001.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Supporting data and code for the published paper: Machine Learning Nonadiabatic Dynamics: Eliminating Phase Freedom of Nonadiabatic Couplings with the State-Interaction State-Averaged Spin-Restricted Ensemble-Referenced Kohn–Sham Approach

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Machine learning based methods to generate conformational ensembles of disordered proteins (len54)

<p>data is in rep_1 for all sequences, which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy)</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Machine learning based methods to generate conformational ensembles of disordered proteins (len18, point mutation)

<p>data is organized by bin number (0-9) and mutation location (4, 8, 12). (Note: in the manuscript, we used the nomenclature bins 1-10 and mutation&nbsp;locations 5,8,13. We simply used a 0-index convention when naming our folders). All data is in rep_1,&nbsp;which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy). However, for bin_num=9/mutation_loc=12, the data used is in rep_2, not rep_1.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Machine learning based methods to generate conformational ensembles of disordered proteins (len36)

<p>data is in rep_1 for all sequences, which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy)</p>

opencc-by-4.0Oct 2023View details →
zenodo28/100

Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Multiclass Classification of Decisions: A Study of the Hibernate Developer Mailing List"

<p>This is the replication package for the paper: &quot;A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List&quot;.&nbsp;It contains the source code and dataset of our experiment for the&nbsp;replication&nbsp;by&nbsp;other&nbsp;researchers. In the meanwhile, we provide brief description of the files in the replication&nbsp;package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py&nbsp;&nbsp;</em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0.&nbsp;<strong>Note that you may&nbsp;get slightly</strong>&nbsp;<strong>different experiment&nbsp;results when conducting the experiments&nbsp;on different environment configurations.</strong></li> <li><em>requirement.txt</em>&nbsp; records all the installation packages and their version numbers needed for the current program to run.&nbsp;You&nbsp;can use &quot;<em>pip install -r requirement.txt</em>&quot; to rebuild the project and install all dependencies. <strong>Note that you may&nbsp;get slightly different experiment&nbsp;results when using different packages or versions.&nbsp;</strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx&nbsp;&nbsp;</em>contains 844&nbsp;labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Supplementary Data for "Predicting Thermodynamic Stability of Inorganic Compounds Using Ensemble Machine Learning Based on Electron Configuration"

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Enhancing Streamflow Prediction through Multi-model Ensemble Framework and Machine Learning Techniques

<p>This file contains python code used in this study and data used to plot figures.&nbsp;</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record