Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.9.0
Dataset results
26 results for “ensemble methods”
EPTGODD-WHU: Ensemble Precipitation and Temperature from CMIP6 GCMs optimized by OLS-DT-DNN methods integration (1850-2100)
<p>This monthly global climate dataset EPTGODD-WHU (precipitation and mean temperature variables with grid size of 0.5°×0.5°) was ensembled from 16 selected CMIP6 GCMs. The published dataset was optimized by OLS (Ordinary Linear Square)-DT (Decision Tree)-DNN (Deep Neural Network) methods integration. The CF (Climate and Forecast) v1.6 was employed as the guideline for NetCDF4 format. The periods of temperature files can be divided into historical (1850-1900) and future (2015-2100) periods. For precipitation, this product provides future (2015-2100) period. Three future scenarios (SSP1-2.6, SSP2-4.5 and SSP5-8.5) were selected for both variables. The units of this dataset are degrees Celsius and mm/month for temperature and precipitation, respectively. Each NetCDF4 file in this dataset includes three dimensions (time, latitude (-89.75°N to 89.75°N) and longitude (-179.75°E to 179.75°E)).</p>
R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods
<p>This is the R-code as well as the underlying data needed to reproduce the results of the springer book chapter: "Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods"</p> <p>For more information contact: dlieske@mta.ca</p> <p> </p>
Figure 4 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 4. Mean suitability for Arborophila crudigularis under baseline climate conditions and future climate scenarios. cccma and csiro represent two general circulation models; RCP2.6, and RCP8.5 represent two greenhouse gas emission scenarios; EN is entire suitable habitat; PR is presence records.
Figure 2 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 2. Performance of each model for predicting the suitable habitat for Arborophila crudigularis. GLM: Generalized linear model; GBM: generalized boosting model; GAM: generalized additive model; CTA: classification tree analysis; ANN: artificial neural network; FDA: flexible discriminant analysis; MARS: multiple adaptive regression splines; RF: random forest; MAXENT: maximum entropy model.
Figure 6 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 6. Changes in suitable habitat for Arborophila crudigularis under the RCP8.5 emission scenario. cccma and csiro represent two general circulation models.
Figure 5 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 5. Changes in suitable habitat for Arborophila crudigularis under the RCP2.6 emission scenario. cccma and csiro represent two general circulation models.
Datasets for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants"
<p>All data used for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants", Criswell, A.W., et al. (2022). The code used to create the paper results from this data can be found at <a href="https://github.com/criswellalexander/hbpm_paper">https://github.com/criswellalexander/hbpm_paper</a> and the underlying software package can be found at <a href="https://github.com/criswellalexander/bayestack">https://github.com/criswellalexander/bayestack</a>.</p>
SAXS calculations for Refinement of 𝛼-synuclein ensembles against SAXS data: Comparison of force fields and methods
<p>SAXS curves calculated from MD simulations used as input to reweighting as described in preprint Refinement of 𝛼-synuclein ensembles against SAXS data: Comparison of force fields and methods</p>
Supplementary Data for "Comparison of Ensemble-Based Data Assimilation Methods for Sparse Oceanographic Data"
<p>This data set represents the supplementary data for the paper Comparison of Ensemble-Based Data Assimilation Methods for Sparse Oceanographic Data (Section 4) by Florian Beiser, Håvard Heitlo Holm, and Jo Eidsvik.</p><p>It contains the data that is plotted in the manuscript. The code for plotting is provided in the supplementary software.</p>
EnGRaiN : A Supervised Ensemble Learning Method for Recovery of Large-scale Gene Regulatory Networks
<p>EnGRaiN is a supervised machine learning method to construct ensemble networks. To benefit from the typical accuracy advantages of supervised learning methods while taking into account the impossibility of knowing true networks for training, we devised a method that uses small training datasets of true positives and true negatives among gene pairs.</p> <p>The datasets used to evaluate the performance of EnGaiN include (i) simulated datasets generated from Yeast networks and (ii) A. thaliana gene expression datasets.</p>
Figure 1 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 1. Modeled range and presence records for Arborophila crudigularis.
Molecular dynamics-generated ensemble dataset of ubiquitin; for "PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles"
<p>The molecular dynamics-generated ensemble dataset (229Mb zip file) for ubiquitin, used in the manuscript "PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles", submitted to the Journal of Chemical Information and Modeling (JCIM). The dataset consists of 6 .dcd files, and one .pdb file. </p>
Data release for 'Ensemble Forecasting of Major Solar Flares: Methods for Combining Models'
<p>This is a release of the data that were used for validation in the paper 'Ensemble Forecasting of Major Solar Flares: Methods for Combining Models' by J. A. Guerra, S. A. Murray, D. S. Bloomfield, and P. T. Gallagher, that has been submitted to the Journal of Space Weather and Space Climate.</p> <p> </p> <ul> </ul> <p>The naming scheme for the files is in the format:</p> <pre><code>class_type_metric.dat</code></pre> <ul> <li>'class' denotes whether the forecast is for M- or X- class flares.</li> <li>'type' is what kind of forecast, i.e., the original ensemble members, an ensemble created from probabilistic validation metrics, or an ensemble created from categorical validation metrics.</li> <li>'metric' specifies the metric used to create the ensemble in the case of 'probabilistic' or 'categorical' types as above (see paper for further details), or in the case of the original ensemble members the name of the operational forecasting method.</li> </ul> <p> </p> <p>The data files are in the format:</p> <pre><code>obs,prob</code></pre> <ul> <li>'obs' denotes whether or not a flare was observed within 24 hours of the forecast issue time (1 for yes and 0 for no).</li> <li>'prob' gives the probabilistic forecast value (between 0.0 and 1.0).</li> </ul> <p> </p> <p>These data files can easily be read into the <a href="https://cran.r-project.org/web/packages/verification/verification.pdf">R verification package</a> to replicate the results presented in the paper.</p>
ALCC: Temporally consistent annual land cover maps over China from 1985 to 2022 based on an ensemble change detection method
<p><span><span>We develope <span><span> </span>a consistent annual land cover product for China (ALCC) from 1985 to 2022. First, change areas for each year was detected baesd on an ensemble change detection method combing CCDC, BFASTm, and Chow Test. Then, stable training samples were derived from both CLCD and CLUDs. The Random Forest classifier was locally trained and used to classify change areas year by year. Finally, ALCC was generated by updating land cover classification results for change areas and remaining class label the same as the base map forthe unchanged areas. ALCC achieved a mean overall accuracy of 81.11±0.67% nationwide and exceeded 72.00% across seven geographical regions.</span></span></span></p>
Deep Ensemble Learning and Transfer Learning Methods for Classification of Senescent Cells from Nonlinear Optical Microscopy Images
<p>This Dataset contains the train and test NLO images in pickle format used for the following publication: Deep Ensemble Learning and Transfer Learning Methods for Classification of Senescent Cells from Nonlinear Optical Microscopy Images</p>
Machine learning based methods to generate conformational ensembles of disordered proteins (len54)
<p>data is in rep_1 for all sequences, which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy)</p>
Machine learning based methods to generate conformational ensembles of disordered proteins (len18, point mutation)
<p>data is organized by bin number (0-9) and mutation location (4, 8, 12). (Note: in the manuscript, we used the nomenclature bins 1-10 and mutation locations 5,8,13. We simply used a 0-index convention when naming our folders). All data is in rep_1, which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy). However, for bin_num=9/mutation_loc=12, the data used is in rep_2, not rep_1. </p>
Machine learning based methods to generate conformational ensembles of disordered proteins (len36)
<p>data is in rep_1 for all sequences, which contains the trajectory (xtc) file, Rg information (in the file Rg.out), pairwise distance information (in the file traj_analysis_data/pairwise_distance_matrix.csv) and the bspline coefficients (in the file bspline_info/xyz_coeff.npy)</p>
Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Multiclass Classification of Decisions: A Study of the Hibernate Developer Mailing List"
<p>This is the replication package for the paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py </em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx </em>contains 844 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>
HIV Tropism Ensemble Methods
<p>Prediction of HIV-1 subtype C tropism using ensemble learning with genotypic algorithms. Made using R version 3.6.0.</p> <p>Original data used to generate models is in /originalData</p> <p>--------------------------</p> <p>Previsão do tropismo do HIV-1 subtipo C com stacking de testes genotípicos. Feito com R, versão 3.6.0.</p> <p>Dados originais usados para gerar os modelos em /originalData</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.