Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forest”
Data for: A new paradigm for medium-range severe weather forecasts: Probabilistic random forest-based predictions
Open the record for dataset details and reuse information.
River metabolism in the contiguous United States: Random forest model code, inputs and outputs
Open the record for dataset details and reuse information.
Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids
Open the record for dataset details and reuse information.
Random forest modelling of multi-scale, multi-species habitat associations within KAZA transfrontier conservation area using spoor data
Open the record for dataset details and reuse information.
Fluoro-Forest CODEX data for random Forest-based cell type annotation
Open the record for dataset details and reuse information.
MetaComNet: A random forest-based framework for making spatial prediction of plant-pollinator interactions
Open the record for dataset details and reuse information.
IIb-RAD-seq coupled with random forest classification indicates regional population structuring and sex-specific differentiation in salmon lice (Lepeophtheirus salmonis)
Open the record for dataset details and reuse information.
Comparing mixed models and Random Forest association tests using naturalGWAS and a Striped Bass SNP dataset
Open the record for dataset details and reuse information.
Training Data For Random Forest Classification
<p>This is the training data for land use/cover classification and it was collected by doing visual interpretation from Google Earth Pro.</p>
Data from: Integration of Random Forest with population-based outlier analyses provides insight on the genomic basis and evolution of run timing in Chinook salmon (Oncorhynchus tshawytscha)
Anadromous Chinook salmon populations vary in the period of river entry at the initiation of adult freshwater migration, facilitating optimal arrival at natal spawning. Run timing is a polygenic trait that shows evidence of rapid parallel evolution in some lineages, signifying a key role for this phenotype in the ecological divergence between populations. Studying the genetic basis of local adaptation in quantitative traits is often impractical in wild populations. Therefore, we used a novel approach, Random Forest, to detect markers linked to run timing across 14 populations from contrasting environments in the Columbia River and Puget Sound, USA. The approach permits detection of loci of small effect on the phenotype. Divergence between populations at these loci was then examined using both principle component analysis and FST outlier analyses, to determine whether shared genetic changes resulted in similar phenotypes across different lineages. Sequencing of 9107 RAD markers in 414 individuals identified 33 predictor loci explaining 79.2% of trait variance. Discriminant analysis of principal components of the predictors revealed both shared and unique evolutionary pathways in the trait across different lineages, characterized by minor allele frequency changes. However, genome mapping of predictor loci also identified positional overlap with two genomic outlier regions, consistent with selection on loci of large effect. Therefore, the results suggest selective sweeps on few loci and minor changes in loci that were detected by this study. Use of a polygenic framework has provided initial insight into how divergence in a trait has occurred in the wild.
Data from: Applications of random forest feature selection for fine-scale genetic population assignment
Genetic population assignment used to inform wildlife management and conservation efforts requires panels of highly informative genetic markers and sensitive assignment tests. We explored the utility of machine-learning algorithms (random forest, regularized random forest, and guided regularized random forest) compared with FST ranking for selection of single nucleotide polymorphisms (SNP) for fine-scale population assignment. We applied these methods to an unpublished SNP dataset for Atlantic salmon (Salmo salar) and a published SNP data set for Alaskan Chinook salmon (Oncorhynchus tshawytscha). In each species, we identified the minimum panel size required to obtain a self-assignment accuracy of at least 90% using each method to create panels of 50-700 markers Panels of SNPs identified using random forest-based methods performed up to 7.8 and 11.2 percentage points better than FST-selected panels of similar size for the Atlantic salmon and Chinook salmon data, respectively. Self-assignment accuracy ≥90% was obtained with panels of 670 and 384 SNPs for each dataset, respectively, a level of accuracy never reached for these species using FST-selected panels. Our results demonstrate a role for machine-learning approaches in marker selection across large genomic datasets to improve assignment for management and conservation of exploited populations.
Dataset for Ha and Aylward 'Automated classification of giant virus genomes using a random forest model built on trademark protein families'
<ul><li>Genome sets used for model training and testing</li><li>Custom Python script that generated fragmented genomes at random completeness levels</li></ul>
HRFMD (Hydrological model based Random Forest Model Diagnostics) results
<p>Results accompanying the publication titled: Advancing Hydrological Model Diagnostics: An Exploratory Approach Using Random Forest Models and Large-sample Catchment Dataset</p>
A Grid Model for Vertical Correction of Precipitable Water Vapor over the Chinese Mainland and Surrounding Areas Using Random Forest
<p>Code to reproduce the work in the manuscript ' A Grid Model for Vertical Correction of Precipitable Water Vapor over the Chinese Mainland and Surrounding Areas Using Random Forest', Junyu Li, Yuxin Wang, Lilong Liu, Yibin Yao, Liangke Huang, Feijuan Li, submitted to GMD.</p>
Datasets and relevant code in the Manuscript "Estimation of fire counts and fire radiative power using satellite optical and microwave vegetation indices with random forest method"
<p>1. multiyears_season_fire_ndvi_fwi_edvi_0.25.mat<br>Multiyear averages of ln (FC), ln (FRP), DMC, ISI, EDVI10-18, EDVI18-36, and NDVI over East Asia in 2003–2010</p> <p>2. RF_edvi_data.mat<br>Estimated FC and FRP based on RF model with EDVIs and NDVI </p> <p>3.RF_fwis_data.mat<br>Estimated FC and FRP based on RF model without EDVIs and NDVI </p> <p>4. temporal_variations.mat<br>East Asia Regional Time Series Dataset</p> <p>5.rf_train_cv_forest_review.py<br>Random forest model python code</p>
Data for Streamflow Prediction: Comparison of SWAT vs. Random Forest Models in Diverse Catchments
<p>This study introduces a time-lag-informed Random Forest (RF) framework for streamflow time series prediction across diverse catchments, and compares its results against SWAT predictions. We found strong evidence of RF's better performance by adding historical flows and time-lags for meteorological values over using only actual meteorological values. On a daily scale, RF demonstrated robust performance (Nash–Sutcliffe efficiency [<em>NSE</em>] > 0.5), whereas SWAT generally yielded unsatisfactory results (<em>NSE</em> < 0.5) and tended to overestimate daily streamflow by up to 27% (<em>PBIAS</em>). However, SWAT provided better monthly predictions, particularly in catchments with irregular flow patterns. Although both models faced challenges in predicting peak flows in snow-influenced catchments, RF outperformed SWAT in an arid catchment. RF also exhibited a notable advantage over SWAT in terms of computational efficiency. Overall, RF is a good choice for daily predictions with limited data, whereas SWAT is preferable for monthly predictions and understanding hydrological processes in depth.</p> <p>This repository contains the input data used for building the RF and SWAT models and the files describing the modeling results.</p> <p>The corresponding Zenodo code repository is available at <a href="../doi/10.5281/zenodo.11064973" target="_blank" rel="noopener">https://zenodo.org/doi/10.5281/zenodo.11064973</a>.</p>
Trained Random Forest model and scaler parameters on new physical and tsfel features from seismic data of 150s length.
Open the record for dataset details and reuse information.
Trained random forest models on 5000 traces per class based on updated features
Open the record for dataset details and reuse information.
A global 0.05° gross primary productivity of sunlit and shaded leaves dataset via combining two-leaf light use efficiency model with random forest over 2002~2020
<p>The TL-CRF model generated a global 0.05´0.05° product for eight-day gross primary productivity (GPP) of sunlit and shaded canopies from 2002 to 2020 by embedding the random forest (RF) submodule into the two-leaf light use efficiency (TL-LUE) model while considering the seasonal differences in the clumping index. The RF technique was used to integrate various environmental stress factors including meteorological, hydrological, soil properties, and elevation, thereby improving the overall scale of the complex environmental conditions to the maximum LUE. This novel GPP product could support further research on spatial and temporal patterns of the carbon cycle and its association with climate change.</p> <p> </p> <p> </p> <p>Variable: GPP, GPP<sub>sh</sub>, and GPP<sub>su</sub></p> <p>Spatial coverage: global</p> <p>Temporal coverage: 2002 to 2020</p> <p>Spatial resolution: 0.05×0.05°</p> <p>Temporal resolution: eight-day</p> <p>Unite: g C m<sup>−2</sup> d<sup>−1</sup></p> <p>Data format: raster (.tif)</p>
Exploring the Multidimensional Impact of Organizational Incentives on the Loyalty of Returned Overseas Faculty: An SEM and Random Forest Analysis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.