Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “Random Forest”

Learn how ShareScore rates datasets ↗
dryad36/100

Data for: A new paradigm for medium-range severe weather forecasts: Probabilistic random forest-based predictions

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad36/100

River metabolism in the contiguous United States: Random forest model code, inputs and outputs

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad36/100

Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad36/100

Random forest modelling of multi-scale, multi-species habitat associations within KAZA transfrontier conservation area using spoor data

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad36/100

Fluoro-Forest CODEX data for random Forest-based cell type annotation

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad36/100

MetaComNet: A random forest-based framework for making spatial prediction of plant-pollinator interactions

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad36/100

IIb-RAD-seq coupled with random forest classification indicates regional population structuring and sex-specific differentiation in salmon lice (Lepeophtheirus salmonis)

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Comparing mixed models and Random Forest association tests using naturalGWAS and a Striped Bass SNP dataset

Open the record for dataset details and reuse information.

publicAug 2022View details →
zenodo32/100

Training Data For Random Forest Classification

<p>This is the training data for land use/cover classification and it was collected by doing visual interpretation from Google Earth Pro.</p>

opencc-by-4.0Jul 2020View details →
dryad32/100

Data from: Integration of Random Forest with population-based outlier analyses provides insight on the genomic basis and evolution of run timing in Chinook salmon (Oncorhynchus tshawytscha)

Anadromous Chinook salmon populations vary in the period of river entry at the initiation of adult freshwater migration, facilitating optimal arrival at natal spawning. Run timing is a polygenic trait that shows evidence of rapid parallel evolution in some lineages, signifying a key role for this phenotype in the ecological divergence between populations. Studying the genetic basis of local adaptation in quantitative traits is often impractical in wild populations. Therefore, we used a novel approach, Random Forest, to detect markers linked to run timing across 14 populations from contrasting environments in the Columbia River and Puget Sound, USA. The approach permits detection of loci of small effect on the phenotype. Divergence between populations at these loci was then examined using both principle component analysis and FST outlier analyses, to determine whether shared genetic changes resulted in similar phenotypes across different lineages. Sequencing of 9107 RAD markers in 414 individuals identified 33 predictor loci explaining 79.2% of trait variance. Discriminant analysis of principal components of the predictors revealed both shared and unique evolutionary pathways in the trait across different lineages, characterized by minor allele frequency changes. However, genome mapping of predictor loci also identified positional overlap with two genomic outlier regions, consistent with selection on loci of large effect. Therefore, the results suggest selective sweeps on few loci and minor changes in loci that were detected by this study. Use of a polygenic framework has provided initial insight into how divergence in a trait has occurred in the wild.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Applications of random forest feature selection for fine-scale genetic population assignment

Genetic population assignment used to inform wildlife management and conservation efforts requires panels of highly informative genetic markers and sensitive assignment tests. We explored the utility of machine-learning algorithms (random forest, regularized random forest, and guided regularized random forest) compared with FST ranking for selection of single nucleotide polymorphisms (SNP) for fine-scale population assignment. We applied these methods to an unpublished SNP dataset for Atlantic salmon (Salmo salar) and a published SNP data set for Alaskan Chinook salmon (Oncorhynchus tshawytscha). In each species, we identified the minimum panel size required to obtain a self-assignment accuracy of at least 90% using each method to create panels of 50-700 markers Panels of SNPs identified using random forest-based methods performed up to 7.8 and 11.2 percentage points better than FST-selected panels of similar size for the Atlantic salmon and Chinook salmon data, respectively. Self-assignment accuracy ≥90% was obtained with panels of 670 and 384 SNPs for each dataset, respectively, a level of accuracy never reached for these species using FST-selected panels. Our results demonstrate a role for machine-learning approaches in marker selection across large genomic datasets to improve assignment for management and conservation of exploited populations.

opencc-zeroDec 2016View details →
zenodo32/100

Dataset for Ha and Aylward 'Automated classification of giant virus genomes using a random forest model built on trademark protein families'

<ul><li>Genome sets used for model training and testing</li><li>Custom Python script that generated fragmented genomes at random completeness levels</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo32/100

HRFMD (Hydrological model based Random Forest Model Diagnostics) results

<p>Results accompanying the publication titled: Advancing Hydrological Model Diagnostics: An Exploratory Approach Using Random Forest Models and Large-sample Catchment Dataset</p>

openapache2.0Nov 2023View details →
zenodo32/100

A Grid Model for Vertical Correction of Precipitable Water Vapor over the Chinese Mainland and Surrounding Areas Using Random Forest

<p>Code to reproduce the work in the manuscript ' A Grid Model for Vertical Correction of Precipitable Water Vapor over the Chinese Mainland and Surrounding Areas Using Random Forest', Junyu Li, Yuxin Wang, Lilong Liu, Yibin Yao, Liangke Huang, Feijuan Li, submitted to GMD.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Datasets and relevant code in the Manuscript "Estimation of fire counts and fire radiative power using satellite optical and microwave vegetation indices with random forest method"

<p>1. multiyears_season_fire_ndvi_fwi_edvi_0.25.mat<br>Multiyear averages of ln (FC), ln (FRP), DMC, ISI, EDVI10-18, EDVI18-36, and NDVI over East Asia in 2003&ndash;2010</p> <p>2. RF_edvi_data.mat<br>Estimated FC and FRP based on RF model with EDVIs and NDVI&nbsp;</p> <p>3.RF_fwis_data.mat<br>Estimated FC and FRP based on RF model without EDVIs and NDVI&nbsp;</p> <p>4. temporal_variations.mat<br>East Asia Regional Time Series Dataset</p> <p>5.rf_train_cv_forest_review.py<br>Random forest model python code</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Data for Streamflow Prediction: Comparison of SWAT vs. Random Forest Models in Diverse Catchments

<p>This study introduces a time-lag-informed Random Forest (RF) framework for streamflow time series prediction across diverse catchments, and compares its results against SWAT predictions. We found strong evidence of RF's better performance by adding historical flows and time-lags for meteorological values over using only actual meteorological values. On a daily scale, RF demonstrated robust performance (Nash&ndash;Sutcliffe efficiency [<em>NSE</em>] &gt; 0.5), whereas SWAT generally yielded unsatisfactory results (<em>NSE</em> &lt; 0.5) and tended to overestimate daily streamflow by up to 27% (<em>PBIAS</em>). However, SWAT provided better monthly predictions, particularly in catchments with irregular flow patterns. Although both models faced challenges in predicting peak flows in snow-influenced catchments, RF outperformed SWAT in an arid catchment. RF also exhibited a notable advantage over SWAT in terms of computational efficiency. Overall, RF is a good choice for daily predictions with limited data, whereas SWAT is preferable for monthly predictions and understanding hydrological processes in depth.</p> <p>This repository contains the input data used for building the RF and SWAT models and the files describing the modeling results.</p> <p>The corresponding Zenodo code repository is available at <a href="../doi/10.5281/zenodo.11064973" target="_blank" rel="noopener">https://zenodo.org/doi/10.5281/zenodo.11064973</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Trained Random Forest model and scaler parameters on new physical and tsfel features from seismic data of 150s length.

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Trained random forest models on 5000 traces per class based on updated features

Open the record for dataset details and reuse information.

openmit-licenseJul 2024View details →
zenodo32/100

A global 0.05° gross primary productivity of sunlit and shaded leaves dataset via combining two-leaf light use efficiency model with random forest over 2002~2020

<p>The TL-CRF model generated a global&nbsp;0.05&acute;0.05&deg; product for eight-day gross primary productivity (GPP) of sunlit and shaded canopies from 2002 to 2020 by embedding the random forest (RF) submodule into the two-leaf light use efficiency (TL-LUE) model while considering the seasonal differences in the clumping index. The RF technique was used to integrate various environmental stress factors including meteorological, hydrological, soil properties, and elevation, thereby improving the overall scale of the complex environmental conditions to the maximum LUE. This novel GPP product could support further research on spatial and temporal patterns of the carbon cycle and its association with climate change.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Variable: GPP, GPP<sub>sh</sub>, and GPP<sub>su</sub></p> <p>Spatial coverage: global</p> <p>Temporal coverage: 2002 to 2020</p> <p>Spatial resolution: 0.05&times;0.05&deg;</p> <p>Temporal resolution: eight-day</p> <p>Unite: g C m<sup>&minus;2</sup> d<sup>&minus;1</sup></p> <p>Data format: raster (.tif)</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Exploring the Multidimensional Impact of Organizational Incentives on the Loyalty of Returned Overseas Faculty: An SEM and Random Forest Analysis

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record