Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “Random Forest”

Learn how ShareScore rates datasets ↗
zenodo32/100

Random Forest fused MODIS and Landsat snow cover from spectral mixture analysis in the Sierra Nevada, USA

<p>This data is snow cover fraction from the Snow Covered Area and Grain Size (SCAG) model for Landsat OLI and Terra MODIS and well as a 2-stage random forest model to fuse the 2 datasets for improved temporal/spatial resolution. There are 170 scenes in 2001 to 2012.&nbsp;It was used in the a publication for Remote Sensing of the Environment titled: Multi-sensor fusion using random forests for daily fractional snow cover at 30&nbsp;m,&nbsp;doi: to be assigned.</p> <p><strong>Inputs</strong>:&nbsp;[YYYYMMDD is year month day of month, $num is 5 or 7 for Landsat platform, $sens is sensor TM or ETM+]</p> <p>Landsat.zip:</p> <p>Snow cover from Landsat: SSN.p042r034_YYYYMMDD.Landsat$num-$sens.canopyadjusted_mask.v01.tif&nbsp;</p> <p>&nbsp;</p> <p>MODIS.zip:</p> <p>Snow cover from MODIS: SSN.SN_W$YYYYMMDD_$YYYYMMDD.Terra-MODIS.snow_cover_percent.v01.tif</p> <p>&nbsp;</p> <p>Predictors.zip<strong>&nbsp;</strong></p> <p>Static predictors (see RSE publication Table 2): SouthernSierraNevada*.tif [* here is the variable name]</p> <p><strong>Outputs [</strong>&nbsp;[YYYYMMDD is year month day of month]</p> <p>ProbabilityNot0Not100.zip</p> <p>SSN.prob.btwn.YYYYMMDD.v3.tif - from classification random forest, probability of being between 0 and 100</p> <p>&nbsp;</p> <p>Probability100fSCA</p> <p>SSN.pro.hundred.YYYYMMDD.v3.tif - from classification random forest, probability of being 100</p> <p>&nbsp;</p> <p>RegressionResult.zip</p> <p>SSN.regression.YYYYMMDD.v3.tif - from prediction random forest</p> <p>&nbsp;</p> <p>Final_Downscaled.zip</p> <p>SSN.downscaled.YYYYMMDD.v3.3e+05.tif - final product (combination of classification and prediction)</p>

opencc-by-4.0Jul 2021View details →
dryad32/100

Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm

<p><strong>Aims</strong>: Remote sensing approaches could be beneficial for monitoring and compiling essential biodiversity data because it is cost-effective and allows for coverage of large areas over a short period. This study investigated the relationship between multispectral remote sensing data from Landsat 8 and Sentinel 2 and species richness and diversity in mountainous and protected grasslands.</p> <p><strong>Locations</strong>: Golden Gate Highlands National Park, Free State, South Africa. </p> <p><strong>Methods</strong>: In-situ data of plant species composition and cover from 142 plots with 16 releves each were distributed across the study site and used to calculate species richness and Shannon-wiener species diversity index (species diversity. We used a machine-learning random forest algorithm to optimise the prediction of species richness and diversity. The algorithm was used to identify the optimal spectral bands and vegetation indices for estimating species richness and diversity. Subsequently, the selected bands and vegetation indices were used to estimate species richness through random forest regression. </p> <p><strong>Results</strong>: This research found weak relationships between remote sensing vegetation indices and the diversity metrics, but significant relationships were found between some spectral bands and diversity metrics. Moreover, using machine learning random forest, the multispectral datasets exhibited strong predictive powers. In this investigation, for both sensors, near-infrared (NIR) seemed to be the most selected band to explain species diversity in mountainous grasslands.</p> <p><strong>Main</strong> <strong>conclusions</strong>: This finding further ascertains the efficiency of using NIR in vegetation mapping.  This research shows that NIR, SAVI and EVI are the most adequate for predicting species richness and diversity in mountainous grasslands with relatively good accuracies.</p>

opencc-zeroMay 2023View details →
zenodo32/100

Output - Results of Random Forest and Multiple Linear regression analysis.

<p><strong>Hybrid streamflow modelling using machine learning and multi-model combination.</strong></p> <p>&nbsp;</p> <p><strong>Structure:</strong></p> <p><strong>MLR_output:</strong></p> <ul> <li>Validate <ul> <li>Different setups</li> </ul> </li> </ul> <p><strong>RF_output:</strong></p> <ul> <li>tune <ul> <li><em>all_stations</em></li> </ul> </li> <li>train <ul> <li><em>Different setups</em></li> </ul> </li> <li>Validate <ul> <li><em>Different setups</em></li> </ul> </li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Presence of perennial halophilous scrubs in 2016 in the former saltworks of Salins de Giraud, random forest classification on WV2

<p>The presence of perennial halophilous scrubs (<em>S. fruticosa</em>&nbsp; and <em>A. macrostachyum</em>) was mapped using a Worldview 2 (WV2) very-high resolution imagery acquired the 31st of August 2016, strictly cloudless. Multispectral indices, among the most used, were adapted to the bands of WV2. The Random Forest model (RF) package into the R software was used to perform image classification. RF OBB and omission errors for the class &ldquo;presence&rdquo; and the class &ldquo;absence&rdquo; were less important when the classification was done for each species individually, hence this raster is a combination of both classifications. The OBB error for <em>S.fruticosa</em> is 1.43% and for <em>A. macrostachyum</em> is 2.18%.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Effect of hyperparameters on variable selection in random forests

<p>These are data used to generate the results presented in simulation studies conducted in&nbsp;Fouodo et al. 2023.&nbsp;Each dataset is an R object in RDS format with&nbsp;100 lists. For each element of the list,&nbsp;parameters used to generate the data are presented, followed by the simulated data.&nbsp;&nbsp;<a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection">Here</a>&nbsp;is the git repository of the R code used to conduct simulations in the&nbsp;manuscript. The files&nbsp;<a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection/tree/main/R-code/01-scenario1">01-data-only.R</a>&nbsp;and&nbsp;<a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection/tree/main/R-code/02-scenario2">02-data-only.R</a>&nbsp;contain the functions and more details about how the data have been simulated for studies&nbsp;1 and 2.</p>

opencc-by-4.0Aug 2023View details →
ClinicalTrials.gov32/100

A Cluster Randomized Trial to Evaluate Long Lasting Insecticidal Hammocks to Prevent Forest Malaria in Vietnam

ClinicalTrials.gov study NCT00853281. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Progress toward forecasting excessive rainfall with random forests based on a deterministic convection-allowing model

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad32/100

Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad32/100

Data from: Novel, continuous monitoring of fine-scale movement using fixed-position radiotelemetry arrays and random forest location fingerprinting

Open the record for dataset details and reuse information.

publicJan 2018View details →
dryad32/100

Data from: Integration of Random Forest with population-based outlier analyses provides insight on the genomic basis and evolution of run timing in Chinook salmon (Oncorhynchus tshawytscha)

Open the record for dataset details and reuse information.

publicApr 2015View details →
dryad32/100

Data from: Applications of random forest feature selection for fine-scale genetic population assignment

Open the record for dataset details and reuse information.

publicJul 2017View details →
dryad32/100

Dataset of: Using random forest for brain tissue identification by Raman spectroscopy

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad28/100

Data from: A practical introduction to random forest for genetic association studies in ecology and evolution

Large genomic studies are becoming increasingly common with advances in sequencing technology, and our ability to understand how genomic variation influences phenotypic variation between individuals has never been greater. The exploration of such relationships first requires the identification of associations between molecular markers and phenotypes. Here we explore the use of Random Forest (RF), a powerful machine learning algorithm, in genomic studies to discern loci underlying both discrete and quantitative traits, particularly when studying wild or non-model organisms. RF is becoming increasingly used in ecological and population genetics because, unlike traditional methods, it can efficiently analyze thousands of loci simultaneously and account for non-additive interactions. However, understanding both the power and limitations of Random Forest is important for its proper implementation and the interpretation of results. We therefore provide a practical introduction to the algorithm and its use for identifying associations between molecular markers and phenotypes, discussing such topics as data limitations, algorithm initiation and optimization, as well as interpretation. We also provide short R tutorials as examples, with the aim of providing a guide to the implementation of the algorithm. Topics discussed here are intended to serve as an entry point for molecular ecologists interested in employing Random Forest to identify trait associations in genomic data sets.

opencc-zeroDec 2017View details →
zenodo28/100

High-resolution snow depth prediction using Random Forest algorithm with topographic parameters

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Figure 8 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies

Figure 8. Relationship of diagnosability to Nei's net nucleotide divergence (dA, Nei and Kumar 2000). Colors indicate taxonomic level of comparison: species (blue), subspecies (green), and populations (red). Vertical lines indicate the binomial 95% CI around diagnosability. Note that x-axis for dA is log10-scaled.

opencc-by-4.0Jun 2017View details →
zenodo28/100

Files associated with Rui Jin, Anand Gnanadesikan and Christopher Holder, Using Random Forests to Compare the Sensitivity of Observed Particulate Inorganic and Particulate Organic Carbon to Environmental Conditions

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo28/100

A global gross primary productivity dataset of sunlit and shaded leaves via combining two-leaf light use efficiency model with random forest from 2002 to 2020

<p><span>The TL-CRF model generated a global </span><span>0.05</span><span><span>&acute;</span></span><span>0.05<span>&deg;</span></span><span> product for eight-day gross primary productivity (GPP) of sunlit and shaded canopies from 2002 to 2020 by embedding the random forest (RF) submodule into the two-leaf light use efficiency (TL-LUE) model while considering the seasonal differences in the clumping index. The RF technique was used to integrate various environmental stress factors including meteorological, hydrological, soil properties, and elevation, thereby improving the overall scale of the complex environmental conditions to the maximum LUE. Eight-day GPP was then aggregated into monthly, seasonal, and annual GPP. This novel GPP product could support further research on spatial and temporal patterns of the carbon cycle and its association with climate change. </span></p>

opencc-by-4.0Aug 2024View details →
zenodo28/100

The data and code for "Gap-Filling of Turbulent Heat Fluxes over Rice–Wheat-Rotation Croplands Using the Random Forest Model""

<p>This file contains the dataset and code for the paper &quot;Gap-Filling of Turbulent Heat Fluxes over Rice&ndash;Wheat-Rotation Croplands Using the Random Forest Model&quot;.</p>

opencc-by-4.0Mar 2023View details →
zenodo28/100

Random Similarity Isolation Forest - precalculated distances matrices

<p>Precalculated distance matrices for Random Similarity Isolation Forest algorithm.</p>

opencc-by-4.0Sep 2023View details →
dryad28/100

Data from: A practical introduction to random forest for genetic association studies in ecology and evolution

Open the record for dataset details and reuse information.

publicMar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record