Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forest”
Random Forest fused MODIS and Landsat snow cover from spectral mixture analysis in the Sierra Nevada, USA
<p>This data is snow cover fraction from the Snow Covered Area and Grain Size (SCAG) model for Landsat OLI and Terra MODIS and well as a 2-stage random forest model to fuse the 2 datasets for improved temporal/spatial resolution. There are 170 scenes in 2001 to 2012. It was used in the a publication for Remote Sensing of the Environment titled: Multi-sensor fusion using random forests for daily fractional snow cover at 30 m, doi: to be assigned.</p> <p><strong>Inputs</strong>: [YYYYMMDD is year month day of month, $num is 5 or 7 for Landsat platform, $sens is sensor TM or ETM+]</p> <p>Landsat.zip:</p> <p>Snow cover from Landsat: SSN.p042r034_YYYYMMDD.Landsat$num-$sens.canopyadjusted_mask.v01.tif </p> <p> </p> <p>MODIS.zip:</p> <p>Snow cover from MODIS: SSN.SN_W$YYYYMMDD_$YYYYMMDD.Terra-MODIS.snow_cover_percent.v01.tif</p> <p> </p> <p>Predictors.zip<strong> </strong></p> <p>Static predictors (see RSE publication Table 2): SouthernSierraNevada*.tif [* here is the variable name]</p> <p><strong>Outputs [</strong> [YYYYMMDD is year month day of month]</p> <p>ProbabilityNot0Not100.zip</p> <p>SSN.prob.btwn.YYYYMMDD.v3.tif - from classification random forest, probability of being between 0 and 100</p> <p> </p> <p>Probability100fSCA</p> <p>SSN.pro.hundred.YYYYMMDD.v3.tif - from classification random forest, probability of being 100</p> <p> </p> <p>RegressionResult.zip</p> <p>SSN.regression.YYYYMMDD.v3.tif - from prediction random forest</p> <p> </p> <p>Final_Downscaled.zip</p> <p>SSN.downscaled.YYYYMMDD.v3.3e+05.tif - final product (combination of classification and prediction)</p>
Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm
<p><strong>Aims</strong>: Remote sensing approaches could be beneficial for monitoring and compiling essential biodiversity data because it is cost-effective and allows for coverage of large areas over a short period. This study investigated the relationship between multispectral remote sensing data from Landsat 8 and Sentinel 2 and species richness and diversity in mountainous and protected grasslands.</p> <p><strong>Locations</strong>: Golden Gate Highlands National Park, Free State, South Africa. </p> <p><strong>Methods</strong>: In-situ data of plant species composition and cover from 142 plots with 16 releves each were distributed across the study site and used to calculate species richness and Shannon-wiener species diversity index (species diversity. We used a machine-learning random forest algorithm to optimise the prediction of species richness and diversity. The algorithm was used to identify the optimal spectral bands and vegetation indices for estimating species richness and diversity. Subsequently, the selected bands and vegetation indices were used to estimate species richness through random forest regression. </p> <p><strong>Results</strong>: This research found weak relationships between remote sensing vegetation indices and the diversity metrics, but significant relationships were found between some spectral bands and diversity metrics. Moreover, using machine learning random forest, the multispectral datasets exhibited strong predictive powers. In this investigation, for both sensors, near-infrared (NIR) seemed to be the most selected band to explain species diversity in mountainous grasslands.</p> <p><strong>Main</strong> <strong>conclusions</strong>: This finding further ascertains the efficiency of using NIR in vegetation mapping. This research shows that NIR, SAVI and EVI are the most adequate for predicting species richness and diversity in mountainous grasslands with relatively good accuracies.</p>
Output - Results of Random Forest and Multiple Linear regression analysis.
<p><strong>Hybrid streamflow modelling using machine learning and multi-model combination.</strong></p> <p> </p> <p><strong>Structure:</strong></p> <p><strong>MLR_output:</strong></p> <ul> <li>Validate <ul> <li>Different setups</li> </ul> </li> </ul> <p><strong>RF_output:</strong></p> <ul> <li>tune <ul> <li><em>all_stations</em></li> </ul> </li> <li>train <ul> <li><em>Different setups</em></li> </ul> </li> <li>Validate <ul> <li><em>Different setups</em></li> </ul> </li> </ul>
Presence of perennial halophilous scrubs in 2016 in the former saltworks of Salins de Giraud, random forest classification on WV2
<p>The presence of perennial halophilous scrubs (<em>S. fruticosa</em> and <em>A. macrostachyum</em>) was mapped using a Worldview 2 (WV2) very-high resolution imagery acquired the 31st of August 2016, strictly cloudless. Multispectral indices, among the most used, were adapted to the bands of WV2. The Random Forest model (RF) package into the R software was used to perform image classification. RF OBB and omission errors for the class “presence” and the class “absence” were less important when the classification was done for each species individually, hence this raster is a combination of both classifications. The OBB error for <em>S.fruticosa</em> is 1.43% and for <em>A. macrostachyum</em> is 2.18%.</p>
Effect of hyperparameters on variable selection in random forests
<p>These are data used to generate the results presented in simulation studies conducted in Fouodo et al. 2023. Each dataset is an R object in RDS format with 100 lists. For each element of the list, parameters used to generate the data are presented, followed by the simulated data. <a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection">Here</a> is the git repository of the R code used to conduct simulations in the manuscript. The files <a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection/tree/main/R-code/01-scenario1">01-data-only.R</a> and <a href="https://github.com/imbs-hl/RF-hyperparameters-and-variable-selection/tree/main/R-code/02-scenario2">02-data-only.R</a> contain the functions and more details about how the data have been simulated for studies 1 and 2.</p>
A Cluster Randomized Trial to Evaluate Long Lasting Insecticidal Hammocks to Prevent Forest Malaria in Vietnam
ClinicalTrials.gov study NCT00853281. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Progress toward forecasting excessive rainfall with random forests based on a deterministic convection-allowing model
Open the record for dataset details and reuse information.
Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm
Open the record for dataset details and reuse information.
Data from: Novel, continuous monitoring of fine-scale movement using fixed-position radiotelemetry arrays and random forest location fingerprinting
Open the record for dataset details and reuse information.
Data from: Integration of Random Forest with population-based outlier analyses provides insight on the genomic basis and evolution of run timing in Chinook salmon (Oncorhynchus tshawytscha)
Open the record for dataset details and reuse information.
Data from: Applications of random forest feature selection for fine-scale genetic population assignment
Open the record for dataset details and reuse information.
Dataset of: Using random forest for brain tissue identification by Raman spectroscopy
Open the record for dataset details and reuse information.
Data from: A practical introduction to random forest for genetic association studies in ecology and evolution
Large genomic studies are becoming increasingly common with advances in sequencing technology, and our ability to understand how genomic variation influences phenotypic variation between individuals has never been greater. The exploration of such relationships first requires the identification of associations between molecular markers and phenotypes. Here we explore the use of Random Forest (RF), a powerful machine learning algorithm, in genomic studies to discern loci underlying both discrete and quantitative traits, particularly when studying wild or non-model organisms. RF is becoming increasingly used in ecological and population genetics because, unlike traditional methods, it can efficiently analyze thousands of loci simultaneously and account for non-additive interactions. However, understanding both the power and limitations of Random Forest is important for its proper implementation and the interpretation of results. We therefore provide a practical introduction to the algorithm and its use for identifying associations between molecular markers and phenotypes, discussing such topics as data limitations, algorithm initiation and optimization, as well as interpretation. We also provide short R tutorials as examples, with the aim of providing a guide to the implementation of the algorithm. Topics discussed here are intended to serve as an entry point for molecular ecologists interested in employing Random Forest to identify trait associations in genomic data sets.
High-resolution snow depth prediction using Random Forest algorithm with topographic parameters
Open the record for dataset details and reuse information.
Figure 8 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 8. Relationship of diagnosability to Nei's net nucleotide divergence (dA, Nei and Kumar 2000). Colors indicate taxonomic level of comparison: species (blue), subspecies (green), and populations (red). Vertical lines indicate the binomial 95% CI around diagnosability. Note that x-axis for dA is log10-scaled.
Files associated with Rui Jin, Anand Gnanadesikan and Christopher Holder, Using Random Forests to Compare the Sensitivity of Observed Particulate Inorganic and Particulate Organic Carbon to Environmental Conditions
Open the record for dataset details and reuse information.
A global gross primary productivity dataset of sunlit and shaded leaves via combining two-leaf light use efficiency model with random forest from 2002 to 2020
<p><span>The TL-CRF model generated a global </span><span>0.05</span><span><span>´</span></span><span>0.05<span>°</span></span><span> product for eight-day gross primary productivity (GPP) of sunlit and shaded canopies from 2002 to 2020 by embedding the random forest (RF) submodule into the two-leaf light use efficiency (TL-LUE) model while considering the seasonal differences in the clumping index. The RF technique was used to integrate various environmental stress factors including meteorological, hydrological, soil properties, and elevation, thereby improving the overall scale of the complex environmental conditions to the maximum LUE. Eight-day GPP was then aggregated into monthly, seasonal, and annual GPP. This novel GPP product could support further research on spatial and temporal patterns of the carbon cycle and its association with climate change. </span></p>
The data and code for "Gap-Filling of Turbulent Heat Fluxes over Rice–Wheat-Rotation Croplands Using the Random Forest Model""
<p>This file contains the dataset and code for the paper "Gap-Filling of Turbulent Heat Fluxes over Rice–Wheat-Rotation Croplands Using the Random Forest Model".</p>
Random Similarity Isolation Forest - precalculated distances matrices
<p>Precalculated distance matrices for Random Similarity Isolation Forest algorithm.</p>
Data from: A practical introduction to random forest for genetic association studies in ecology and evolution
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.