Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forest”
Lake Bathymetry Maps derived from Landsat and Random Forest Modeling, North Slope, AK
This dataset provides lake bathymetry maps derived from Landsat surface reflectance products for a portion of the North Slope area of Alaska. A random forest regression algorithm was used to generate depths for each point identified as being part of a lake, creating depth prediction files for each Landsat scene available for the study period: 2016-07-01 to 2018-08-31. These products are fitted to the ABoVE standard projection and reference grid to make them easily scalable and geometrically compatible with other products in the ABoVE study domain. The data are provided in cloud-optimized GeoTIFF (COG) format.
tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction
GEO Series GSE189510. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
A Random-Forest Based Algorithm for Prediction of Enhancers From Histone Modifications
GEO Series GSE37858. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Random forest-based modelling to detect novel biomarkers for prostate cancer progression
GEO Series GSE127985. Homo sapiens. 70 samples. Type: Methylation profiling by genome tiling array.
Data for: "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study"
<p>Data and software related to the manuscript "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study" (2020). </p> <p>========</p> <p>Typo in README_DataRep.txt:</p> <p>' 2) "one_config" [...] Selected results requiring these data are shown in Fig. <strong>13</strong>'<strong> </strong>(not 11).</p>
Random Forest Cloud Model for Predicting Liquid Cloud Microphysical Properties from A-Train Data
<p>Code for creating and analyzing the performance of a random forest model to predict cloud optical depth and cloud top effective radius from A-train satellite observations. Because of storage limitations, this directory does not include full satellite dataset, but the CloudSat data is available from the CloudSat Data Processing center (https://www.cloudsat.cira.colostate.edu/) and the CALIPSO data from NASA's Atmospheric Science Data Center (https://asdc.larc.nasa.gov/project/CALIPSO). </p>
An adaptable Random Forest model for the declustering of earthquake catalogs
<p>Random Forest models trained on different proportion of the synthetic earthquake catalog data set.</p> <p>The name of the model is defined as RF_XP_Y_Z.sav where:</p> <p> - X is the number of neighbor considered in the training stage;</p> <p> - Y is the number of feature considered in the training stage;</p> <p> - Z is the percentage (0-100) of the total data set used for the training stage. </p> <p>More details on how to use the model in <a href="https://github.com/florentaden/mldeclustering">this Github page</a>.</p>
Toward emulating an explicit organic chemistry mechanism with a random forest model: dataset and training code
<p>This repository contains the dataset created with the GECKO-A model and the code (training_gecko_rf_final.py) used to train and test random forests for predicting secondary organic aerosol formation.</p> <p>For each simulation, results are distributed in two separate files identified as such:</p> <ul> <li><precursor>_library_<id>_predictors.csv and <precursor>_library_<id>_outcomes.csv.</li> <li><precursor> is either ARO1 (toluene) or dodecane_4gen (dodecane).</li> <li><id> is a unique simulation identifier.</li> <li>the *predictors.csv files contain the state of the predictors for each timestep at the beginning of the chemical solver integration step.</li> <li>the *outcomes.csv files contain the state of the outcomes at the end of the chemical solver integration step.</li> </ul> <p>The TRAINING_* directories contain training simulations. TRAINING_ALL contains all the training data, used for the default random forest configuration. TRAINING_*NOX contain sorted training data matching LOW, MID and HIGH NOx initial regimes (see associated article) to train the specialized random forests.</p> <p>Similarly, the VALIDATION_* directories contain validation simulations, used to test the random forests after training.</p> <p>The TESTING* directories contain the results of testing the random forest for comparison with the VALIDATION simulations.</p>
Results of cross-validation by Mlflow for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"
<p>This dataset of results contains a summary of the evaluation metrics for each fold, including information about hyperparameter tuning.</p>
Datasets for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"
<p>This dataset contains evaluations of understandability incorporating audience characteristics and model characteristics.<br> The dataset has been collected in two conducted experiments. Audience' data was collected using students computing science and business and management students.</p> <p>The main_dataset.csv file contains the dataset to build the predictive model.</p> <p>The final_validation_dataset.csv file contains the dataset to evaluate the performance of the predictive model.</p>
Effects of Forest Therapy on Physical and Psychological Parameters in the General Population - a Randomized Controlled Trial
ClinicalTrials.gov study NCT05562128. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: Identification of any structure-specific hepatotoxic potential of different pyrrolizidine alkaloids using Random Forest and artificial Neural Network
Open the record for dataset details and reuse information.
miRWoods: enhanced precursor detection and stacked random forests for the sensitive detection of microRNAs.
GEO Series GSE125279. Bos taurus; Felis catus. 14 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Random Forest models and maps of heavy metal and nitrogen concentrations in moss in 2010 across Europe, link to research data and scientific software
<p>Research data and scientific software related to a study exploring the statistical relations between the concentration of nine heavy metals (As, Cd, Cr, Cu, Hg, Ni, Pb, V, Zn) and N in moss specimens collected in 2010 throughout Europe and a set potential explanatory variables (such as the atmospheric deposition calculated by use of two chemical transport models, distance from emission sources, density of different land uses, population density, elevation, precipitation, clay content of soils). Statistical analysis and modelling relies on Random Forest (RF). RF-models in conjunction with a Geographical Information System (GIS) were then used for mapping spatial patterns of element concentrations in moss across Europe.</p>
UNDERSTANDING SEVERE WEATHER PROCESSES THROUGH SPATIOTEMPORAL RELATIONAL RANDOM FORESTS
UNDERSTANDING SEVERE WEATHER PROCESSES THROUGH SPATIOTEMPORAL RELATIONAL RANDOM FORESTS AMY MCGOVERN, TIMOTHY SUPINIE, DAVID JOHN GAGNE II, NATHANIEL TROUTMAN, MATTHEW COLLIER, RODGER A. BROWN, JEFFREY BASARA, AND JOHN K. WILLIAMS Abstract. Major severe weather events can cause a significant loss of life and property. We seek to revolutionize our understanding of and ability to predict such events through the mining of severe weather data. Because weather is inherently a spatiotemporal phenomenon, mining such data requires a model capable of representing and reasoning about complex spatiotemporal dynamics, including temporally and spatially varying attributes and relationships. We introduce an augmented version of the Spatiotemporal Relational Random Forest, which is a Random Forest that learns with spatiotemporally varying relational data. Our algorithm maintains the strength and performance of Random Forests but extends their applicability, including the estimation of variable importance, to complex spatiotemporal relational domains. We apply the augmented Spatiotemporal Relational Random Forest to three severe weather data sets. These are: predicting atmospheric turbulence across the continental United States, examining the formation of tornadoes near strong frontal boundaries, and understanding the translation of drought across the southern plains of the United States. The results on such a wide variety of real-world domains demonstrate the extensive applicability of the Spatiotemporal Relational Random Forest. Our long-term goal is to significantly improve the ability to predict and warn about severe weather events.
Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multisource data (2001-2002)
<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:<span>T<sub>ave</sub>, </span><span>R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%</span><span>; T<sub>max</sub>, </span><span>R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%</span><span>; T<sub>min</sub>, </span><span>R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%</span><span>).</span></p>
Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2013-2014)
<div> <p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>, R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>, R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>, R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p> </div>
Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2017-2018)
<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>, R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>, R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>, R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>
Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2011-2012)
<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>, R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>, R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>, R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>
Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2009-2010)
<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>, R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>, R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>, R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.