Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “Random Forest”

Learn how ShareScore rates datasets ↗
nasa28/100

Lake Bathymetry Maps derived from Landsat and Random Forest Modeling, North Slope, AK

This dataset provides lake bathymetry maps derived from Landsat surface reflectance products for a portion of the North Slope area of Alaska. A random forest regression algorithm was used to generate depths for each point identified as being part of a lake, creating depth prediction files for each Landsat scene available for the study period: 2016-07-01 to 2018-08-31. These products are fitted to the ABoVE standard projection and reference grid to make them easily scalable and geometrically compatible with other products in the ABoVE study domain. The data are provided in cloud-optimized GeoTIFF (COG) format.

restrictednotspecifiedApr 2025View details →
geo24/100

tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction

GEO Series GSE189510. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo24/100

A Random-Forest Based Algorithm for Prediction of Enhancers From Histone Modifications

GEO Series GSE37858. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMay 2012View details →
geo24/100

Random forest-based modelling to detect novel biomarkers for prostate cancer progression

GEO Series GSE127985. Homo sapiens. 70 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenOct 2019View details →
zenodo24/100

Data for: "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study"

<p>Data and software&nbsp;related to the manuscript &quot;Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study&quot; (2020).&nbsp;</p> <p>========</p> <p>Typo in README_DataRep.txt:</p> <p>&#39; 2) &quot;one_config&quot; [...] Selected results requiring these data are shown in Fig. <strong>13</strong>&#39;<strong>&nbsp;</strong>(not 11).</p>

opencc-by-4.0Sep 2020View details →
zenodo24/100

Random Forest Cloud Model for Predicting Liquid Cloud Microphysical Properties from A-Train Data

<p>Code for creating and analyzing the performance of a random forest model to predict cloud optical depth and cloud top effective radius from A-train satellite observations. Because of storage limitations, this directory does not include full satellite dataset, but the CloudSat data is available from the CloudSat Data Processing center (https://www.cloudsat.cira.colostate.edu/) and the CALIPSO data from NASA's Atmospheric Science Data Center (https://asdc.larc.nasa.gov/project/CALIPSO).&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo24/100

An adaptable Random Forest model for the declustering of earthquake catalogs

<p>Random Forest models trained on different proportion of the synthetic earthquake catalog&nbsp;data set.</p> <p>The name of the model is defined as RF_XP_Y_Z.sav where:</p> <p>&nbsp; &nbsp;- X is the number of neighbor considered in the training stage;</p> <p>&nbsp; &nbsp;- Y is the number of feature considered in the training stage;</p> <p>&nbsp; &nbsp;- Z is the percentage (0-100) of the total data set used for the training stage.&nbsp;</p> <p>More details on how to use the model in <a href="https://github.com/florentaden/mldeclustering">this Github&nbsp;page</a>.</p>

opencc-by-4.0Oct 2021View details →
zenodo24/100

Toward emulating an explicit organic chemistry mechanism with a random forest model: dataset and training code

<p>This repository contains the dataset created with the GECKO-A model and the code (training_gecko_rf_final.py) used to train and test random forests for predicting secondary organic aerosol formation.</p> <p>For each simulation, results are distributed in two separate files identified as such:</p> <ul> <li>&lt;precursor&gt;_library_&lt;id&gt;_predictors.csv and &lt;precursor&gt;_library_&lt;id&gt;_outcomes.csv.</li> <li>&lt;precursor&gt; is either ARO1 (toluene) or dodecane_4gen (dodecane).</li> <li>&lt;id&gt; is a unique simulation identifier.</li> <li>the *predictors.csv files contain the state of the predictors for each timestep at the beginning of the chemical solver integration step.</li> <li>the *outcomes.csv files contain the state of the outcomes at the end of the chemical solver integration step.</li> </ul> <p>The TRAINING_* directories contain training simulations. TRAINING_ALL contains all the training data, used for the default random forest configuration. TRAINING_*NOX contain sorted training data matching LOW, MID and HIGH NOx initial regimes (see associated article) to train the specialized random forests.</p> <p>Similarly, the VALIDATION_* directories contain validation simulations, used to test the random forests after training.</p> <p>The TESTING* directories contain the results of testing the random forest for comparison with the VALIDATION simulations.</p>

opencc-by-4.0Nov 2022View details →
zenodo24/100

Results of cross-validation by Mlflow for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"

<p>This dataset of results contains a summary of the evaluation metrics for each fold, including information about hyperparameter tuning.</p>

restrictedcc-by-4.0May 2023View details →
zenodo24/100

Datasets for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"

<p>This dataset contains evaluations of understandability incorporating audience characteristics and model characteristics.<br> The dataset has been collected in two conducted experiments. Audience&#39; data was collected using students computing science and business and management students.</p> <p>The main_dataset.csv file contains the dataset&nbsp;to build the predictive model.</p> <p>The final_validation_dataset.csv file contains the dataset&nbsp;to evaluate the performance of the predictive model.</p>

restrictedcc-by-4.0May 2023View details →
ClinicalTrials.gov24/100

Effects of Forest Therapy on Physical and Psychological Parameters in the General Population - a Randomized Controlled Trial

ClinicalTrials.gov study NCT05562128. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad24/100

Data from: Identification of any structure-specific hepatotoxic potential of different pyrrolizidine alkaloids using Random Forest and artificial Neural Network

Open the record for dataset details and reuse information.

publicSep 2017View details →
geo24/100

miRWoods: enhanced precursor detection and stacked random forests for the sensitive detection of microRNAs.

GEO Series GSE125279. Bos taurus; Felis catus. 14 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenAug 2019View details →
zenodo20/100

Random Forest models and maps of heavy metal and nitrogen concentrations in moss in 2010 across Europe, link to research data and scientific software

<p>Research data and scientific software related to a study exploring the statistical relations between the concentration of nine heavy metals (As, Cd, Cr, Cu, Hg, Ni, Pb, V, Zn) and N in moss specimens collected in 2010 throughout Europe and a set potential explanatory variables (such as the atmospheric deposition calculated by use of two chemical transport models, distance from emission sources, density of different land uses, population density, elevation, precipitation, clay content of soils). Statistical analysis and modelling relies on Random Forest (RF). RF-models in conjunction with a Geographical Information System (GIS) were then used for mapping spatial patterns of element concentrations in moss across Europe.</p>

restrictedMar 2017View details →
nasa20/100

UNDERSTANDING SEVERE WEATHER PROCESSES THROUGH SPATIOTEMPORAL RELATIONAL RANDOM FORESTS

UNDERSTANDING SEVERE WEATHER PROCESSES THROUGH SPATIOTEMPORAL RELATIONAL RANDOM FORESTS AMY MCGOVERN, TIMOTHY SUPINIE, DAVID JOHN GAGNE II, NATHANIEL TROUTMAN, MATTHEW COLLIER, RODGER A. BROWN, JEFFREY BASARA, AND JOHN K. WILLIAMS Abstract. Major severe weather events can cause a significant loss of life and property. We seek to revolutionize our understanding of and ability to predict such events through the mining of severe weather data. Because weather is inherently a spatiotemporal phenomenon, mining such data requires a model capable of representing and reasoning about complex spatiotemporal dynamics, including temporally and spatially varying attributes and relationships. We introduce an augmented version of the Spatiotemporal Relational Random Forest, which is a Random Forest that learns with spatiotemporally varying relational data. Our algorithm maintains the strength and performance of Random Forests but extends their applicability, including the estimation of variable importance, to complex spatiotemporal relational domains. We apply the augmented Spatiotemporal Relational Random Forest to three severe weather data sets. These are: predicting atmospheric turbulence across the continental United States, examining the formation of tornadoes near strong frontal boundaries, and understanding the translation of drought across the southern plains of the United States. The results on such a wide variety of real-world domains demonstrate the extensive applicability of the Spatiotemporal Relational Random Forest. Our long-term goal is to significantly improve the ability to predict and warn about severe weather events.

restrictednotspecifiedApr 2025View details →
zenodo16/100

Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multisource data (2001-2002)

<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:<span>T<sub>ave</sub>, </span><span>R<sup>2</sup> = 0.97, RMSE = 1.61℃ and rRMSE = 13.24%</span><span>; T<sub>max</sub>, </span><span>R<sup>2</sup> = 0.94, RMSE = 2.35℃ and rRMSE = 13.02%</span><span>; T<sub>min</sub>, </span><span>R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%</span><span>).</span></p>

embargoedcc-by-4.0Mar 2024View details →
zenodo16/100

Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2013-2014)

<div> <p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>,&nbsp;R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p> </div>

embargoedcc-by-4.0Apr 2024View details →
zenodo16/100

Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2017-2018)

<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>,&nbsp;R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>

embargoedcc-by-4.0Apr 2024View details →
zenodo16/100

Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2011-2012)

<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>,&nbsp;R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>

embargoedcc-by-4.0Apr 2024View details →
zenodo16/100

Near Surface Air Temperature Dataset for China with high temporal and spatial resolution generated using random forest and multi-source data (2009-2010)

<p>The dataset presents the daily near surface air temperature of China with 1km spatial resolution, including daily average air temperature, maximum temperature and minimum air temperature. The dataset was generated using machine learning and multiple variables, the accuracy was:T<sub>ave</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.97, RMSE = 1.61℃ and rRMSE = 13.24%; T<sub>max</sub>,&nbsp;R<sup>2</sup>&nbsp;= 0.94, RMSE = 2.35℃ and rRMSE = 13.02%; T<sub>min</sub>,&nbsp;R<sup>2</sup> = 0.95, RMSE = 2.04℃ and rRMSE = 27.09%).</p>

embargoedcc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record