Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
217
datasets available to search
ShareScore release 0.9.0
Dataset results
217 results for “forest model”
Lake Bathymetry Maps derived from Landsat and Random Forest Modeling, North Slope, AK
This dataset provides lake bathymetry maps derived from Landsat surface reflectance products for a portion of the North Slope area of Alaska. A random forest regression algorithm was used to generate depths for each point identified as being part of a lake, creating depth prediction files for each Landsat scene available for the study period: 2016-07-01 to 2018-08-31. These products are fitted to the ABoVE standard projection and reference grid to make them easily scalable and geometrically compatible with other products in the ABoVE study domain. The data are provided in cloud-optimized GeoTIFF (COG) format.
Random forest-based modelling to detect novel biomarkers for prostate cancer progression
GEO Series GSE127985. Homo sapiens. 70 samples. Type: Methylation profiling by genome tiling array.
Dastaset for: "Rodriguez-Galiano, V.F., Sanchez-Castillo, M., Dash, J., Atkinson, P. and Ojeda-Zujar, J. (2016). Modelling interannual variation in the spring and autumn land surface phenology of the European forest, Biogeosciences, 13
<p>Dastaset for: "Rodriguez-Galiano, V.F., Sanchez-Castillo, M., Dash, J., Atkinson, P. and Ojeda-Zujar, J. (2016). Modelling interannual variation in the spring and autumn land surface phenology of the European forest, Biogeosciences, 13</p>
Random Forest Cloud Model for Predicting Liquid Cloud Microphysical Properties from A-Train Data
<p>Code for creating and analyzing the performance of a random forest model to predict cloud optical depth and cloud top effective radius from A-train satellite observations. Because of storage limitations, this directory does not include full satellite dataset, but the CloudSat data is available from the CloudSat Data Processing center (https://www.cloudsat.cira.colostate.edu/) and the CALIPSO data from NASA's Atmospheric Science Data Center (https://asdc.larc.nasa.gov/project/CALIPSO). </p>
An adaptable Random Forest model for the declustering of earthquake catalogs
<p>Random Forest models trained on different proportion of the synthetic earthquake catalog data set.</p> <p>The name of the model is defined as RF_XP_Y_Z.sav where:</p> <p> - X is the number of neighbor considered in the training stage;</p> <p> - Y is the number of feature considered in the training stage;</p> <p> - Z is the percentage (0-100) of the total data set used for the training stage. </p> <p>More details on how to use the model in <a href="https://github.com/florentaden/mldeclustering">this Github page</a>.</p>
Toward emulating an explicit organic chemistry mechanism with a random forest model: dataset and training code
<p>This repository contains the dataset created with the GECKO-A model and the code (training_gecko_rf_final.py) used to train and test random forests for predicting secondary organic aerosol formation.</p> <p>For each simulation, results are distributed in two separate files identified as such:</p> <ul> <li><precursor>_library_<id>_predictors.csv and <precursor>_library_<id>_outcomes.csv.</li> <li><precursor> is either ARO1 (toluene) or dodecane_4gen (dodecane).</li> <li><id> is a unique simulation identifier.</li> <li>the *predictors.csv files contain the state of the predictors for each timestep at the beginning of the chemical solver integration step.</li> <li>the *outcomes.csv files contain the state of the outcomes at the end of the chemical solver integration step.</li> </ul> <p>The TRAINING_* directories contain training simulations. TRAINING_ALL contains all the training data, used for the default random forest configuration. TRAINING_*NOX contain sorted training data matching LOW, MID and HIGH NOx initial regimes (see associated article) to train the specialized random forests.</p> <p>Similarly, the VALIDATION_* directories contain validation simulations, used to test the random forests after training.</p> <p>The TESTING* directories contain the results of testing the random forest for comparison with the VALIDATION simulations.</p>
Results of cross-validation by Mlflow for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"
<p>This dataset of results contains a summary of the evaluation metrics for each fold, including information about hyperparameter tuning.</p>
Datasets for "Can the audience understand a business process model? Using Random Forest to classify level understandability based on personal and model factors"
<p>This dataset contains evaluations of understandability incorporating audience characteristics and model characteristics.<br> The dataset has been collected in two conducted experiments. Audience' data was collected using students computing science and business and management students.</p> <p>The main_dataset.csv file contains the dataset to build the predictive model.</p> <p>The final_validation_dataset.csv file contains the dataset to evaluate the performance of the predictive model.</p>
Forest-BGC Model (OTTER)
Steve Running's Forest-BGC Model
Random Forest models and maps of heavy metal and nitrogen concentrations in moss in 2010 across Europe, link to research data and scientific software
<p>Research data and scientific software related to a study exploring the statistical relations between the concentration of nine heavy metals (As, Cd, Cr, Cu, Hg, Ni, Pb, V, Zn) and N in moss specimens collected in 2010 throughout Europe and a set potential explanatory variables (such as the atmospheric deposition calculated by use of two chemical transport models, distance from emission sources, density of different land uses, population density, elevation, precipitation, clay content of soils). Statistical analysis and modelling relies on Random Forest (RF). RF-models in conjunction with a Geographical Information System (GIS) were then used for mapping spatial patterns of element concentrations in moss across Europe.</p>
Near-term modeled forest fire area and fuel aridity projections
Open the record for dataset details and reuse information.
A pigtailed macaque model for Kyasanur Forest disease virus and Alkhurma hemorrhagic disease virus pathogenesis
GEO Series GSE185797. Macaca nemestrina. 30 samples. Type: Expression profiling by high throughput sequencing.
Trend analysis and random forests models assessing spatial and temporal patterns of wildfire probability for the eastern United States
<p>We used historic fire perimeters from Monitoring Trends in Burn Severity to assess trends and drivers of wildfires in the eastern United States. We used a suite of predictor variables relating to weather, vegetation cover, and human infrastructure to parameterize random forests models predicting fire occurrence. Models were used to project annual burned areas using all selected predictors, and to project the marginal response of annual burned areas to the most important weather predictors. This dataset includes Python scripts, raster maps of fire probability, and tables summarizing analysis results. </p>
Dataset for Water Resource / Forest Restoration Modeling in Tahoe-Central Sierra Region
<p>This dataset includes materials related to a water resource modeling project in the Tahoe-Central Sierra Initiative (TCSI) region. The project used the DHSVM and LANDIS-II models to predict the potential hydrological impacts of various forest management scenarios related to partial or full restoration of the historic forest disturbance return interval.</p> <p>This dataset includes:</p> <ol> <li>Folders for each figure in the submitted manuscript, including data and scripts needed to reproduce the figures.</li> <li>Model results, including CSV files of raw streamflow outputs from DHSVM, a combined CSV file of mass/energy balance outputs from DHSVM, and geoTIFF rasters of various snowpack states from DHSVM. Additionally, post-processed data are included, such as aggregated streamflow products and comparisons between management scenarios.</li> <li>Ancillary data related to the model setup, calibration, and scenario implementation, including meteorology, configuration files, watershed boundaries, output vegetation maps from LANDIS-II used as inputs to DHSVM, and calibration reference data.</li> </ol>
Forest cover and biomass model estimations for Maasai Mau region, Kenya - Beta version
<p>We are releasing preliminary forest cover and biomass model outputs from IBM’s Prithvi Earth Observation Foundation Model. The outputs to be released are snippets with limited spatial and temporal coverage and generated from a preliminary, non-optimized version of IBM’s Prithvi Earth Observation Foundation Model. The outputs cover Maasai Mau forest region in Kenya across several yearly timestamps. The outputs are provided in geo-referenced file format (geotiff) under a CDLA license.</p>
Thesis: Locate and understand the mortality risk of forest tree under climate change. A multi-scale approach combining statistical and mechanistic modeling
<p>Data and tables in chapters 2 and 3 of the thesis</p>
Individual risk prediction: comparing Random Forests with Cox proportional-hazards model by a simulation study
<p>Results provided for reproducibility revision</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.