Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forest”
Random forest modelling of multi-scale, multi-species habitat associations within KAZA transfrontier conservation area using spoor data
<p>As landscape-scale conservation models grow in prominence, assessments of how wildlife utilise multiple-use landscapes are required to inform effective conservation and management planning. Such efforts should strive to incorporate multi-species perspectives to maximise value for conservation, and should account for scale to accurately capture species-environment relationships. We show that the random forest machine learning algorithm can be used to model large-scale sign-based data in a multi-scale framework. We used this method to investigate scale-dependent habitat associations for 16 mammal species of high conservation importance across the southern Kavango Zambezi (KAZA) Transfrontier Conservation Area in Botswana and Zimbabwe. Our findings revealed substantial variation in the factors shaping habitat use across species, and illustrate that different species often have divergent responses to the same environmental and anthropogenic factors, and differ in the scales at which they respond to them. For all variables across all species, scale optimisation most often selected our largest scale. Precipitation, soil nutrients, and vegetation appeared to be the most important factors determining mammal distributions, likely through their associations with food resources for herbivores and, in turn, prey availability for carnivores. Anthropogenic pressures also had an important influence on habitat use, with many species selecting against areas with high cattle density. The variety of relationships with human density indicated that species vary in their tolerance of humans. We found a consistent positive relationship with areas under high protection, and negative relationship with unprotected and less-strictly protected areas. Policy implications: This study highlights the importance of adopting a multi-scale, multi-species approach for critical decision-making processes that depend on understanding wildlife distributions and habitat associations, such as protected area, corridor, and buffer zone prioritisation. We use our findings to identify changing rainfall patterns and increasing livestock numbers as two emerging trends that may impact wildlife distributions, both within sub-Saharan Africa and on a global scale.</p>
Comparing mixed models and Random Forest association tests using naturalGWAS and a Striped Bass SNP dataset
<p>In this study, we used the phenotype simulation package naturalGWAS to test the performance of Zhao's Random Forest method in comparison to an uncorrected Random Forest test, latent factor mixed models (LFMM), genome-wide efficient mixed models (GEMMA), and confounder adjusted linear regression (CATE). We created 400 sets of phenotypes, corresponding to five effect sizes and 2, 5, 15, or 30 causal loci, simulated from two empirical datasets containing SNPs from Striped Bass representing three and 13 populations. All association methods were evaluated for their ability to detect genotype-phenotype associations based on power, false discovery rates, and number of false positives. Genomic inflation was highest for uncorrected Random Forest and LFMM tests and lowest for Gemma and Zhao's Random Forest. All association tests had similar power to detect causal loci, and Zhao's Random Forest had the lowest false discovery rate in all scenarios. To measure the performance of association tests in small datasets with few loci surrounding a causal gene we also ran analyses again after removing causal loci from each dataset. All association tests were only able to find true positives, defined as loci located within 30k bp of a causal locus, in 3%–18% of simulations. In contrast, at least one false positive was found in 17%–44% of simulations. Zhao's Random Forest again identified the fewest false positives of all association tests studied. The ability to test the power of association tests for individual empirical datasets can be an extremely useful first step when designing a GWAS study.</p>
Using spectral reflectance and random forest method for modeling soil surface changes induced by simulated rainfall - datasets
<p>Using spectral reflectance and random forest method for modeling soil surface changes induced by simulated rainfall - datasets</p> <p>The impact of simulated rainfall on the soil surface roughness of different soil types with various initial surface states and the differences between their spectral characteristics were studied under laboratory conditions. The soil samples were collected from a horizon of fields near Poznań, western Poland. The physical and physicochemical properties of each soil sample were determined. Then, the part of the soil materials, consisting of natural aggregates, were used to form three soil surface roughness. </p> <p>An explanation of the table column names in the “soils properties.csv” file:</p> <p> </p> <ul> <li> <p>“textural classification” - Name of the granulometric group. Soil texture was determined by the hydrometer method according to standard PN-R-04032.</p> </li> <li> <p>“sand” - Sand content in the soil sample in %.</p> </li> <li> <p>“silt” – Silt content in the soil sample in %.</p> </li> <li> <p>“clay” – Clay content in the soil sample in %.</p> </li> <li> <p>pHH2O” - The pH of the soil sample determined in water. The soil pH was determined by the potentiometry method.</p> </li> <li> <p>“pHKCl” – The pH of the soil sample determined in KCl. The soil pH was determined by the potentiometry method.</p> </li> <li> <p>“SOC” – Organic matter content in soil was determined by oxidation titration using K2Cr2O7 with H2SO4 on the block mineralization.</p> </li> </ul> <p> </p> <p>An explanation of the table column names in the “rainfall doses.csv” file:</p> <p> </p> <ul> <li> <p>“rainfall simulation” - Rainfall simulation number.</p> </li> <li> <p>“rainfall dose” - One-time amount of rainfall dose expressed in millimeters.</p> </li> <li> <p>“accumulated rainfall” – Summation of rainfall after each successive dose expressed in millimeters.</p> </li> </ul> <p> </p> <p>An explanation of the table column names in the “soil measurements” file:</p> <p> </p> <ul> <li> <p>“textural classification” - Name of the granulometric group. Soil texture was determined by the hydrometer method according to standard PN-R-04032.</p> </li> <li> <p>“rainfall simulation” - Rainfall simulation number.</p> </li> <li> <p> “reflectance” - The amount of radiation reflected from the soil surface under the influence of successive rainfalls and expressed in nanometres. </p> </li> <li> <p>“roughness state” - The size of the roughness: R1 is the lowest soil roughness state, R2 represents medium soil roughness, and R3 represents the greatest roughness.</p> </li> <li> <p>“T3D” - Tortuosity index is a surface roughness index. It was calculated from DEM (Digital Elevation Model). It expresses the ratio between the true surface of DEM and its flat horizontal area.</p> </li> <li> <p>“HSD” - Height Standard Deviation is the second surface roughness index. It was calculated from DEM and expressed in millimeters. </p> </li> </ul> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><br> </p> <p> </p>
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
<p><span>Kill rates are a central parameter to assess the impact of predation on prey species. An accurate estimation of kill rates requires correct identification of kill sites, often achieved by field-checking GPS location clusters (GLCs). However, there are potential sources of error included in kill site identification, such as failing to detect GLCs that are kill sites and misclassifying the generated GLCs (e.g. kill for non-kill) that were not field-checked. Here, we address these two sources of error using a large GPS dataset of collared Eurasian lynx, an apex predator of conservation concern in Europe, in three multi-prey systems, with different combinations of wild, semi-domestic, and domestic prey. We first used a subsampling approach to investigate how different GPS-fix schedules affect the detection of GLCs indicating kill sites. Then, we evaluated the potential of the random forest algorithm to classify GLCs as non-kills, small prey kills, and ungulate kills. We show that the number of fixes can be reduced to from 7 to 3 fixes/night without missing more than 5% of the ungulate kills, in a system composed of wild prey. Reducing the number of fixes per 24-h decreased the probability of detecting GLCs connected with kill sites, particularly those of semi-domestic or domestic prey, and small prey. Random forest successfully predicted between 73%-90% of ungulate kills but failed to classify most small prey in all systems, with sensitivity (true positive rate) lower than 65%. Additionally, removing domestic prey improved the algorithm's overall accuracy. We provide a set of recommendations for studies focusing on kill site detection, which can be considered for other large carnivore species besides the Eurasian lynx. We recommend caution when working in systems including domestic prey, as the odds of underestimating kill rates are higher.</span></p>
iterative Random Forests data and analyses
<p>This repository contains scripts to run the simulations and case studies described in: <em>iterative Random Forests to discover predictive and stable high-order interactions</em>.</p>
Can ingredients based forecasting be learned? Disentangling a random forest's severe weather predictions
<p>Machine learning (ML)-based models have been rapidly integrated into forecast practices across the weather forecasting community in recent years. While ML tools introduce additional data to forecasting operations, there is a need for explainability to be available alongside the model output, such that the guidance can be transparent and trustworthy for the forecaster. This work makes use of the algorithm tree interpreter (TI) to disaggregate the contributions of meteorological features used in the Colorado State University Machine Learning Probabilities (CSU-MLP) system, a random forest-based ML tool that produces real-time probabilistic forecasts for severe weather using inputs from the Global Ensemble Forecast System v12. TI feature contributions are analyzed in time and space for CSU-MLP day-2 and 3 individual hazard (tornado, wind, and hail) forecasts and day-4 aggregate severe forecasts over a 2-yr period. For individual forecast periods, this work demonstrates that feature contributions derived from TI can be interpreted in an ingredients-based sense, effectively making the CSU-MLP probabilities physically interpretable. When investigated in an aggregate sense, TI illustrates that the CSU-MLP system's predictions use meteorological inputs in ways that are consistent with the spatiotemporal patterns seen in meteorological fields that pertain to severe storms climatology. This work concludes with a discussion on how these insights could be beneficial for model development, real-time forecast operations, and retrospective event analysis.</p>
Input data files for Dietrich et al. Chl-a and nutrient random forest modeling
<p>Input data for the models originally from:</p> <p>EPA, U. S. <em>WSIO Indicator Data Library</em>, <<a href="https://www.epa.gov/wsio/wsio-indicator-data-library">https://www.epa.gov/wsio/wsio-indicator-data-library</a>> (2023).</p> <p>Platt, L. R., Spaulding, S.A., Covert, A., Murphy, J.C., and Raynor, N. A national harmonized dataset of discrete chlorophyll from lakes and streams (2005-2022). (2023). <a href="https://doi.orghttps">https://doi.org:https://doi.org/10.5066/P9J0ZIOF</a></p> <p>Saad, D. A., Argue, D.M., Schwarz, G.E., Anning, D.W., Ator, S.W., Hoos, A.B., Preston, S.D., Robertson, D.M., and Wise, D.R., 2019. Water-quality and streamflow datasets used for estimating long-term mean daily streamflow and annual loads to be considered for use in regional streamflow, nutrient and sediment SPARROW models, United States, 1999-2014. (2019). <a href="https://doi.orghttps">https://doi.org:https://doi.org/10.5066/F7DN436B</a></p> <p> </p>
Trained Random Forest Model for PNW Seismic Event Classification Trained on 150s waveforms (P-50, P+100), 50 Hz, and 1-10 Hz BP Filtered
<p>This dataset contains three trained random forest models named as following - </p> <ul> <li>P_10_100_F_1_10_50.joblib - This is a model trained on 110s long waveforms (origin time - 10, origin time +100) in case of earthquakes and explosions and (first arrival pick -10, first arrival pick + 100) in case of surface events, the waveforms are tapered using 10% cosine taper, bandpass filtered between 1-10 Hz using Butterworth four corner filter, normalized and resampled to 50 Hz. </li> <li>P_50_100_F_1_10_50.joblib </li> <li>P_10_30_F_1_15_50.joblib. </li> </ul> <p>And also the standard scaler parameters for each features that will be used to normalize them. </p>
Supplementary Data: iterative Random Forests to discover predictive and stable high-order interactions
<p>This repository contains scripts and data for the simulations and case studies described in <em>iterative Random Forests to discover predictive and stable high-order interactions.</em></p>
Random Forest Ranks
<p>Details of Random Forest classifiers based on genotype to feature & feature to <em>Prakriti </em>analysis</p>
Mapping of glacial lakes using Sentinel-1 and Sentinel-2 data and a random forest classifier: Strengths and challenges
<p>The water body detection and mapping algorithm named 'glakemap' that I designed was aimed at specifically mapping glacial lakes across alpine regions where their detection and mapping are challenged by many factors such as shadows, cloud cover, turbidity, and ice surface. The algorithm uses Copernicus Sentinel-1 and -2 satellites data and machine learning model (random forest) in an integrated manner to automatically classify glacial lakes from other surface features. In specific, the algorithm takes Sentinel-1 and -2 satellites data as the main inputs. It calculates radar backscatter and Normalised Difference Water Indices (NDWIs) using these datasets, respectively. The radar backscatter and NDWIs products (images) are segmented using a set of rules producing many polygons including lake polygons. Lake polygons are then automatically separated/retained using the random forest model which is trained using features relevant to lakes.</p> <p>The dataset is also available at https://www.mountcryo.org/</p>
Full-coverage 1 km daily ambient PM2.5 and O3 concentrations of China in 2005-2017 based on multi-variable random forest model
<p>The aim of our study was to construct random forest models with high-performance, and estimate daily average PM<sub>2.5</sub> concentration and O<sub>3</sub> daily maximum 8h average concentration (O<sub>3</sub>-8hmax) of China in 2005-2017 at a spatial resolution of 1km×1km. The model variables included meteorological variables, satellite data, chemical transport model output, geographic variables and socioeconomic variables. Random forest model based on ten-fold cross validation was established, and spatial and temporal validations were performed to evaluate the model performance. According to our sample-based division method, the daily, monthly and yearly simulations of PM<sub>2.5</sub> gave average model fitting R<sup>2</sup> values of 0.85, 0.88 and 0.90, respectively; these R<sup>2</sup> values were 0.77, 0.77, and 0.69 for O<sub>3</sub>-8hmax, respectively. The meteorological variables and their lagged values can significantly affect both PM<sub>2.5</sub> and O<sub>3</sub>-8hmax simulations. During 2005-2017, PM<sub>2.5</sub> exhibited an overall downward trend, while ambient O<sub>3</sub> experienced an upward trend. Whilst the spatial patterns of PM<sub>2.5</sub> and O<sub>3</sub>-8hmax barely changed between 2005 and 2017, the temporal trend had spatial characteristic.</p> <p>Each dataset is the annual mean concentration of PM<sub>2.5</sub> or O<sub>3</sub>-8hmax based on the standard grid (Grid.csv) for that year. The coordinate system of the grid is WGS-84.</p>
Limitations of using surrogates for behaviour classification of accelerometer data: refining methods using random forest models in Caprids
<p>Animal-attached devices can be used on cryptic species to measure their movement and behaviour, enabling unprecedented insights into fundamental aspects of animal ecology and behaviour. However, direct observations of subjects are often still necessary to translate biologging data accurately into meaningful behaviours. As many elusive species cannot easily be observed in the wild, captive or domestic surrogates are typically used to calibrate data from devices. However, the utility of this approach remains equivocal. </p> <p>Here, we assess the validity of using captive conspecifics, and phylogenetically-similar domesticated counterparts (surrogate species) for calibrating behaviour classification. Tri-axial accelerometers and tri-axial magnetometers were used with behavioural observations to build random forest models to predict the behaviours. We applied these methods using captive Alpine ibex (Capra ibex) and a domestic counterpart, pygmy goats (Capra aegagrus hircus), to predict the behaviour including terrain slope for locomotion behaviours of captive Alpine ibex. </p> <p>Behavioural classification of captive Alpine ibex and domestic pygmy goats was highly accurate (> 98%). Model performance was reduced when using data split per individual, i.e., classifying behaviour of individuals not used to train models (mean ± sd = 56.1 ± 11%). Behavioural classifications using domestic counterparts, i.e., pygmy goat observations to predict ibex behaviour, however, were not sufficient to predict all behaviours of a phylogenetically similar species accurately (> 55%).</p> <p>We demonstrate methods to refine the use of random forest models to classify behaviours of both captive and free-living animal species. We suggest there are two main reasons for reduced accuracy when using a domestic counterpart to predict the behaviour of a wild species in captivity; domestication leading to morphological differences and the terrain of the environment in which the animals were observed. We also identify limitations when behaviour is predicted in individuals that are not used to train models. Our results demonstrate that biologging device calibration needs to be conducted using: (i) with similar conspecifics, and (ii) in an area where they can perform behaviours on terrain that reflects that of species in the wild.</p>
Houska_et_al_Dataset_Identifying the drivers of discharge and instream nitrate concentrations using Random Forest
<p>Dataset and Python code to perform a Random Forest regression analysis on mesured discharge and nitrate concentration data at sixteen points within the Sschwingbach Earth Observatory (SEO) located in Hüttenberg, Hesse, Germany.</p>
Data for Bradter, Altringham, Kunin, Thom, O'Connell & Benton: Variable ranking and selection with random forest for unbalanced data. Environmental Data Science
<p>The data are used in 'Bradter, Altringham, Kunin, Thom, O'Connell & Benton: Variable ranking and selection with random forest for unbalanced data. Environmental Data Science' and are described in the ReadMe file and in the manuscript and Supporting information.</p>
Data for: A new paradigm for medium-range severe weather forecasts: Probabilistic random forest-based predictions
<p>Historical observations of severe weather and simulated severe weather environments (i.e., features) from the Global Ensemble Forecast System v12 (GEFSv12) Reforecast Dataset (GEFS/R) are used in conjunction to train and test random forest (RF) machine learning (ML) models to probabilistically forecast severe weather out to days 4–8. RFs are trained with ~9 years of the GEFS/R and severe weather reports to establish statistical relationships. Feature engineering is briefly explored to examine alternative methods for gathering features around observed events, including simplifying features using spatial averaging and increasing the GEFS/R ensemble size with time-lagging. Validated RF models are tested with ~1.5 years of real-time forecast output from the operational GEFSv12 ensemble and are evaluated alongside expert human-generated outlooks from the Storm Prediction Center (SPC). Both RF-based forecasts and SPC outlooks are skillful with respect to climatology at days 4 and 5 with diminishing skill thereafter. The RF-based forecasts exhibit tendencies to slightly underforecast severe weather events, but they tend to be well-calibrated at lower probability thresholds. Spatially averaging predictors during RF training allows for prior-day thermodynamic and kinematic environments to generate skillful forecasts, while time-lagging acts to expand the forecast areas, increasing resolution but decreasing overall skill. The results highlight the utility of ML-generated products to aid SPC forecast operations into the medium range.</p>
Can ingredients based forecasting be learned? Disentangling a random forest's severe weather predictions
Open the record for dataset details and reuse information.
Data from: Demographic model selection using random forests and the site frequency spectrum
Open the record for dataset details and reuse information.
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
Open the record for dataset details and reuse information.
Random forests for predicting species identity of forensically important blow flies (Diptera: Calliphoridae) and flesh flies (Sarcophagidae) using geometric morphometric data: proof of concept
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.