Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forests”
Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany
<p>The dataset consists of particulate matter pollution concentration, measured in three localities - Hermsdorf, Charlottenburg and Adlershof, in Berlin, Germany.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer_rd_30s.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer_rd_30s.geojson</a> shows the observed PM2.5 concentration in a 30 second interval.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> shows the concentrations shown is the local concentration (observed concentration - background concentration) in a 30 second interval. The background concentration is calculated as the lowest 5 percentile of the measured concentration for each measurement round. </p> <p><a href="../api/records/10076056/draft/files/PM2.5_lc_max.geojson/content" target="_blank" rel="noopener noreferrer">PM2.5_lc_max.geojson</a> contains the information from <a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> in a 25m resolution. Additionally, it contains the land use information for each coordinate.</p> <p>The original publication providing all necessary background information on study sites, methodology and data processing is the following: Venkatraman Jagatha, J., T. Sauter, C. Schneider (2024): Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany. MDPI Sensors, 24(13), 4193, DOI: 10.3390/s24134193. The paper is fully open access and can be downloaded at <a href="https://doi.org/10.3390/s24134193">https://doi.org/10.3390/s24134193</a>.</p> <p>Information on working with geojson file can be found under <a href="https://geojson.readthedocs.io/en/latest/">GeoJSON</a> .</p>
Exploring AdaBoost and Random Forests machine learning approaches for infrared pathology on unbalanced data sets
<p>The use of infrared spectroscopy to augment decision-making in histopathology is a promising direction for the diagnosis of many disease types. Hyperspectral images of healthy and diseased tissue, generated by infrared spectroscopy, are used to build chemometric models that can provide objective metrics of disease state. It is important to build robust and stable models to provide confidence to the end user. The data used to develop such models can have a variety of characteristics which can pose problems to many model-building approaches. Here we have compared the performance of two machine learning algorithms – AdaBoost and Random Forests – on a variety of non-uniform data sets. Using samples of breast cancer tissue, we devised a range of training data capable of describing the problem space. Models were constructed from these training sets and their characteristics compared. In terms of separating infrared spectra of cancerous epithelium tissue from normal-associated tissue on the tissue microarray, both AdaBoost and Random Forests algorithms were shown to give excellent classification performance (over 95% accuracy) in this study. AdaBoost models were more robust when datasets with large imbalance were provided. The outcomes of this work are a measure of classification accuracy as a function of training data available, and a clear recommendation for choice of machine learning approach.</p>
Code for Random Forest models that predict pharmaceutical and water chemistry measurements in Baltimore Ecosystem Study streams
This file contains code to model the relationship between the water chemistry measurements and discharge measured as part of BES routine sampling and the pharmaceuticals measured in WY 2018. We use Random Forest models to predict 1) total (i.e., summed) concentration of the pharmaceuticals for which we screened, 2) total nutrient concentrations (TN & TP), 3) whether or not the antibiotic trimethoprim was detected in a given sample, and 4) whether or not nitrate and TP were above or below environmentally-relevant threshold concentrations. We also use RF models to predict N and P concentrations over a longer period, in order to compare models for nutrients to pharma. Code and analyses here rely on data processed in the file "BESPharma_WY2018.Rmd", published on EDI (doi:10.6073/pasta/610cb67fcbc8982c2af8ed946dce8ea5) and BES water chemistry data published on EDI (doi:10.6073/pasta/ce7f30e6013e003bfe28c5fd7d4aed23 )
Live Fuel Moisture Content Mapping in the Mediterranean Basin Using Random Forests and Combining MODIS Spectral and Thermal Data
<p>Live fuel moisture content (LFMC), defined as the mass of water in the foliage and small twigs relative to its total dry mass, is a key factor affecting fire potential and determining wildfire danger and activity. Fuel moisture is directly related to the amount of energy needed to evaporate water before ignition. Consequently, high moisture values reduce, or even inhibit, ignitability and subsequent fire spread.</p> <p>To cover the absence of a specific model to estimate LFMC for the Mediterranean Basin at the sub-continental scale, we built an empirical model based on Random Forests (LFMC<sub>RF</sub>) and combining MODIS spectral bands, vegetation indices, land surface temperature, and the day of year as predictors. The details on the modeling and validation methods, and the accuracy of the estimates are in the related publication <strong><a href="https://doi.org/10.3390/rs14133162">Cunill Camprubí et al., 2022</a></strong>.</p> <p>This dataset contains a collection of weekly LFMC maps from February 2000 to December 2021. The maps cover the Mediterranean and part of the Temperate biomes of the Mediterranean Basin. File <em>mapping_area_LFMC-RF_W-1.0.png</em> shows the target mapping areas.</p> <p>Metadata:</p> <ul> <li>Spectral Information: MODIS MCD43A4 C.6</li> <li>Land Surface Temperature: MODIS MOD11A2 C.6</li> <li>Land Cover Mask: MODIS MCD12Q1 C.6</li> <li>Coordinate Reference System: Native MODIS Sinusoidal</li> <li>Temporal Resolution: Weekly (W)</li> <li>Spatial Resolution: ~500 m</li> <li>File Format: NetCDF v.4</li> <li>Scale Factor: 0.01</li> </ul> <p>Fundings:</p> <p>The study was funded by the MICINN (RTI2018-094691-B-C31), European Union’s Horizon 2020-Research and Innovation Framework Programme under grant agreement no. 101003890 project FirEUrisk, the National Natural Science Foundation of China (U20A20179, 31850410483), and the talent proposals in Sichuan Province (2020JDRC0065) from Southwest University of Science and Technology (18ZX7131).</p>
Datasets and results from: "Random Forest Classification and Solar Flares Data: Analysis and Validation"
<p><strong>Instructions for the data and code repository</strong></p> <p>Results, post-processing workflow, and datasets for the research paper titled "Random Forest Classification and Solar Flares Data: Analysis and Validation".</p> <p>The folder contains three .csv files: the complete dataset (dataset.csv), the balanced training dataset (train_dataset.csv), and the testing dataset (test_dataset.csv).</p> <p>The folder also contains the result files from the research (.csv output files with predictions and .html files with evaluation metrics, etc.) exported from the JASP software. The number in each file name corresponds to the number of trees utilized in Random Forest modelling.</p> <p>In addition, the Python script for the post-processing workflow is provided, with comments located in the script.</p> <p>The soft range X-ray irradiance and VLF amplitude data were obtained from:<br> National Centers for Environmental Information (NCEI) Available online: https://www.ncei.noaa.gov/. Accessed on: 24th June 2023. <br> Worldwide archive of low-frequency data and observations (WALDO) Available online: https://waldo.world/. Accessed on: 24th June 2023.</p>
Sensor and nutrient data associated with the article Harrison et al. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression
This document describes a dataset used to produce Random Forests Regression models of stream nitrogen and phosphorus concentrations from high-frequency sensor data, as reported in: Harrison, J.W., Lucius, M.A., Farrell, J.L., Eichler, L.W., and Relyea, R.A. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression. Science of the Total Environment: https://doi.org/10.1016/j.scitotenv.2020.143005. The dataset consists of paired values of stream nitrogen and phosphorus concentrations and various high-frequency sensor parameters (water temperature, specific conductance, pH, fluorescent dissolved organic matter, turbidity, hydrostatic pressure, soil moisture) collected during baseflow and storm events from 2018 to 2019 as part of routine monitoring of eleven tributaries of Lake George, New York. This dataset does not include raw data; two levels of processing were performed: (1) erroneous values (extreme or otherwise outlying values with no apparent environmental cause) were removed from the sensor data as part of the routine QA/QC process of the Jefferson Project, and (2) one-hour rolling medians of the raw sensor data were calculated at a 1-minute timestep to maximize pairing of sensor data with nutrient concentrations. The resultant dataset was used to train and test the models presented in Harrison et al. 2020.
ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)
<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>
A computational intelligence approach to predict energy demand using Random Forest in a Cloudera cluster
<p>Society’s energy consumption has shot up in recent years, making the prediction of its demand a current challenge to ensure an efficient and responsible use. Artificial intelligence techniques have proven to be potential tools in handling tedious tasks and making sense of large-scale data to make better business decisions in different areas of knowledge. In this article, the use of random forests algorithms in a Big Data environment is proposed for households energy demand forecasting. The predictions are based on the use of information from different sources, confirming a fundamental role of socioeconomic data in consumer’s behaviours. On the other hand, the use of Big Data architectures is proposed to perform horizontal and vertical scaling of the solution to be used in real environments. Finally, a tool for high-resolution predictions with great efficiency is introduced, which enables energy management in a very accurate way.</p> <p>Raw data is incuded in data.csv. This file contains half hourly home electricity consumption registers for 4404 households with fix tariffs (not subject to dynamic time of use) for a period between November 2011 and February 2014. Original information was acquired from the Low Carbon London project led by UK Power Networks (https://data.london.gov.uk/dataset/smartmeter-energy-use-data-in-london-households)</p> <p>RFResults.zip contains the energy predictions for each ACORN group using the generated Random Forest algorithm. For this purpose, the first 613 days of a total of 818 observations of each group were considered for training and the last 205 days for testing.</p> <p>Meteorological data was adquired from the darksky app (https://darksky.net). These data are included in the weather_hourly_darksky.csv</p> <p>uk_bank_holidays. xlsx contains the dated of UK bank holidays for the studied period, used as additional variable related to occupancy</p>
Global gross primary production (GPP) product generated by data fusion based on random forest
<p>Improving the ability of gross primary production (GPP) estimates to capture extreme climate perturbations and reduce the uncertainty of GPP response processes to extreme climate is a new challenge. Based on the random forest algorithm, we integrated the multimodel GPP simulation results published by the Multiscale Synthesis and Terrestrial Model Intercomparison Project, the FLUXNET flux-site-observed GPP, the standardized precipitation index (SPI) and the standardized temperature index (STI) to generate a set of global GPP time-series data products from 2001 to 2010. The new GPP product was named DFRF-GPP, referring to the GPP generated by data fusion based on random forest. DFRF-GPP is highly reliable and can be used as a valuable data source for various applications, especially in high-temperature and drought-related studies.</p>
Random forest regression for fuzzy grades
<p>The used dataset contained information about:</p> <ul> <li>students' gender (1 = male, 2 = female);</li> <li>schools' type (1 = scientific lyceums, 2 = other lyceums, 3 = technical schools, 4 = vocational schools);</li> <li>schools' macreregion (1 = Northwester Italy, 2 = Northeastern Italy, 3 = Central Italy, 4 = Southern Italy, 5 = Southern Italy and Isles);</li> <li>students' origin (1 = native Italian student, 2 = first-generation immigrant student, 3 = second-generation immigrant student);</li> <li>students' ESCS (continuous data);</li> <li>teacher-given grades in mathematics (from 1 to 10);</li> <li>students' achievements on the INVALSI mathematics test (continuous data);</li> <li>students' fuzzy grade, obtained as a combination of teacher-given grades and achievements in mathematics (from 1 to 10).</li> </ul>
Identification of high-wind features within extratropical cyclones using a probabilistic random forest - Part 2: Climatology - Dataset
<p>This dataset provides output of RAMEFI for the wind feature climatology presented in Eisenstein et al. (2023; 10.5194/wcd-2023-10) for the winter months October to March 2000-2019 using COSMO-REA6 (https://reanalysis.meteo.uni-bonn.de/?COSMO-REA6).</p> <p><strong>rf_crea_<yyyymm>.nc</strong> include the unfiltered probabilities for 'no feature' (p0), warm jet (p1), cold-frontal convection (p2), cold jet (p3) and cold-sector winds (p5) for each month.</p> <p>To filter for cyclone tracks, use <strong>cyclone_tracks.csv</strong>. The<strong> </strong>file includes interpolated ERA5 cyclone tracks for the investigated area and time period (see Section 2.4 of the paper).</p> <p><strong>mask.nc</strong> includes a land sea mask, height of surface level and a mask to exclude certain grid points as discussed in the manuscript (e.g., grid points with an altitude over 800m and the Balkans) for further filtering.</p>
American Goshawk habitat data from nest stands and random points within the Minidoka Ranger District, Sawtooth National Forest, USA
This data supported analysis of American Goshawk (Astur atricapillus) nest stand habitat and was collected within the Minidoka Ranger District of the Sawtooth National Forest in southern Idaho and northern Utah from 2017-2020. The central goal of this research was to develop management tools that demonstrate the utility of conducting analyses at multiple spatial scales as well as using both parametric and machine learning approaches. The stand-level dataset includes variables collected by hand in the field at nest stands and paired random forested sites 300 meters away. It also includes some terrain variables based on remote sensing data. Variables included in the stand-level data table include nest, distance to edge, distance to road, distance to water, division, dominant tree species, canopy closure, Stand Density Index (SDI), Trees per hectare, elevation, slope, Topographic Position Index (TPI), northness, eastness, Diameter at Breast Height (DBH), DBH variance, tree height, tree height variance, and crown depth. We recommend that the stand-level data be used to identify relevant variables and their thresholds for forest managers due to its high resolution. The forest-wide dataset includes only variables collected using various remote sensing datasets at nests and random forested points throghout the Minidoka Ranger District of the Sawtooth National Forest. Variables included in the forest-wide data table include nest, canopy closure, elevation, slope, TPI, northness, eastness, distance to road, distance to water, distance to edge, tree height, and crown depth. We recommend that the forest-wide data be used to identify areas of high suitability for goshawk occupancy across the study area along with sites that could become suitable habitat with management intervention. Latitude and longitude data, while used in our analyses, are excluded from the data tables to protect breeding goshawks from disturbance.
A Dataset of Pull Requests and A Trained Random Forest Model for predicting Pull Request Acceptance
<p>A Curated Dataset of 470,925 pull requests for 3349 popular NPM packages, description of the variables, code snippet for creating a Random Forest model for predicting pull request acceptance, and a pre-trained Random Forest model (in R). The dataset is for the ESEM-2020 paper: "Impact of Technical and Social Factors on Pull Request Quality for the NPM Ecosystem" (<a href="https://arxiv.org/abs/2007.04816">https://arxiv.org/abs/2007.04816</a>). </p> <p>Citation:</p> <pre>@inproceedings{dey2020effect, title={Effect of technical and social factors on pull request quality for the npm ecosystem}, author={Dey, Tapajit and Mockus, Audris}, booktitle={Proceedings of the 14th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM)}, pages={1--11}, year={2020} }</pre>
Dataset: Random forest models of ultra-low frequency magnetospheric wave power.
<p>Predictive models of ground-based ultra-low frequency (ULF, 1-15 mHz) wave power, corresponding to magnetospheric waves. The series of decision tree ensembles (random forests) are dependent on solar wind properties, latitude and azimuthal angle around the Earth (magnetic local time, MLT).</p>
Figure 7 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 7. Summary of Random Forests classifications for each empirical comparison. Each row shows results from the stratum with the smallest fraction of individuals correctly classified, with comparisons labeled by their taxonomic codes as listed in Table 2. Colors identify comparison type as species (blue), subspecies (green), and populations (red). Points show the fraction of individuals correctly classified with probabilities> 50% (PD50, circles), and> 95% (PD95, triangles). Thin colored lines show 95% confidence intervals (CI) around PD50 estimates. Gray bars show range of a priori random classification rates based on individual size (left) to maximum possible classification rates based on shared haplotypes (right).
Figure 6 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 6. Frequency distributions of the change in observed diagnosability (x-axis) in the simulated data for increasing levels of the probability of misstratification (vertical panels). Figures on the left and right columns are censored by data sets for original diagnosability ≤50% and>50%, respectively.
Figure 5 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 5. Two-dimensional GAM fits of theta (Ɵ), number of migrants (Nem), and divergence time in generations (T) from Model 2 simulated data. From left to right, columns show results from models without migration (m = 0), with migration and Nem <1, and Nem ≥ 1. Colors indicate model prediction of percent correctly classified.
Figure 4 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 4. GAM fit of number of migrants (Nem) from Model 2 parameters. Solid line shows median value of predicted percent correctly classified, and shaded area shows 95% CI. The switch from bimodal distribution to a normal distribution occurs at Nem = 1 (log10Nem = 0).
Figure 3 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 3. Two-dimensional GAM fits of effective population size (Ne), divergence time in generations (T), and mutation rate (µ) from Model 1 simulated data. Results from models without migration to the left and those with migration to the right. Colors indicate model prediction of percent correctly classified.
Figure 1 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 1. (A) Distribution of a hypothetical character for two putative subspecies (red and blue) demonstrating minimum overlap necessary to satisfy 75% rule of Amadon (1949). Character is continuous on the x-axis. Dashed lines indicate the point at which 75% of each distribution is outside of 99%+ of the other. Solid line indicates point of overlap where 97% of both distributions are outside one another. (B) Probability of membership to subspecies for specimens having values along the character axis. Probability is based on the ratio of the distribution frequencies at each point along the x-axis, with a 50:50 probability occurring at the threshold point.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.