Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “random forest regression”
Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany
<p>The dataset consists of particulate matter pollution concentration, measured in three localities - Hermsdorf, Charlottenburg and Adlershof, in Berlin, Germany.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer_rd_30s.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer_rd_30s.geojson</a> shows the observed PM2.5 concentration in a 30 second interval.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> shows the concentrations shown is the local concentration (observed concentration - background concentration) in a 30 second interval. The background concentration is calculated as the lowest 5 percentile of the measured concentration for each measurement round. </p> <p><a href="../api/records/10076056/draft/files/PM2.5_lc_max.geojson/content" target="_blank" rel="noopener noreferrer">PM2.5_lc_max.geojson</a> contains the information from <a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> in a 25m resolution. Additionally, it contains the land use information for each coordinate.</p> <p>The original publication providing all necessary background information on study sites, methodology and data processing is the following: Venkatraman Jagatha, J., T. Sauter, C. Schneider (2024): Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany. MDPI Sensors, 24(13), 4193, DOI: 10.3390/s24134193. The paper is fully open access and can be downloaded at <a href="https://doi.org/10.3390/s24134193">https://doi.org/10.3390/s24134193</a>.</p> <p>Information on working with geojson file can be found under <a href="https://geojson.readthedocs.io/en/latest/">GeoJSON</a> .</p>
Sensor and nutrient data associated with the article Harrison et al. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression
This document describes a dataset used to produce Random Forests Regression models of stream nitrogen and phosphorus concentrations from high-frequency sensor data, as reported in: Harrison, J.W., Lucius, M.A., Farrell, J.L., Eichler, L.W., and Relyea, R.A. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression. Science of the Total Environment: https://doi.org/10.1016/j.scitotenv.2020.143005. The dataset consists of paired values of stream nitrogen and phosphorus concentrations and various high-frequency sensor parameters (water temperature, specific conductance, pH, fluorescent dissolved organic matter, turbidity, hydrostatic pressure, soil moisture) collected during baseflow and storm events from 2018 to 2019 as part of routine monitoring of eleven tributaries of Lake George, New York. This dataset does not include raw data; two levels of processing were performed: (1) erroneous values (extreme or otherwise outlying values with no apparent environmental cause) were removed from the sensor data as part of the routine QA/QC process of the Jefferson Project, and (2) one-hour rolling medians of the raw sensor data were calculated at a 1-minute timestep to maximize pairing of sensor data with nutrient concentrations. The resultant dataset was used to train and test the models presented in Harrison et al. 2020.
Random forest regression for fuzzy grades
<p>The used dataset contained information about:</p> <ul> <li>students' gender (1 = male, 2 = female);</li> <li>schools' type (1 = scientific lyceums, 2 = other lyceums, 3 = technical schools, 4 = vocational schools);</li> <li>schools' macreregion (1 = Northwester Italy, 2 = Northeastern Italy, 3 = Central Italy, 4 = Southern Italy, 5 = Southern Italy and Isles);</li> <li>students' origin (1 = native Italian student, 2 = first-generation immigrant student, 3 = second-generation immigrant student);</li> <li>students' ESCS (continuous data);</li> <li>teacher-given grades in mathematics (from 1 to 10);</li> <li>students' achievements on the INVALSI mathematics test (continuous data);</li> <li>students' fuzzy grade, obtained as a combination of teacher-given grades and achievements in mathematics (from 1 to 10).</li> </ul>
Output - Results of Random Forest and Multiple Linear regression analysis.
<p><strong>Hybrid streamflow modelling using machine learning and multi-model combination.</strong></p> <p> </p> <p><strong>Structure:</strong></p> <p><strong>MLR_output:</strong></p> <ul> <li>Validate <ul> <li>Different setups</li> </ul> </li> </ul> <p><strong>RF_output:</strong></p> <ul> <li>tune <ul> <li><em>all_stations</em></li> </ul> </li> <li>train <ul> <li><em>Different setups</em></li> </ul> </li> <li>Validate <ul> <li><em>Different setups</em></li> </ul> </li> </ul>
Use of Vegetation Change Tracker, Spatial Analysis, and Random Forest Regression to Assess the Evolution of Plantation Stand Age in Southeast China
<p>It is crucial to determine the spatio-temporal distribution patterns of forest ages across wide regions, as forest management plans and practices, and ecosystem carbon budgeting are highly dependent on these. However, given frequent deforestation events (e.g., harvesting) and rapid recovery of plantation stands in Southern China, field-based forest age measurements over wide regions are time-consuming, labour-intensive, and costly. In the current study, we mapped the spatio-temporal patterns of forest stand ages across three typical plantations in Southern China. This was accomplished by using two new feasible and accurate methods, 1) integrating vegetation change tracker (VCT) algorithm and spatial analysis (VCT-SA) for the pixels that were disturbed at least once from 1987 to 2017, and 2) integrating VCT and random forest (VCT-RF) for the pixels were not disturbed during the study period. The results revealed the spatio-temporal distribution of age structure, which indicated that the plantation stands in our large study area were increasingly aging.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.