Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
17
datasets available to search
ShareScore release 0.9.0
Dataset results
17 results for “Random regression”
Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany
<p>The dataset consists of particulate matter pollution concentration, measured in three localities - Hermsdorf, Charlottenburg and Adlershof, in Berlin, Germany.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer_rd_30s.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer_rd_30s.geojson</a> shows the observed PM2.5 concentration in a 30 second interval.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> shows the concentrations shown is the local concentration (observed concentration - background concentration) in a 30 second interval. The background concentration is calculated as the lowest 5 percentile of the measured concentration for each measurement round. </p> <p><a href="../api/records/10076056/draft/files/PM2.5_lc_max.geojson/content" target="_blank" rel="noopener noreferrer">PM2.5_lc_max.geojson</a> contains the information from <a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> in a 25m resolution. Additionally, it contains the land use information for each coordinate.</p> <p>The original publication providing all necessary background information on study sites, methodology and data processing is the following: Venkatraman Jagatha, J., T. Sauter, C. Schneider (2024): Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany. MDPI Sensors, 24(13), 4193, DOI: 10.3390/s24134193. The paper is fully open access and can be downloaded at <a href="https://doi.org/10.3390/s24134193">https://doi.org/10.3390/s24134193</a>.</p> <p>Information on working with geojson file can be found under <a href="https://geojson.readthedocs.io/en/latest/">GeoJSON</a> .</p>
Sensor and nutrient data associated with the article Harrison et al. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression
This document describes a dataset used to produce Random Forests Regression models of stream nitrogen and phosphorus concentrations from high-frequency sensor data, as reported in: Harrison, J.W., Lucius, M.A., Farrell, J.L., Eichler, L.W., and Relyea, R.A. 2020. Prediction of stream nitrogen and phosphorus concentrations from high-frequency sensors using Random Forests Regression. Science of the Total Environment: https://doi.org/10.1016/j.scitotenv.2020.143005. The dataset consists of paired values of stream nitrogen and phosphorus concentrations and various high-frequency sensor parameters (water temperature, specific conductance, pH, fluorescent dissolved organic matter, turbidity, hydrostatic pressure, soil moisture) collected during baseflow and storm events from 2018 to 2019 as part of routine monitoring of eleven tributaries of Lake George, New York. This dataset does not include raw data; two levels of processing were performed: (1) erroneous values (extreme or otherwise outlying values with no apparent environmental cause) were removed from the sensor data as part of the routine QA/QC process of the Jefferson Project, and (2) one-hour rolling medians of the raw sensor data were calculated at a 1-minute timestep to maximize pairing of sensor data with nutrient concentrations. The resultant dataset was used to train and test the models presented in Harrison et al. 2020.
Random forest regression for fuzzy grades
<p>The used dataset contained information about:</p> <ul> <li>students' gender (1 = male, 2 = female);</li> <li>schools' type (1 = scientific lyceums, 2 = other lyceums, 3 = technical schools, 4 = vocational schools);</li> <li>schools' macreregion (1 = Northwester Italy, 2 = Northeastern Italy, 3 = Central Italy, 4 = Southern Italy, 5 = Southern Italy and Isles);</li> <li>students' origin (1 = native Italian student, 2 = first-generation immigrant student, 3 = second-generation immigrant student);</li> <li>students' ESCS (continuous data);</li> <li>teacher-given grades in mathematics (from 1 to 10);</li> <li>students' achievements on the INVALSI mathematics test (continuous data);</li> <li>students' fuzzy grade, obtained as a combination of teacher-given grades and achievements in mathematics (from 1 to 10).</li> </ul>
Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.
<p>Cloud computing is one of the most popular and sophisticated technologies adopted by organizations worldwide. Some world-leading organizations enhance their efficiency and effectiveness by using cloud computing technology. Working from home (WFH) has been a popular trend among organizations during the coronavirus (COVID-19) pandemic. The COVID-19 saw a breakthrough in work cultures and environments where working from home was a remarkable success in remote working environments, despite being a rare phenomenon in Sri Lanka. Yet, it is argued that the deployment of work from home has not been effective among Sri Lankan business organizations due to a lack of IT infrastructure, facilities, and knowledge. The purpose of the study is to investigate the impact of cloud computing, embracing the service models (Infrastructure as a Service, Platform as a Service, and Software as a Service) as theoretical lenses and testing the COVID-19 as the moderator. The study has been conducted based on a deductive approach and adopted a stratified random sampling method. The sample consisted of 384 IT employees among those who had experienced working from home. The study utilized multiple regression and found that cloud computing service models significantly impact work from home with the moderating effect of COVID-19.</p>
Data for: Leveraging spatio-temporal genomic breeding value estimates of dry matter yield and herbage quality in ryegrass via random regression models
<p>Joint modeling of correlated multi-environment and multi-harvest data of perennial crop species may offer advantages in prediction schemes and a better understanding of the underlying dynamics in space and time. The goal of the present study was to investigate the relevance of incorporating the longitudinal dimension of within-season multiple measurements of forage perennial ryegrass traits in a reaction norm model setup that additionally accounts for genotype-environment interactions (G×E). Genetic parameters and accuracy of genomic breeding value (gEBV) predictions were investigated by fitting three random regression models (gRRM) using Legendre polynomial functions to the data. Genomic DNA sequencing of family pools of diploid perennial ryegrass was performed using DNA nanoball-based technology and yielded 56,645 single nucleotide polymorphisms which were used to calculate the allele frequency-based genomic relationship matrix. Biomass yield's estimated additive genetic variance and heritability values were higher in later harvests. The additive genetic correlations were moderate to low in early measurements and peaked at intermediates, with fairly stable values across the environmental gradient, except for the initial harvest data collection. This led to the conclusion that complex (G×E) arises from spatial and temporal dimensions in the early season, with lower re-ranking trends thereafter. In general, modeling the temporal dimension with a second-order orthogonal polynomial improved the accuracy of gEBV prediction for nutritive quality traits, but no gain in prediction accuracy was detected for dry matter yield. This study leverages the flexibility and usefulness of gRRM models for perennial ryegrass breeding and can be readily extended to other multi-harvest crops.</p>
Randomized Trial of Imaging Versus Risk Factor-Based Therapy for Plaque Regression
ClinicalTrials.gov study NCT01212900. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Data for: Leveraging spatio-temporal genomic breeding value estimates of dry matter yield and herbage quality in ryegrass via random regression models
Open the record for dataset details and reuse information.
Single-incision laparoscopic cholecystectomy versus conventional multi-port laparoscopic cholecystectomy: A systematic review, meta-analysis, and meta-regression of randomized controlled trials
<p>Single-incision laparoscopic cholecystectomy versus conventional multi-port laparoscopic cholecystectomy: A systematic review, meta-analysis, and meta-regression of randomized controlled trials</p>
Output - Results of Random Forest and Multiple Linear regression analysis.
<p><strong>Hybrid streamflow modelling using machine learning and multi-model combination.</strong></p> <p> </p> <p><strong>Structure:</strong></p> <p><strong>MLR_output:</strong></p> <ul> <li>Validate <ul> <li>Different setups</li> </ul> </li> </ul> <p><strong>RF_output:</strong></p> <ul> <li>tune <ul> <li><em>all_stations</em></li> </ul> </li> <li>train <ul> <li><em>Different setups</em></li> </ul> </li> <li>Validate <ul> <li><em>Different setups</em></li> </ul> </li> </ul>
Comparing Re-TACE Versus SABR for Post-prior-TACE Incompletely Regressed HCC: a Randomized Controlled Trial (TASABR)
ClinicalTrials.gov study NCT02921139. IPD Sharing: UNDECIDED. Countries: 1. Publications: 7.
Effect of Timolol on Refractive Outcomes in Eyes With Myopic Regression After LASIK: a Randomized Clinical Trial
ClinicalTrials.gov study NCT01506635. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Risk of cancer with angiotensin-receptor blockers increases with increasing cumulative exposure: Meta-regression analysis of randomized trials
<p>Angiotensin-receptor blockers (ARBs) are a class of drugs approved for the treatment of several common conditions, such as hypertension and heart failure. Recently, regulatory agencies have started to identify possibly carcinogenic nitrosamines and azido compounds in a multitude of formulations of several ARBs, resulting in progressive recalls. Furthermore, data from several randomized controlled trials suggested that there is also a clinically increased risk of cancer and specifically lung cancer with ARBs; whereas other trials suggested no increased risk. The purpose of this analysis was to provide additional insight into the ARB-cancer link by examining whether there is a relationship between degree of cumulative exposure to ARBs and risk of cancer in randomized trials. Trial-level data from ARB Trialists Collaboration including 15 randomized controlled trials was extracted and entered into meta-regression analyses. The two co-primary outcomes were the relationship between cumulative exposure to ARBs and risk of all cancers combined and the relationship between cumulative exposure and risk of lung cancer. A total of 74,021 patients were randomized to an ARB resulting in a total cumulative exposure of 172,389 person-years of exposure to daily high dose (or equivalent). 61,197 patients were randomized to control. There was a highly significant correlation between the degree of cumulative exposure to ARBs and risk of all cancers combined (slope=0.07 [95% CI 0.03 to 0.11], p<0.001), and also lung cancer (slope=0.16 [95% CI 0.05 to 0.27], p=0.003). Accordingly, in trials where the cumulative exposure was greater than 3 years of exposure to daily high dose, there was a statistically significant increase in risk of all cancers combined (I<sup>2</sup>=31.4%, RR 1.11 [95% CI 1.03 to 1.19], p=0.006). There was a statistically significant increase in risk of lung cancers in trials where the cumulative exposure was greater than 2.5 years (I<sup>2</sup>=0%, RR 1.21 [95% CI 1.02 to 1.44], p=0.03). In trials with lower cumulative exposure to ARBs, there was no increased risk of all cancers combined or lung cancer. Cumulative exposure-risk relationship with ARBs was independent of background angiotensin-converting enzyme inhibitor treatment or the type of control (i.e. placebo or non-placebo control). Since this is a trial-level analysis. the effects of patient characteristics such as age and smoking status could not be examined due to lack of patient-level data. In conclusion, this analysis, for the first time, reveals that risk of cancer with ARBs (and specifically lung cancer) increases with increasing cumulative exposure to these drugs. The excess risk of cancer with long-term ARB use has public health implications.</p>
Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait
In prediction of genomic values, single-step method has been demonstrated to outperform multi-step methods. In statistical analyses of longitudinal traits, random regression test-day model (RR-TDM) has clear advantages over other models. Our goal in this study was to evaluate the performance of the model integrating both single-step and RR-TDM prediction methods, called single-step random regression test-day model (SS RR-TDM), in comparison with the pedigree-based RR-TDM and genomic best linear unbiased prediction (GBLUP) model. We performed extensive simulations to exploit potential advantages of SS RR-TDM over the other two models under various scenarios with different level of heritability, the number of QTL as well as the selection scheme. SS RR-TDM was found to achieve the highest accuracy and unbiasedness under all scenarios, exhibiting robust prediction ability in longitudinal trait analyses. Moreover, SS RR-TDM showed better persistency of accuracy over generations than GBLUP model. In addition, we also found that the SS RR-TDM had advantages over RR-TDM and GBLUP in terms of a real dataset of human contributed by the GAW18 workshop. The findings in our study firstly proved the feasibility and advantages of the SS RR-TDM, and further enhanced strategies for the genomic prediction of longitudinal traits in the future.
Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait
Open the record for dataset details and reuse information.
Data from: Risk of cancer with angiotensin-receptor blockers increases with increasing cumulative exposure: Meta-regression analysis of randomized trials
Open the record for dataset details and reuse information.
The Effectiveness of Lumber Regression Technique on Disc Bulge Randomized Control Trial
ClinicalTrials.gov study NCT06548334. IPD Sharing: NO. Countries: 1. Publications: 0.
Use of Vegetation Change Tracker, Spatial Analysis, and Random Forest Regression to Assess the Evolution of Plantation Stand Age in Southeast China
<p>It is crucial to determine the spatio-temporal distribution patterns of forest ages across wide regions, as forest management plans and practices, and ecosystem carbon budgeting are highly dependent on these. However, given frequent deforestation events (e.g., harvesting) and rapid recovery of plantation stands in Southern China, field-based forest age measurements over wide regions are time-consuming, labour-intensive, and costly. In the current study, we mapped the spatio-temporal patterns of forest stand ages across three typical plantations in Southern China. This was accomplished by using two new feasible and accurate methods, 1) integrating vegetation change tracker (VCT) algorithm and spatial analysis (VCT-SA) for the pixels that were disturbed at least once from 1987 to 2017, and 2) integrating VCT and random forest (VCT-RF) for the pixels were not disturbed during the study period. The results revealed the spatio-temporal distribution of age structure, which indicated that the plantation stands in our large study area were increasingly aging.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.