Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.9.0
Dataset results
10 results for “Random Forest algorithm”
Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm
<p>Dataset from the results of data crawling via Twitter which discusses the Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>
Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"
<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm. </p> <p>Dataset S1. Seafloor density measurements. Columns are labeled with a header and include associated drilling project and measurement type for each sample. File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p>
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
<p><span>Kill rates are a central parameter to assess the impact of predation on prey species. An accurate estimation of kill rates requires correct identification of kill sites, often achieved by field-checking GPS location clusters (GLCs). However, there are potential sources of error included in kill site identification, such as failing to detect GLCs that are kill sites and misclassifying the generated GLCs (e.g. kill for non-kill) that were not field-checked. Here, we address these two sources of error using a large GPS dataset of collared Eurasian lynx, an apex predator of conservation concern in Europe, in three multi-prey systems, with different combinations of wild, semi-domestic, and domestic prey. We first used a subsampling approach to investigate how different GPS-fix schedules affect the detection of GLCs indicating kill sites. Then, we evaluated the potential of the random forest algorithm to classify GLCs as non-kills, small prey kills, and ungulate kills. We show that the number of fixes can be reduced to from 7 to 3 fixes/night without missing more than 5% of the ungulate kills, in a system composed of wild prey. Reducing the number of fixes per 24-h decreased the probability of detecting GLCs connected with kill sites, particularly those of semi-domestic or domestic prey, and small prey. Random forest successfully predicted between 73%-90% of ungulate kills but failed to classify most small prey in all systems, with sensitivity (true positive rate) lower than 65%. Additionally, removing domestic prey improved the algorithm's overall accuracy. We provide a set of recommendations for studies focusing on kill site detection, which can be considered for other large carnivore species besides the Eurasian lynx. We recommend caution when working in systems including domestic prey, as the odds of underestimating kill rates are higher.</span></p>
Eurasian lynx GLCs' characteristics for classification with random forest algorithm
Open the record for dataset details and reuse information.
Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm
<p><strong>Aims</strong>: Remote sensing approaches could be beneficial for monitoring and compiling essential biodiversity data because it is cost-effective and allows for coverage of large areas over a short period. This study investigated the relationship between multispectral remote sensing data from Landsat 8 and Sentinel 2 and species richness and diversity in mountainous and protected grasslands.</p> <p><strong>Locations</strong>: Golden Gate Highlands National Park, Free State, South Africa. </p> <p><strong>Methods</strong>: In-situ data of plant species composition and cover from 142 plots with 16 releves each were distributed across the study site and used to calculate species richness and Shannon-wiener species diversity index (species diversity. We used a machine-learning random forest algorithm to optimise the prediction of species richness and diversity. The algorithm was used to identify the optimal spectral bands and vegetation indices for estimating species richness and diversity. Subsequently, the selected bands and vegetation indices were used to estimate species richness through random forest regression. </p> <p><strong>Results</strong>: This research found weak relationships between remote sensing vegetation indices and the diversity metrics, but significant relationships were found between some spectral bands and diversity metrics. Moreover, using machine learning random forest, the multispectral datasets exhibited strong predictive powers. In this investigation, for both sensors, near-infrared (NIR) seemed to be the most selected band to explain species diversity in mountainous grasslands.</p> <p><strong>Main</strong> <strong>conclusions</strong>: This finding further ascertains the efficiency of using NIR in vegetation mapping. This research shows that NIR, SAVI and EVI are the most adequate for predicting species richness and diversity in mountainous grasslands with relatively good accuracies.</p>
Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm
Open the record for dataset details and reuse information.
High-resolution snow depth prediction using Random Forest algorithm with topographic parameters
Open the record for dataset details and reuse information.
tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction
GEO Series GSE189510. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
A Random-Forest Based Algorithm for Prediction of Enhancers From Histone Modifications
GEO Series GSE37858. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Data for: "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study"
<p>Data and software related to the manuscript "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study" (2020). </p> <p>========</p> <p>Typo in README_DataRep.txt:</p> <p>' 2) "one_config" [...] Selected results requiring these data are shown in Fig. <strong>13</strong>'<strong> </strong>(not 11).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.