Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “Random Forest algorithm”

Learn how ShareScore rates datasets ↗
zenodo40/100

Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm

<p>Dataset from the results of data crawling via Twitter which discusses the&nbsp;Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>

openother-openSep 2022View details →
zenodo36/100

Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"

<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm.&nbsp;</p> <p>Dataset S1.&nbsp;Seafloor density measurements.&nbsp; Columns are labeled with a header and include associated drilling project and measurement type for each sample.&nbsp;&nbsp;File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5&times;5-arc minute resolution.&nbsp; Units are g/cm^3.&nbsp; File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5&times;5-arc minute resolution.&nbsp; Units are g/cm^3.&nbsp; File format: netCDF (.nc)</p>

opencc-by-4.0Sep 2020View details →
dryad36/100

Eurasian lynx GLCs' characteristics for classification with random forest algorithm

<p><span>Kill rates are a central parameter to assess the impact of predation on prey species. An accurate estimation of kill rates requires correct identification of kill sites, often achieved by field-checking GPS location clusters (GLCs). However, there are potential sources of error included in kill site identification, such as failing to detect GLCs that are kill sites and misclassifying the generated GLCs (e.g. kill for non-kill) that were not field-checked. Here, we address these two sources of error using a large GPS dataset of collared Eurasian lynx, an apex predator of conservation concern in Europe, in three multi-prey systems, with different combinations of wild, semi-domestic, and domestic prey. We first used a subsampling approach to investigate how different GPS-fix schedules affect the detection of GLCs indicating kill sites. Then, we evaluated the potential of the random forest algorithm to classify GLCs as non-kills, small prey kills, and ungulate kills. We show that the number of fixes can be reduced to from 7 to 3 fixes/night without missing more than 5% of the ungulate kills, in a system composed of wild prey. Reducing the number of fixes per 24-h decreased the probability of detecting GLCs connected with kill sites, particularly those of semi-domestic or domestic prey, and small prey. Random forest successfully predicted between 73%-90% of ungulate kills but failed to classify most small prey in all systems, with sensitivity (true positive rate) lower than 65%. Additionally, removing domestic prey improved the algorithm's overall accuracy. We provide a set of recommendations for studies focusing on kill site detection, which can be considered for other large carnivore species besides the Eurasian lynx. We recommend caution when working in systems including domestic prey, as the odds of underestimating kill rates are higher.</span></p>

opencc-zeroOct 2022View details →
dryad36/100

Eurasian lynx GLCs' characteristics for classification with random forest algorithm

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad32/100

Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm

<p><strong>Aims</strong>: Remote sensing approaches could be beneficial for monitoring and compiling essential biodiversity data because it is cost-effective and allows for coverage of large areas over a short period. This study investigated the relationship between multispectral remote sensing data from Landsat 8 and Sentinel 2 and species richness and diversity in mountainous and protected grasslands.</p> <p><strong>Locations</strong>: Golden Gate Highlands National Park, Free State, South Africa. </p> <p><strong>Methods</strong>: In-situ data of plant species composition and cover from 142 plots with 16 releves each were distributed across the study site and used to calculate species richness and Shannon-wiener species diversity index (species diversity. We used a machine-learning random forest algorithm to optimise the prediction of species richness and diversity. The algorithm was used to identify the optimal spectral bands and vegetation indices for estimating species richness and diversity. Subsequently, the selected bands and vegetation indices were used to estimate species richness through random forest regression. </p> <p><strong>Results</strong>: This research found weak relationships between remote sensing vegetation indices and the diversity metrics, but significant relationships were found between some spectral bands and diversity metrics. Moreover, using machine learning random forest, the multispectral datasets exhibited strong predictive powers. In this investigation, for both sensors, near-infrared (NIR) seemed to be the most selected band to explain species diversity in mountainous grasslands.</p> <p><strong>Main</strong> <strong>conclusions</strong>: This finding further ascertains the efficiency of using NIR in vegetation mapping.  This research shows that NIR, SAVI and EVI are the most adequate for predicting species richness and diversity in mountainous grasslands with relatively good accuracies.</p>

opencc-zeroMay 2023View details →
dryad32/100

Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm

Open the record for dataset details and reuse information.

publicMay 2023View details →
zenodo28/100

High-resolution snow depth prediction using Random Forest algorithm with topographic parameters

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
geo24/100

tRForest: a novel random forest-based algorithm for tRNA-derived fragment target prediction

GEO Series GSE189510. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2022View details →
geo24/100

A Random-Forest Based Algorithm for Prediction of Enhancers From Histone Modifications

GEO Series GSE37858. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMay 2012View details →
zenodo24/100

Data for: "Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study"

<p>Data and software&nbsp;related to the manuscript &quot;Random forest algorithms for recognizing daily life activities using plantar pressure information: A smart-shoe study&quot; (2020).&nbsp;</p> <p>========</p> <p>Typo in README_DataRep.txt:</p> <p>&#39; 2) &quot;one_config&quot; [...] Selected results requiring these data are shown in Fig. <strong>13</strong>&#39;<strong>&nbsp;</strong>(not 11).</p>

opencc-by-4.0Sep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record