Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

43

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

43 results for “decision trees”

Learn how ShareScore rates datasets ↗
zenodo44/100

Image Repository Decision Tree - Where do I deposit my imaging data

<p>Depositing data in quality data repositories is one crucial step towards FAIR (Findable, Accessible, Interoperable, and Reusable) data. Accordingly, Euro-BioImaging strongly encourages sharing scientific imaging data in established, thematic repositories.&nbsp;</p> <p>To guide you in the selection of appropriate repositories, we have created an overview of available repositories for different types of image data, including their scope and requirements. This decision tree guides you through questions about your data and directs you to the correct repository, and/or provides instructions for further processing to meet the critera of the repositories.&nbsp;</p> <p>Three seperate trees are provided for different classes of imaging data: open bioimage data, preclinical data, and human imaging data. These versions with three trees can be used for web-view. Update: also the editable versions in powerpoint format (.pptx) are now provided. Please be aware that opening the versions with another program might lead to shifted formatting.</p> <p>Update: we now also provide ready-to-print versions designed to be printed on A3 format. One page shows the open bioimaging data tree and one page combines the preclinical and human imaging data trees. Also the editable versions of these are provided.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Full decision tree of aquaZone

<p>The figure shows the full decision tree of aquaZone, which includes eight out of 30 environmental criteria.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Source Data for Manuscript: "Retrievals Applied To A Decision Tree Framework Can Characterize Earth-like Exoplanet Analogs"

<p>This dataset accompanies the manuscript entitled: "Retrievals Applied To A Decision Tree Framework Can Characterize Earth-like Exoplanet Analogs", which was accepted for publication in the Planetary Science Journal. Included are the source files for all figures included in the paper.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Advancements in QSAR modelling: Decision trees and rotation forest for prediction of Aspergillus anti-inflammatory metabolites

<p>This study presents applications of advancements in QSAR modelling for predicting nitric oxide (NO) inhibitors and anti-inflammatory metabolites from the <em>Aspergillus</em> genus. Inflammation-related diseases remain a pressing concern, necessitating the identification of effective anti-inflammatory compounds. Using decision trees, the Ranker method, and CorrelationAtrributeEval as a base classifier for attribute selection together with Rotation Forest and Adaboost as enhancers, we explored their potential with different classifiers including Artificial Neural Networks and J48 Trees. The proposed QSAR models employed an ensemble approach with Rotation Forest and Adaboost.M1, applying an automated KNIME workflow. Seven molecular descriptors were selected and trained on a comprehensive dataset of diverse anti-inflammatory <em>Aspergillus</em> specialised metabolites. Results showed that the Rotation Forest-enhanced version outperformed other models, capturing complex structure-activity relationships and improving predictive performance. Chemical characteristics of electrotopological state, topological distances, and functional groups including secondary amides and alcohols contribute to important anti-inflammatory effects. The developed QSAR model showed good predictive performance for anti-inflammatory <em>Aspergillus</em> metabolites, focusing on their NO inhibitory activity. These results can contribute to the discovery of novel anti-inflammatory drugs based on computational techniques.</p> <p>&nbsp;</p>

opencc-bySep 2023View details →
zenodo36/100

OD-1 Open Data Decision Tree Version 0.4

<p>This is a prototype/proof of concept upload to Zenodo to test and demonstrate the concept of working in a 'Commons Compliant' way. This involves conducting all stages of research and research communication openly (Open by Design and Open by Default). </p> <p>Decision trees record a set of actions that can be taken with scholarly objects to make them as Open, FAIR and Citable as possible. They represent the most practical framework for putting the Principles of the Scholarly Commons into practice.</p> <p>This tree was uploaded using a Draw.io - Dropbox - Zenodo workflow.</p> <p> </p>

opencc-by-4.0May 2017View details →
zenodo36/100

Digging for Decision Trees: A Case Study in Strategy Sampling and Learning (experimental reproduction package)

<p>This artifact permits to reproduce the experimental results obtained with the techniques and algorithms presented in the article "Digging for Decision Trees: A Case Study in Strategy Sampling and Learning" by Carlos E. Budde, Pedro R. D'Argenio, and Arnd Hartmanns (2024).</p> <p>The contents include all data and software (formal models, software tools, Python &amp; bash scripts) used in the experimental evaluation presented in &sect;7 of the article. Detailed instructions on how to reproduce the results are bundled in the artifact. Execution has been tested in the Virtual Machine available at https://zenodo.org/records/7113223.</p> <p>&nbsp;</p>

opengpl-3.0-or-laterAug 2024View details →
zenodo36/100

Training Data for Decision Tree Practical in reading-ml-chemistry repo

<p>This is a dataset that can be used to train a decision tree to predict the band gap of a material. The data is associated with a notebook for running the practical and it can be found at https://github.com/keeeto/reading-ml-chemistry.</p> <p>&nbsp;</p> <p>The data originally comes from the Materials Project.</p> <p>&nbsp;</p> <p>A new muon spectroscopy dataset is added. It is from - Machine learning approach to muon spectroscopy analysis - https://iopscience.iop.org/article/10.1088/1361-648X/abe39e/meta</p>

opencc-by-4.0Jan 2021View details →
dryad36/100

Linking animal behaviour and tree recruitment: Caching decisions by a scatter hoarder corvid determine seed fate in a Mediterranean agroforestry system

<p><span>1. Seed dispersal by scatter-hoarder corvids is key for the establishment of important tree species from the Holarctic region such as the walnut (<em>Juglans regia</em>). However, the factors that drive animal decisions to cache seeds in specific locations and the consequences of these decisions on seed fate are poorly understood. </span></p> <p><span>2. We experimentally created four distinct, replicated habitat types in a Mediterranean agricultural landscape where the Eurasian magpie (<em>Pica</em> <em>pica</em>) is a common scatter-hoarder: soft bare soil; compacted bare soil; compacted soil with a dense herbaceous cover; and soft linear bare soil made up of the irrigation furrows that separated the rest of the treatments. We also experimentally placed visual landmarks (stones, sticks, and bunches of dry plants) to test if magpies use them to place seed caches. Walnut dispersal from feeders to the habitats was monitored by radio-tracking and camera traps. </span></p> <p><span>3. A sowing experiment simulating natural caches tested the effect of caching type on seed germination and seedling emergence. Seed mass was controlled for the dispersal and sowing experiments.</span></p> <p><span>4. Magpies selected the two habitats with soft soil, and avoided the one with compacted soil, to cache nuts. Seed mass did not affect dispersal distance, germination, or emergence; however, heavier seeds were cached more often under litter and in the habitat with herbaceous cover, whereas lighter seeds were more often buried in the soft bare soil habitat. Seed burial under soil or litter determined seed fate, as there was virtually no emergence from unburied nuts. There was no evidence of any effect of the visual landmarks.</span></p> <p>5. Synthesis. The consequences of seed caching for seedling early establishment are driven by a fine decision-making process of the disperser. Magpies seemed to ponder the characteristics of the habitat and the seed itself to determine where and how to cache each nut. By doing so, magpies reinforced the quality of seed dispersal effectiveness, as they cached walnuts in locations that enhanced both seed survival and seedling emergence.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Research Data Management Decision Tree

<p>Researchers often face the same stumbling blocks. To support them, we have developed a <strong>RDM Decisional Tree</strong> starting from the fundamental bricks of the data lifecycle and posing a series of questions to help researchers navigate:</p> <p>1) the domain specific nature and origin of the data they are handling;</p> <p>2) Privacy/Ethics requirements (e.g. GDPR);</p> <p>3) Intellectual Property Rights;</p> <p>4) active data storage;</p> <p>5) long-term deposit and preservation.</p> <p><span>This diagram has been created by the team of data steward at the University of Bologna (UNIBO, Alma Mater Studiorum - Universit&agrave; di Bologna) in October 2022.&nbsp;</span>It has been developed in parallel to the Research Data Management: Data Lifecycle, available here:&nbsp;<a href="https://doi.org/10.5281/zenodo.7249050">10.5281/zenodo.7249050</a>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Statistical Decision Trees for Marine Ecological Data

<p>Both decision trees can be used to choose an appropriate statistical analysis when dealing with ecological (marine) data. For each analysis, an example publication using/explaining the method, some of the R commands required and a tutorial are listed. Please note that (1) these statistical analyses should be seen as suggestions and other approaches are available, (2) papers listed are only provided as examples but appropriate referencing of the method is expected and (3) some analyses have specific assumptions that the users should check before analysing their data.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Clustering Tasks and Decision Trees with Elegiac Poets

<p>The dataset contains files generated during a Natural Language Processing (NLP) and automatic text analysis task. Attached is a <strong>Jupyter notebook</strong> with the complete code, along with several <strong>Excel files (.xlsx)</strong> containing organized information. Additionally, there are three folders that include files generated during the Silhouette calculation, K-means clustering, and feature extraction using decision trees.</p> <p>The three folders are:<br>1. <strong>Silhouette Calculation:</strong> Contains PNG images of Silhouette plots for various analysis configurations.<br>2.<strong> K-means Clustering: </strong>Contains pickle (.pkl) files with features and labels for each combination of excluded author, n-gram type, n-gram range, and matrix type.<br>3. <strong>Feature Extraction: </strong>Contains CSV files with lists of documents by cluster and the most important features along with information gain and information gain ratio metrics.</p> <p>Other file formats included in the dataset are:<br>- CSV files containing Silhouette scores, optimal clustering results, cluster assignments, and optimal cluster assignments.<br>- PNG images of scatter plots colored by author and by cluster.<br>- Pickle files containing the top features extracted during the analysis.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Decision Tree Datasets

<p>Raw datasets of some of the monetary economic methods used in environmental economics and used here to build predictive models using the decision tree technique. The datasets contain some qualitative choice variables of monetary economic methods for evaluating the protective service provided by the forest against rockfalls.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

On Constructing Limits-of-Acceptability in Watershed Hydrology using Decision Trees

<p>This submission contains the hydrological data used in the study.</p> <p>Use the following code in python to access the data</p> <p>Load_data = pickle.load( open(filename,&#39;rb&#39;)) # filename should be specified along with full file path; pandas version 1.3.5&nbsp; might be required to open this file</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Dataset: Recruitment responses of shade-tolerant and heliophilous trees in human-degraded areas: the necessity of knowing the recruitment autecology of species for making reforestation decisions

<p>This repository contains the files associated with the following article:</p> <p>Johanna Croce, Ernesto I. Badano, Andr&eacute;s T&aacute;lamo. Recruitment responses of shade-tolerant and heliophilous trees in human-degraded areas: the necessity of knowing the recruitment autecology of species for making reforestation decisions Submitted to <em>Land Degradation &amp; Development</em>.</p> <p>The first Microsoft Excel file (Dataset 01 - Microclimate.xlsx) contains two sheets with the microclimatic data (average, maximum and minimum soil temperatures, and volumetric soil water content) measured in Cerro Chachapoyas and Cerro Fachacano. These measurements were performed on three plots of each trearment, including shrub-protected treatment with high shade, shrub-protected treatment with medium shade, sunny treatment with high herbaceous cover, sunny treatment with medium herbaceous cover and controls. The second Microsoft Excel file (Dataset 02 -Plant responses.xlsx Dataset 02 - Plant responses.xlsx) contains three sheets with the data used to estimate the seedling emergence rates, plant survival rates and net aboveground growth rates of the tree species, including <em>Anadenanthera colubrina</em> and <em>Ceiba chodatii</em> in Cerro Chachapoyas, and <em>Jacaranda mimosifolia</em> in Cerro Fachacano.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Data and code for "Des-q: a quantum algorithm to construct and efficiently retrain decision trees for regression and binary classification"

<p>It contains the code and the data to reproduce the figures in the paper&nbsp;&quot;Des-q: a quantum algorithm to construct and efficiently retrain decision trees for regression and binary classification&quot; published in arXiv:&nbsp;https://arxiv.org/abs/2309.09976</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

Linking animal behaviour and tree recruitment: Caching decisions by a scatter hoarder corvid determine seed fate in a Mediterranean agroforestry system

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad36/100

Understanding smallholder decision-making to increase farm tree diversity: Enablers and barriers for forest landscape restoration in Western Kenya

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad32/100

Data from: Learning to see the wood for the trees: machine learning, decision trees and the classification of isolated theropod teeth

Taxonomic identification of fossils based on morphometric data traditionally relies on the use of standard linear models to classify such data. Machine learning and decision trees offer powerful alternative approaches to this problem but are not widely used in palaeontology. Here, we apply these techniques to published morphometric data of isolated theropod teeth in order to explore their utility in tackling taxonomic problems. We chose two published datasets consisting of 886 teeth from 14 taxa and 3020 teeth from 17 taxa, respectively, each with five morphometric variables per tooth. We also explored the effects that missing data have on the final classification accuracy. Our results suggest that machine learning and decision trees yield superior classification results over a wide range of data permutations, with decision trees achieving accuracies of 96% in classifying test data in some cases. Missing data or attempts to generate synthetic data to overcome missing data seriously degrade all classifiers predictive accuracy. The results of our analyses also indicate that using ensemble classifiers combining different classification techniques and the examination of posterior probabilities is a useful aid in checking final class assignments. The application of such techniques to isolated theropod teeth demonstrate that simple morphometric data can be used to yield statistically robust taxonomic classifications and that lower classification accuracy is more likely to reflect preservational limitations of the data or poor application of the methods.

opencc-zeroSep 2020View details →
dryad32/100

Data from: Foraging decisions with conservation consequences: Interaction between beavers and invasive tree species

<p><span>Herbivore species can either hinder or accelerate the invasion of woody species through selective utilization. Therefore, an exploration of foraging decisions can contribute to the understanding and forecasting of woody plant invasions. Despite the large distribution range and rapidly growing abundance of beaver species across the Northern Hemisphere, only a few studies focus on the interaction between the beaver and invasive woody plants. </span></p> <p><span>We collected data on the woody plant supply and utilization at 20 study sites in Hungary, at two fixed distances from the water. The following parameters were registered: taxon, trunk diameter, type of utilization, and carving depth. Altogether 5401 units (trunks and thick branches) were identified individually. We developed a statistical protocol that uses a dual approach, combining whole-database and transect-level analyses to examine foraging strategy.</span></p> <p><span>Taxon, diameter, and distance from water all had a significant effect on foraging decisions. The order of preference for the four most abundant taxa was: <em>Populus </em>spp. (softwood), <em>Salix </em>spp. (softwood), <em>Fraxinus pennsylvanica</em></span> <span>(invasive hardwood),<em> Acer negundo</em> (invasive hardwood). The diameter influenced the type of utilization, as units with greater diameter were rather carved or debarked than felled. According to the central-place foraging strategy, intensity of the foraging decreased with the distance from the water, while both the taxon and diameter selectivity increased. This suggests stronger modification of the woody vegetation directly along the waterbank, together with a weaker impact further from the water. </span></p> <p><span>In contrast to invasive trees, for which utilization occurred almost exclusively in the smallest diameter class, even the largest softwood trees were utilized by means of carving and debarking. This may lead to the gradual loss of softwoods or the transformation of them into shrubby form. After the return of the beaver, mature stages of softwood stands and thus the structural heterogeneity of floodplain woody vegetation could be supported by the maintenance of sufficiently large active floodplains. </span></p> <p><span>The beaver accelerates the shift of the canopy layer's species composition towards invasive hardwood species, supporting the enemy release hypothesis. However, the long-term impact will also depend on how plants respond to different types of utilization and on their ability to regenerate, which are still unexplored issues in this environment. Our results should be integrated with knowledge about factors influencing the competitiveness of the studied native and invasive woody species to support floodplain conservation and reconstruction.</span></p>

opencc-zeroApr 2022View details →
zenodo32/100

Dataset and Additional Information for the paper A LINEAR-ALGEBRAIC MODEL FOR ESTIMATING ANTI-LEARNING WHEN A DECISION TREE SOLVES THE PARITY BIT PROBLEM, by ALEXEI LISITSA and ALEXEI VERNITSKI (submitted)

<p>This upload contains a dataset and additional information for the paper&nbsp;A LINEAR-ALGEBRAIC MODEL FOR ESTIMATING<br> ANTI-LEARNING WHEN A DECISION TREE SOLVES THE PARITY BIT PROBLEM, by ALEXEI LISITSA and &nbsp;ALEXEI VERNITSKI (submitted)&nbsp;</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record