Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,347
datasets available to search
ShareScore release 0.9.0
Dataset results
4,347 results for “tool”
Short example of Biceps Brachii muscle surface HDEMG decomposition using the DEMUSE Tool
<p>This dataset contains 4 examples of synthetic high density surface EMG signals of the Biceps Brachii muscle and results of their decomposition into separate motor unit activity. It is intended as a demonstration of the DEMUSE Tool software for sEMG decomposition and as a basis for practical example of dataset preparation for the HybridNeuro project webinar on Data management and ethics (<a href="https://www.hybridneuro.feri.um.si/results.html#webinars">https://www.hybridneuro.feri.um.si/results.html#webinars</a>). Two sets of data are included: the raw simulated sEMG signals and the results of decomposition of those signals with the DEMUSE Tool.</p>
Application of eDNA as a tool for assessing fish population abundance, Northern Wisconsin, US, 2017 - 2018
Environmental DNA concentrations, WDNR/GLIFWC mark-recapture population estimates, and abiotic lake data on 24 lakes in Wisconsin's Ceded Territory used to evaluate the relationship between walleye abundance and environmental DNA density and its application as a fisheries management tool.
Hand-selective visual regions represent how to grasp 3D tools for use: brain decoding during real actions
Open the record for dataset details and reuse information.
MiRoR15-P1-Tools used to assess the quality of peer review reports: a methodological systematic review
<p>Database, data extraction form, R codes and protocol related to: Superchi C, González JA, Solà I, Cobo E, Hren D, Boutron I. <em>Tools used to assess the quality of peer review reports: a methodological systematic review</em>. BMC Med Res Methodol. 2019;19(48):1–14. DOI: <a href="https://doi.org/10.1186/s12874-019-0688-x">https://doi.org/10.1186/s12874-019-0688-x</a></p> <p> </p>
MethylDetectR - A Translational Tool for Methylation-Based Health Profiling
<p><strong>** CORRECTION (2025-05-29): Please note that the script <a href="https://zenodo.org/api/records/15548022/draft/files/Script_For_User_To_Generate_Scores.R/content" target="_blank" rel="noopener noreferrer">Script_For_User_To_Generate_Scores.R</a> had an error whereby single-CpG EpiScores were producing the same score for all samples. This has been corrected in the newest version. </strong><br><br>This dataset includes reproducible code for the two applications related to the 'MethylDetectR' software. These .R files are included as 'MethylDetectR - Calculate Your Scores.R' and 'MethylDetectR.R'. An example DNAm file and SexAgeinfo file for upload to 'MethylDetectR - Calculate Your Scores' are included. These are 'DNAm_File_Example.rds' and 'SexAgeinfo_example.csv' respectively. An example output file from this application/for upload to 'MethylDetectR' is included as 'MethylDetectR - Test For Upload.csv'. An example and optional input file for case/control data is also available as 'MethylDetectR_Case_Control_Example.csv'. </p> <p>Furthermore, a script for the user to generate their own DNAm-based estimated values for human traits is included as 'Script_For_User_To_Generate_Scores.R'. A necessary associated file as 'Predictors_Shiny_By_Groups.csv' is also present for the script to run. We have also included separate necessary files to generate the chronological age predictor from Bernabeu <em>et al.</em> in our most recent versions of MethylDetectR. </p> <p>Lastly, an additional file called 'Truncate_to_these_CpGs.csv' is available which allows users to subset their methylation file to those CpG sites used in the 'MethylDetectR - Calculate Your Scores' application. This may substantially reduce the size of the methylation file for upload as well as its upload time. </p>
Data associated with the following publication: Developing the Playground Play Value and Usability Audit Tool (PVUA): An Evaluation of Content Validity via an Expert Panel
<p>This data set contains the supporting data associated with the following publication:</p> <p>Morgenthaler, T., Loebach, J., Lynch, H., Pentland, D., Kottorp, A., & Schulze, C. (in press). Developing the Playground Play Value and Usability Audit Tool (PVUA): An Evaluation of Content Validity via an Expert Panel. Children, Youth and Environments. [DOI was not yet available when the data set was published]</p> <p>The data set includes the following files:</p> <ul> <li>read me file [contains all relevant information to understand and reuse this data set] </li> <li>13 additional files [for description, see read me file]</li> </ul> <p>For more information, please contact the lead researcher, Thomas Morgenthaler (tom.morgenthaler@gmail.com or 121101888@umail.ucc.ie)</p> <p> </p>
DESIRA - inventory of digital tools for agriculture, forestry, and rural areas
<p>Inventory of digital tools for agriculture, forestry, and rural areas collected by the DESIRA consortium.</p>
Dataset for a machine learning tool to improve lymph node staging with FDG-PET/CT
<p>This upload provides Open Data associated with the publication "A machine learning tool to improve prediction of mediastinal lymph node metastases in non-small cell lung cancer using routinely obtainable [<sup>18</sup>F]FDG-PET/CT parameters" by Rogasch JMM <em>et al.</em> (2022).</p> <p>The upload contains the anonymized dataset with 10 features necessary for the final GBM model that was presented in the publication. However, the original full dataset with 40 features was excluded from this Open Data repository because it may not comply with strict rules of data anonymization. The full dataset can be obtained from the corresponding author (julian.rogasch@charite.de) upon reasonable request.</p> <p>Besides the dataset, this upload provides the original python and R scripts that were used as well as their output.</p> <p>A description of all files can be found in "content_description_2022_11_19.txt".</p> <p>A user-friendly web tool that implements the final machine learning model can be found here: <a href="https://baumgagl.github.io/PET_LN_calculator/">PET_LN_calculator</a> </p>
Introductory Motus Prioritization Tool Data (Open first)
Addressing survival and movement of priority migratory avian species of concern along the Pacific Flyway is paramount for their conservation. Yet, the migratory life stage is understudied in many avian species. The Motus radiotelemetry receiver network is an established system for tracking survival and movement of avian species. This network is an international collaborative that successfully identifies stopover site duration, connected migratory routes, post-fledging dispersal and survival, and adult survival and fidelity on a landscape-scale; parameters that cannot be easily estimated using non-tagged birds. While the Motus network is highly connected in eastern North America, the western part of the continent is lagging in coverage and connectivity, limiting the ability to obtain sample sizes large enough to robustly model demographic parameters from tagged birds. Thus, the expansion of the Motus network is a high priority for Pacific Flyway State Agencies. To date, no method exists for determining priority locations for new Motus receiving stations. With collaborations from States and the Canadian Province of British Columbia, we used eBird citizen scientist data to prioritize strategic locations for new Motus receiving stations throughout the Pacific Flyway. We model priority species’ co-occupancy of varying abundance states (i.e., absent, present, abundant, abundant in multiple weeks) with spatially varying Landsat (red and near infrared), water, land cover types, and weather covariates while accounting for variable detection with temporally varying survey effort covariates. Using occupancy model predictions, we identify high-use areas of the Pacific Flyway for establishing new Motus receiving towers that have high probabilities of intercepting high presence and /or abundance of multiple species of interest in a series of predictive occupancy maps. This package contains required files to recreate the data analysis, print out maps based on predictions from the
Software Tools to Collect and Use Provenance in R
The software tools that scientists use to process and analyze data are typically optimized for performance and ease of use. Few if any such tools are designed to capture and record the details of what happens as the tool performs its task. This detailed information, and more generally the history of an item of data from its creation to its present state, is known as provenance. Provenance has the potential to make science more transparent, reliable, and reproducible. This project focused on collecting and using provenance for scripts written in the R statistical language, which is widely used by ecologists and environmental scientists for data analysis and visualization. Our tools include a provenance collector (rdtLite), which collects provenance as an R script executes (or during a console session), as well as other tools that use the collected provenance to document and visualize the execution or to support activites such as script debugging. The R packages included here are also available on CRAN. For more details, see the project website on GitHub (https://end-to-end-provenance.github.io).
Performance of users with Cerebral Palsy playing GABLE Games together with their results to the Left/Right Dynamic balance tool
<p>This dataset contains data generated by users of GABLE platform. The data shows the performance of some users with Cerebral Palsy playing GABLE Games together with their results to the Left/Right Dynamic balance tool. More information about GABLE project can be found at: www.projectgable.eu</p>
MiRoR15-P2-Development of ARCADIA: a tool for assessing the quality of peer-review reports in biomedical research
<p>Survey questionnaire, anonymised survey data, and codebook related to: Superchi C, Hren D, Blanco D, Rius R, Recchioni A, Boutron I, González JA. Development of ARCADIA: a tool for assessing the quality of peer-review reports in biomedical research. BMJ Open 2020;0:e035604. doi:10.1136/bmjopen-2019-035604</p>
Defect Prediction Tool Validation Dataset 2
<p><strong>This dataset is used to address the Research Questions in the study at Transactions on Software Engineering</strong>: <strong>Within-Project</strong> <strong>Defect Prediction of Infrastructure-as-Code using Product and Process Metrics. </strong></p> <p><strong>See also: https://github.com/stefanodallapalma/TSE-2020-05-0217.</strong></p> <p>It provides</p> <p>* <strong>repositories.json</strong> - a list of repositories selected from open-source GitHub repositories based on the Ansible language.</p> <p>* <strong>fixing-commits.json</strong> - a list of defect-fixing commits extracted from those repositories.</p> <p>* <strong>fixed-files.json</strong> - a list of Ansible files fixed in those defect-fixing commits and respective bug-inducing commits.</p> <p>* <strong>failure-prone-files.json</strong> - a list of failure-prone files through the repository's commit history.</p> <p>* <strong>metrics.zip </strong>- csv files consisting of releases (set of files) and their IaC-oriented, delta and process metrics extracted from each analyzed repository</p> <p>* <strong>projects.zip </strong>- for each analyzed project, it contains the data (models, performance, and results of Recursive Feature Elimination) used to answer the Research Questions.</p> <p><strong>Context</strong></p> <p><em>Infrastructure-as-code (IaC)</em> is the DevOps strategy that allows management and provisioning of infrastructure through the definition of machine-readable files and automation around them, rather than physical hardware configuration or interactive configuration tools.</p> <p>On the one hand, although IaC represents an ever-increasing widely adopted practice nowadays, still little is known concerning how to best maintain, speedily evolve, and continuously improve the code behind the IaC strategy in a measurable fashion. <br> On the other hand, source code measurements are often computed and analyzed to evaluate the different quality aspects of the software developed.<br> In particular, Infrastructure-as-Code is simply "code", as such it is prone to defects as any other programming languages.</p> <p>This dataset targets the YAML-based Ansible language to devise <strong>within-project defects prediction</strong> approaches for IaC based on Machine-learning.</p> <p><strong>Content</strong></p> <p>The dataset contains metrics extracted from 85 open-source GitHub repositories based on the Ansible language that satisfied the following criteria:</p> <p>* The repository has at least one push event to its master branch in the last six months;<br> * The repository has at least 2 releases;<br> * At least 10% of the files in the repository are IaC scripts;<br> * The repository has at least 2 core contributors;<br> * The repository has evidence of continuous integration practice, such as the presence of a .travis.yaml file;<br> * The repository has a comments ratio of at least 0.1%;<br> * The repository has commit frequency of at least 2 per month on average;<br> * The repository has an issue frequency of at least 0.01 events per month on average;<br> * The repository has evidence of a license, such as the presence of a LICENSE.md file<br> * The repository has at least 100 source lines of code.</p> <p>Metrics are grouped into three categories:</p> <p>* <strong>IaC-Oriented:</strong> metrics of structural properties derived from the source code of infrastructure scripts. Click [here](https://www.sciencedirect.com/science/article/pii/S0164121220301618) for more info.</p> <p>* <strong>Delta</strong>: metrics that capture the amount of change in a file between two successive releases, collected for each IaC-oriented metric.</p> <p>* <strong>Process</strong>: metrics that capture aspects of the development process rather than aspects about the code itself. Description of the process metrics in this dataset can be found [here](https://pydriller.readthedocs.io/en/latest/processmetrics.html).</p> <p>In addition to the metrics, the dataset contains the pre-trained models (*.joblib) in the folders rq1 and rq2 of projects.zip.</p> <p>You can load the model in Python as follows:</p> <p>```<br> from joblib import load<br> model = load('projects/owner/repository/rq1/random_forest.joblib'), mmap_mode='r')</p> <p>best_estimator = model['estimator'] # The estimator that maximized the AUC-PR</p> <p>cv_results = model['cv_results'] # The results of each step of the validation procedure</p> <p>best_index = mode['best_index'] # The index to access the best cv_results<br> ```</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p> </p> <p>This work is supported by the European Commission grants no. 825040 (RADON H2020).</p> <p><br> <strong>Inspiration</strong></p> <p>What source code properties and properties about the development process are good predictors of defects in Infrastructure-as-Code scripts?</p>
IPBES Data Management Tutorials - Session 5.2: Tools to find and attribute DOIs
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>The session on <em>tools to find and attribute DOIs </em>covers fundamental background information on digital object identifiers and how to resolve and reserve them.</p>
An integrated polygenic tool substantially enhances coronary artery disease prediction
<p>Summary-level CAD GWAS data generated by Genomics plc as presented in:</p> <p>Riveros-Mckay F. et al. An integrated polygenic tool substantially enhances coronary artery disease prediction. Circulation: Genomics and Precision Medicine (in press). </p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at research@genomicsplc.com</p> <p> </p> <p>NOTES<br> -----------------------------<br> These analyses were carried out using the full UK Biobank imputation data release (v3b). Analyses were restricted to a subset of UK Biobank, described as “Group I” in the published paper. Group I, “no PCE/QRISK3 available”, included 114,196 European-ancestry individuals with missing data that prevented PCE or QRISK3 calculation.</p> <p>CAD case phenotypes were defined as described in the “Phenotype definitions” section of the paper’s Supplementary Materials, using both prevalent (pre-baseline) and incident (post-baseline) events.</p> <p>All analyses included Age at assessment, sex, genotyping chip, and 10 principal components as covariates. </p> <p>We used plink2.0 logistic regression. For chromosome X variants males were treated as having 0 or 2 alternative alleles. </p> <p>The results are not adjusted for genomic control.</p> <p> </p> <p>DATA FILE CONTENT DESCRIPTION<br> -----------------------------<br> cpra Variant ID in ‘CPRA’ format. Position reflects position in b37. <br> chrom Chromosome<br> pos Position in base pairs (b37, 1-based)<br> alt Alternative allele (effect allele)<br> beta Effect size (log odds ratio)<br> standard_error Standard error of beta <br> minus_log10_p Minus log(base 10) of P-value<br> ref Reference allele (non-effect allele)<br> ncase Number of cases<br> ncontrol Number of controls</p>
IPBES Data Management Tutorials - Session 5.3: Literature access tools
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>This session on literature access tools introduces Research4Life, a tool which provides experts in middle to low income countries access to scientific and grey literature.</p>
Multi-faceted analyses of Poland's Bronze and Early Iron Age hoards: Fig.5. Pottery (A, C), animal bones (B), a human skull (C, D), and a flint tool (D) excavated from underneath the stone layer in Kaliszany (archaeological site no. 3)
<p>The set contains a figure, with with photographs that show examples of finds discovered during excavations at archaeological site 3 in Kaliszany, Wągrowiec commune, Poland. It is a stone and earth structure in which a hoard of metal objects dating to the Late Bronze Age was discovered in 1943. The photo is from the 2022 survey, when the south-western part of the structure was explored. <br><br>The paper and data were prepared as part of a project funded by the National Science Centre, Poland: <em>A Biography of Late Bronze and Early Iron Ages Hoards. A Multi-Faceted Analysis of Metal Objects Related to Monumental Constructions in Poland</em> (UMO-2021/41/B/HS3/00038)</p>
Grapegenomics.com: a web portal with genomic data and analysis tools for wild and cultivated grapevines
<p><a href="https://grapegenomics.com">Grapegenomics.com</a> is a web portal that provides public access to genome references for grapevine cultivars (<em>Vitis vinifera</em> ssp. <em>vinifera</em>), wild grapevines (<em>Vitis vinifera</em> ssp. <em>sylvestris</em>), various wild grape species (<em>Vitis</em> spp. and <em>Muscadinia</em> spp.), and major fungal pathogens affecting grapes.</p> <p>All genomes are accessible through dedicated genome browsers, and published genomes are available for complete <a href="https://www.grapegenomics.com/download.php">download</a>.</p> <p>The site hosts all genomes produced by the laboratory of Dario Cantù in the Department of Viticulture and Enology at the University of California, Davis, along with published genome references generated by others, such as PN40024 and Pinot noir ENTAV115. Instructions for genome submission are provided <a href="https://www.grapegenomics.com/submit.php">here</a>. The portal is maintained by Noé Cochetel (ndcochetel[at]ucdavis.edu). In this version 2.0, all genome browsers utilize <a href="https://jbrowse.org/jb2/">jbrowse 2</a>. <br><br>Link to the website: <a href="https://www.grapegenomics.com">https://www.grapegenomics.com</a> </p>
Dataset of "Liquid-Jet Photoemission Spectroscopy as a Structural Tool: Site-Specific Acid-Base Chemistry of Vitamin C"
<p>Liquid-jet photoemission spectroscopy (LJ-PES) directly probes the electronic structure of solutes<br>and solvents. It also emerges as a novel tool to explore chemical structure in aqueous solutions, yet<br>the scope of the approach has to be examined. Here, we present a pH-dependent liquid-jet photoelectron<br>spectroscopic investigation of ascorbic acid (vitamin C). We combine core-level photoelectron<br>spectroscopy and ab initio calculations, allowing us to site-specifically explore the acid-base chemistry<br>of the biomolecule. For the first time, we demonstrate the capability of the method to simultaneously<br>assign two deprotonation sites within the molecule. We show that a large change in chemical shift<br>appears even for atoms distant several bonds from the chemically modified group. Furthermore, we<br>present a highly efficient and accurate computational protocol based on a single structure using the<br>maximum overlap method for modeling core-level photoelectron spectra in aqueous environments.<br>This work poses a broader question: To what extent can LJ-PES complement established structural<br>techniques such as nuclear magnetic resonance? Answering this question is highly relevant in view<br>of the large number of incorrect molecular structures published.</p>
A uniaxial hysteretic superelastic constitutive model applied to additive manufactured lattices - data and postprocessing tools
<p>This data set contains all result data obtained during the implementation of an uniaxial hysteretic superelastic constitutive model and its application to additive manufactured lattices.</p> <p>Furthermore, it contains all ABAQUS .inp files, the implemented subroutine of the hysteretic superelastic constitutive model, diagrams generated from the data, as well as postprocessing tools for generating the diagrams.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.