Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
278
datasets available to search
ShareScore release 0.9.0
Dataset results
278 results for “Validated dataset”
Dataset associated to the manuscript "A novel method for characterising the inter- and intra-lake variability of CH4 emissions: validation and application across a latitudinal transect in the Alpine region"
<p>The dataset comprises data collected from nine specifically chosen lakes located across the eastern Alps during the ALCH4 project. The primary objective of this project was to assess the efficacy of a mobile Eddy Covariance platform in characterizing CH4 emissions. The field measurements were conducted throughout the ice-free periods in the years 2018 and 2019. This dataset encompasses: i) the EddyPro 7.0.6 (LI-COR Inc., Lincoln, NE, USA) output for all Eddy Covariance campaigns; ii) outcomes of the chamber measurements; iii) data obtained from surface samples. This research was funded by the Autonomous Province of Bozen/Bolzano.</p>
Data from: Validation of an algorithm for identifying MS cases in administrative health claims datasets
Open the record for dataset details and reuse information.
iValiD-TB: A fully characterized Mycobacterium tuberculosis dataset for antimicrobial resistance bioinformatics workflow validations
Open the record for dataset details and reuse information.
Data from: Establishing macroecological trait datasets: digitalization, extrapolation, and validation of diet preferences in terrestrial mammals worldwide
Open the record for dataset details and reuse information.
Data from: An updated global dataset for diet preferences in terrestrial mammals: testing the validity of extrapolation
Open the record for dataset details and reuse information.
Datasets from: Validated removal of nuclear pseudogenes and sequencing artefacts from mitochondrial metabarcode
Open the record for dataset details and reuse information.
Supplementary material 1 from: Rowley JJL, Callaghan CT (2020) The FrogID dataset: expert-validated occurrence records of Australia's frogs collected by citizen scientists. ZooKeys 912: 139-151. https://doi.org/10.3897/zookeys.912.38253
: Data type: Species data
Figure 2 from: Rowley JJL, Callaghan CT (2020) The FrogID dataset: expert-validated occurrence records of Australia's frogs collected by citizen scientists. ZooKeys 912: 139-151. https://doi.org/10.3897/zookeys.912.38253
Figure 2 Frequency histogram for the 172 species published in our openly accessible dataset, showing the number of records (on a log-scale) and how many species have that associated number of records.
Figure 1 from: Rowley JJL, Callaghan CT (2020) The FrogID dataset: expert-validated occurrence records of Australia's frogs collected by citizen scientists. ZooKeys 912: 139-151. https://doi.org/10.3897/zookeys.912.38253
Figure 1 Photographs of the top six species recorded in the first year FrogID. 1Crinia signifera2Limnodynastes peronii3Litoria peronii4Litoria fallax5Limnodynastes tasmaniensis6Litoria ewingii.
Reproducible Validation and Replication Studies in Nanoscale Physics (problem datasets for validation and replications from Ellis et al., 2016)
<p>Problem folders including all the input files necessary to reproduce the computations of the results related to Validation and replication of Ellis et al. 2016, on the paper: Reproducible Validation and Replication Studies in Nanoscale Physics</p>
Validation results of Non-parametric models of pH Neutralization plant Dataset
<p>This repository contains a table with different validation results of a series of models tested in pH Neutralization plant and Clarke plant.</p> <p>In this case, the naming protocol was as follows: [plant name]_DataSet_[model variable].</p> <p>The content of every CSV file is presented as follows: </p> <p>Column 1: { name: Model_No, type: numeric integer, limit: None, description: Code to identifiy the model (Primary Key)}</p> <p>Column 2: { name: KS, type: numeric float, limits: [0 1], description: Kolgomorov-Smirnov test}</p> <p>Column 3: { name: AD, type: numeric float, limits: [0 1], description: Anderson-Darling test}</p> <p>Column 4: { name: SW, type: numeric float, limits: [0 1], description: Shapiro-Wilk test}</p> <p>Column 5: { name: WX, type: numeric float, limits: [0 1], description: Wilcoxon test}</p> <p>Column 6: { name: FIT, type: numeric float, limits: [-Inf 1], description: Goodness of Fit metric}</p> <p>Column 7: { name: TIC, type: numeric float, limits: [0 1], description: Theil Inequality Index}</p> <p>Column 8: { name: Willmott, type: numeric float, limits: [0 1], description: Willmott metric}</p> <p>Column 9: { name: Russell_Pr, type: numeric float, limits: [0 1], description: Phase of Russell }</p> <p>Column 10: { name: Russell_Mr, type: numeric float, limits: [0 Inf], description: Magnitude of Russell }</p> <p>Column 11: { name: SG_Mr, type: numeric float, limits: [-Inf 1], description: Sprague & Geers}</p> <p>Column 12: { name: Anova, type: numeric float, limits: [0 1], description: Anova test}</p> <p>Column 13: { name: Dvure, type: numeric float, limits: [0 100], description: Dvurecenska metric}</p> <p>Column 14: { name: DTW, type: numeric float, limits: [0 Inf], description: DTW metric}</p> <p>Column 15: { name: D_DTW, type: numeric float, limits: [0 Inf], description: derivative of DTW metric}</p>
Dataset for "Validation of the Astro dataset clustering solutions with external data"
<p><strong>Validation data for the <em>Astro</em> scientific publication clustering benchmark dataset</strong></p> <p>This is the dataset used in the publication Donner, P. "Validation of the Astro dataset clustering solutions with external data", Scientometrics, DOI 10.1007/s11192-020-03780-3</p> <p>Certain data included herein are derived from Clarivate Web of Science. © Copyright Clarivate 2020. All rights reserved.</p> <p>Published with permission from Clarivate.</p> <p>The original <em>Astro</em> dataset is not contained in this data. It can be obtained from http://topic-challenge.info/ and requires permission from Clarivate Analytics for use.</p> <p> </p> <p>This dataset collection consists of four files. Each file contains an independent dataset that relates to the Astro dataset via Web of Science (WoS) record identifiers. These identifiers are called UTs. All files are tabular data in CSV format. In each, at least one column contains UT data. This should be used to link to the <em>Astro </em>dataset or other WoS data. The datasets are discussed in detail in the journal publication.</p> <p> </p>
Dataset on the validation of individuals' intention to engagement in philanthropic activities measure
Open the record for dataset details and reuse information.
Validation of a German translation of the Compassionate Engagement and Action Scales and Sussex-Oxford Compassion Scales in a general population sample: Dataset
Open the record for dataset details and reuse information.
Dataset for Cognitive and Relational Experience in Online Platforms: Scale Development and Validation (PLX)
Open the record for dataset details and reuse information.
Datasets for validating scMulan
Open the record for dataset details and reuse information.
Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection
<p>In the last decade, High-Throughput Sequencing (HTS) has revolutionized biology and medicine. This technology allows the sequencing of huge amount of DNA and RNA fragments at a very low price. In medicine, HTS tests for disease diagnostics are already brought into routine practice. However, the adoption in plant health diagnostics is still limited. One of the main bottlenecks is the lack of expertise and consensus on the standardization of the data analysis. The Plant Health Bioinformatic Network (PHBN) is an Euphresco project aiming to build a community network of bioinformaticians/computational biologists working in plant health. One of the main goals of the project is to develop reference datasets that can be used for validation of bioinformatics pipelines and for standardization purposes.</p> <p>Semi-artificial datasets have been created for this purpose (Datasets 1 to 10). They are composed of a "real" HTS dataset spiked with artificial viral reads. It will allow researchers to adjust their pipeline/parameters as good as possible to approximate the actual viral composition of the semi-artificial datasets. Each semi-artificial dataset allows to test one or several limitations that could prevent virus detection or a correct virus identification from HTS data (<i>i.e.</i> low viral concentration, new viral species, non-complete genome).</p> <p>Eight artificial datasets only composed of viral reads (no background data) have also been created (Datasets 11 to 18). Each dataset consists of a mix of several isolates from the same viral species showing different frequencies. The viral species were selected to be as divergent as possible. These datasets can be used to test haplotype reconstruction software, the goal being to reconstruct all the isolates present in a dataset.</p> <p><span>A GitLab repository (<a href="https://gitlab.com/ilvo/VIROMOCKchallenge">https://gitlab.com/ilvo/VIROMOCKchallenge</a>) is available and provides a complete description of the composition of each dataset, the methods used to create them and their goals.</span></p>
MolClassifier Training and Validation Datasets
<div> <div>The dataset contains 18626 chemical images (15720 for training and 2906 for validation) with annotated classes: `Molecular Structure`, `Markush Structure` and `Background`. <div> <div>Selected chemical images are randomly selected from the outputs of a segmentation module applied to documents from the United States Patent and Trademark Office.</div> <div>This dataset is part of <a href="https://github.com/DS4SD/PatCID">PatCID: an open-access dataset of chemical structures in patent documents</a>.</div> </div> </div> </div>
Magnetic Resonance Fingerprinting DICOM Validation Datasets for Real-Time Automated Quality Control for Quantitative MRI
<p>12 DICOM validation datasets (in DICOMDIR .zip format) from two Magnetic Resonance Fingerprinting sequences at two slice thickness across three test-retest sets per configuration. Also included are the resulting vial extraction reports and data for each dataset. </p>
Dataset accompanying: "Applying and Validating Coulomb Rate-and-State Seismicity Models in Acoustic Emission Experiments")
<p>Dataset accompanying: "Applying and Validating Coulomb Rate-and-State Seismicity Models in Acoustic Emission Experiments" by Heimisson, Naderloon, Chandra and Barnhoorn.</p> <p>Please reference the following publication if the data is used:<br><br>Heimisson, E.R., Naderloo, M., Chandra, D. and Barnhoorn, A., 2024. Applying and validating Coulomb rate-and-state seismicity models in acoustic emission experiments. <em>Tectonophysics</em>, p.230574. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.tecto.2024.230574" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.tecto.2024.230574 </span></span></a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.