Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

278

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

278 results for “Validated dataset”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Validation of an algorithm for identifying MS cases in administrative health claims datasets

Objective: To develop a valid algorithm for identifying multiple sclerosis (MS) cases in administrative health claims (AHC) datasets. Methods: We used 4 AHC datasets from the Veterans Administration (VA), Kaiser Permanente Southern California (KPSC), Manitoba (Canada), and Saskatchewan (Canada). In the VA, KPSC, and Manitoba, we tested the performance of candidate algorithms based on inpatient, outpatient, and disease-modifying therapy (DMT) claims compared to medical records review using sensitivity, specificity, positive and negative predictive values, and interrater reliability (Youden J statistic) both overall and stratified by sex and age. In Saskatchewan, we tested the algorithms in a cohort randomly selected from the general population. Results: The preferred algorithm required ≥3 MS-related claims from any combination of inpatient, outpatient, or DMT claims within a 1-year time period; a 2-year time period provided little gain in performance. Algorithms including DMT claims performed better than those that did not. Sensitivity (86.6%–96.0%), specificity (66.7%–99.0%), positive predictive value (95.4%–99.0%), and interrater reliability (Youden J = 0.60–0.92) were generally stable across datasets and across strata. Some variation in performance in the stratified analyses was observed but largely reflected changes in the composition of the strata. In Saskatchewan, the preferred algorithm had a sensitivity of 96%, specificity of 99%, positive predictive value of 99%, and negative predictive value of 96%. Conclusions: The performance of each algorithm was remarkably consistent across datasets. The preferred algorithm required ≥3 MS-related claims from any combination of inpatient, outpatient, or DMT use within 1 year. We recommend this algorithm as the standard AHC case definition for MS.

opencc-zeroDec 2018View details →
dryad32/100

Data from: An updated global dataset for diet preferences in terrestrial mammals: testing the validity of extrapolation

1. Diet is a key trait of an organism's life history that influences a broad spectrum of ecological and evolutionary processes. Kissling et al. (2014) compiled a species-specific dataset of diet preferences of mammals for 38% of a total of 5364 terrestrial mammalian species assessed for the International Union for Conservation of Nature's Red List, to facilitate future studies. The authors imputed dietary data for the remaining 62% by using extrapolation from phylogenetic relatives. 2. We collected dietary information for 1261 mammalian species for which data were extrapolated by Kissling et al. (2014), in order to evaluate the success with which such extrapolation can predict true diets. 3. The extrapolation method devised by Kissling et al. (2014) performed well for broad dietary categories (consumers of plants and animals). However, the method performed inconsistently, and sometimes poorly, for finer dietary categories, varying in accuracy in both dietary categories and mammalian orders. 4. The results of the extrapolation performance serve as a cautionary tale. Given the large variation in extrapolation performance, we recommend a more conservative approach for inferring mammalian diets, whereby dietary extrapolation is implemented only when there is a high degree of phylogenetic conservatism for dietary traits. Phylogenetic comparative methods can be used to detect and measure phylogenetic signal in diet. If data for species are needed, then only the broadest feeding categories should be used. This would ensure a greater level of accuracy and provide a more robust dataset for further ecological and evolutionary analysis.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Establishing macroecological trait datasets: digitalization, extrapolation, and validation of diet preferences in terrestrial mammals worldwide

Ecological trait data are essential for understanding the broad-scale distribution of biodiversity and its response to global change. For animals, diet represents a fundamental aspect of species' evolutionary adaptations, ecological and functional roles, and trophic interactions. However, the importance of diet for macroevolutionary and macroecological dynamics remains little explored, partly because of the lack of comprehensive trait datasets. We compiled and evaluated a comprehensive global dataset of diet preferences of mammals ("MammalDIET"). Diet information was digitized from two global and cladewide data sources and errors of data entry by multiple data recorders were assessed. We then developed a hierarchical extrapolation procedure to fill-in diet information for species with missing information. Missing data were extrapolated with information from other taxonomic levels (genus, other species within the same genus, or family) and this extrapolation was subsequently validated both internally (with a jack-knife approach applied to the compiled species-level diet data) and externally (using independent species-level diet information from a comprehensive continentwide data source). Finally, we grouped mammal species into trophic levels and dietary guilds, and their species richness as well as their proportion of total richness were mapped at a global scale for those diet categories with good validation results. The success rate of correctly digitizing data was 94%, indicating that the consistency in data entry among multiple recorders was high. Data sources provided species-level diet information for a total of 2033 species (38% of all 5364 terrestrial mammal species, based on the IUCN taxonomy). For the remaining 3331 species, diet information was mostly extrapolated from genus-level diet information (48% of all terrestrial mammal species), and only rarely from other species within the same genus (6%) or from family level (8%). Internal and external validation showed that: (1) extrapolations were most reliable for primary food items; (2) several diet categories ("Animal," "Mammal," "Invertebrate," "Plant," "Seed," "Fruit," and "Leaf") had high proportions of correctly predicted diet ranks; and (3) the potential of correctly extrapolating specific diet categories varied both within and among clades. Global maps of species richness and proportion showed congruence among trophic levels, but also substantial discrepancies between dietary guilds. MammalDIET provides a comprehensive, unique and freely available dataset on diet preferences for all terrestrial mammals worldwide. It enables broad-scale analyses for specific trophic levels and dietary guilds, and a first assessment of trait conservatism in mammalian diet preferences at a global scale. The digitalization, extrapolation and validation procedures could be transferable to other trait data and taxa.

opencc-zeroDec 2013View details →
zenodo32/100

Microphone Comparison Array Validation Dataset

<p>Data generated as part of research to determine an appropriate method to make audio recordings suitable for use in comparing the perceptual characteristics imparted by microphones. &nbsp;Data comprise audio files, listening test interfaces and MATLAB code.</p> <p><strong>References</strong></p> <p>BBC SNN (2015): A.Pearce. T.Brookes, M.Dewhirst, &quot;Timbral differences between microphones&quot;, BBC Sound Now &amp; Next Technology Fair, London, UK, 19-20 May 2015</p> <p>AES139 (2015): A.Pearce. T.Brookes, M.Dewhirst, &quot;Validation of experimental methods to record stimuli for microphone comparisons&quot;, Audio Eng.Soc. 139th Convention, New York, USA, 29 Oct - 1 Nov 2015</p>

opencc-by-nc-4.0Oct 2015View details →
zenodo32/100

Dataset for Validation Experiments of MITMProbe

<p>Dataset for validation experiments of man-in-the-middle-probe video proxy performance evaluation.</p>

opencc-by-4.0May 2017View details →
zenodo32/100

Observed monthly N2O emission dataset used for model carlibaration and validation

<p>The N<sub>2</sub>O emission dataset for calibration sites was extracted from the published figures and tables using GetData Graph Digitizer version 2.24; the other information such as biome, geographic location, experimental period, soil organic carbon content, soil pH and soil texture was selected from corresponding literature. If the data related to soil was not avaliable, we extracted them from the soil database(IGBP-DIS)</p>

opencc-by-nc-nd-4.0Jun 2017View details →
zenodo32/100

FIGURE 2. Phylogenetic results. A, Maximum likelihood tree from COI dataset rooted with Ophelia limacina. B, Maximum likelihood tree from ITS1 in Validation of three sympatric Thoracophelia species (Annelida: Opheliidae) from Dillon Beach, California using mitochondrial and nuclear DNA sequence data

FIGURE 2. Phylogenetic results. A, Maximum likelihood tree from COI dataset rooted with Ophelia limacina. B, Maximum likelihood tree from ITS1 dataset rooted according to the result for the COI dataset. Support values are shown as jackknife from parsimony analysis and bootstrap from maximum likelihood respectively separated by /. * indicates 100% values for each support measure.

opennotspecifiedJan 2013View details →
zenodo32/100

Synthesized training and validation dataset for DeepMB

Open the record for dataset details and reuse information.

openmit-licenseNov 2023View details →
zenodo32/100

Dataset for the Report on 2.3. "Equipment prototypes validation in laboratory conditions"

<p>The system has been tested using signals that were generated using a mathematical model based on physical processes when applying wind lidar using Matlab software.</p><p>3 types of signals were considered:</p><p>1.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Simulation of a set of backscattering events for a signal that does not attenuate (does not exist in reality)</p><p>2.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Simulation of a set of backscattering events for a signal attenuating proportional to the square of the distance (does not exist in reality)</p><p>3.&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Simulation of a set of events to check the RMS of the system resolution period</p><p>To perform the tests, the system FPGA code was modified to allow playback of the test signals downloaded from the PC through the reserved Fire-out and Event out connectors (upper left connectors of the system).</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

A High-Density Land cover Validation dataset in Xinjiang—HDLV-XJ

<p>A 2020 High-Density Land cover Validation dataset in Xinjiang. To ensure a sufficient number of samples in complex areas and to position appropriate sampling points in homogeneous and heterogeneous areas, the equal-area stratified random sampling method based on multiple indicators was utilized.&nbsp;</p><p>The HDLV_XJ includes 20,932 validation samples. It considerably higher sample numbers for each land cover type compared to the other datasets.The HDLV-XJ provides representative validation data with sufficient samples for rare categories, enabling a more accurate assessment of the accuracy of land cover products in the Xinjiang. We provide an xls file of this validation dataset and the code used in constructing the dataset.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

PathoFact 2.0 validation test datasets

<p>Test datasets used to validate PathoFact 2.0 performance</p>

opengpl-3.0-or-laterNov 2024View details →
zenodo32/100

Temporal validity of software datasets for code metrics: an empirical assessment of sampling strategies

<p>This is the repository for the scripts and data of the study "Building and updating software datasets: an empirical assessment".</p> <h2>Data collected</h2> <p>The data generated for the study it can be downloaded as a zip file. Each folder inside the file corresponds to one of the datasets of projects employed in the study (qualitas, currentSample and qualitasUpdated). Every dataset comprised three files "class.csv", "method.csv" and "sample.csv", with class metrics, method metrics and repository metadata of the projects respectively. Here is a description of the datasets:</p> <ul> <li>qualitas: includes code metrics and repository metrics from the projects in the release 20130901r of the Qualitas Corpus.</li> <li>currentSample: includes code metrics and repository metrics from a recent sample collected with our sampling procedure.</li> <li>qualitasUpdated: includes code metrics and repository metrics from an updated version of the Qualitas Corpus applying our maintenance procedure.</li> </ul> <h2>Plot graphics</h2> <p>To plot the results and graphics in the article there is a Jupyter Notebook "Experiment.ipynb". It is initially configured to use the data in "datasets" folder.</p> <h2>Replication Kit</h2> <p>For replication purposes, the datasets containing recent projects from Github can be re-generated. To do so, the virtual environment must have installed the dependencies in "requirements.txt" file, add Github's tokens in "./token" file, re-define or leave as is the paths declared in the constants (variables written in caps) in the main method, and finally run "main.py" script. The portable versions of the source code scanner&nbsp;<a href="https://sourcemeter.com/" target="_blank" rel="noopener">Sourcemeter</a> are located as zip files in "./Sourcemeter/tool" directory. To install Sourcemeter the appropriate zip file must be decompressed excluding the root folder "SourceMeter-10.2.0-x64-&lt;OS&gt;".</p> <p>The script comprise 5 steps:</p> <ol> <li>Project retrieval from Github: at first the sampling frame with projects complying with a specific quality criteria are retrieved from Github's API.</li> <li>Create samples: with the sampling frame retrieved, the current samples are selected (currentSample and qualitasUpdated). In the case of qualitasUpdated, it is important to have first the "sample.csv" file inside the qualitas folder of the dataset originally created for the study. This file contains the metadata of the projects in Qualitas Corpus.</li> <li>Project download and analysis: when all the samples are selected from the sampling frame (currentSample and qualitasUpdated), the repositories are downloaded and scanned with SourceMeter. In the cases in which the analysis is not possible, the projects are replaced with another one with similar size.</li> <li>Outlier detection: once the datasets are collected, it is necessary to manually look for possible outliers in the code metrics under study. In the notebook "Experiment.ipynb" there are specific sections dedicated for it ("Outlier detection (Section 4.2.2)").</li> <li>Outlier replacement: when the outliers are detected, in the same notebook there is also a section for outlier replacement ("Replace Outliers") where the outliers' url have to be listed to find the appropriate replacement.</li> </ol> <ul> <li>If it is required, the metrics from the Qualitas Corpus can also be re-generated. First, it is necessary to download the release 20130901r from its <a href="http://www.qualitascorpus.com/download/" target="_blank" rel="noopener">official webpage</a>. Second, decompress the .tar files downloaded. Third, make sure that the compressed files with source code from the projects (.java files) are placed in the "compressed" folder, in some cases it is necessary to read the "QC_README" file in the project's folder. Finally, run the original main script "Generate metrics for the Qualitas Corpus (QC) dataset" part of the code. &nbsp;</li> </ul>

openmit-licenseApr 2024View details →
zenodo32/100

MSNovelist: scripts and datasets for validation and de novo discovery.

<p>This repository contains all scripts and datasets required to reproduce the validation and the de novo discovery application in the MSNovelist paper. See README.md for details. The MSNovelist software itself is available on Github: https://github.com/meowcat/MSNovelist</p>

opencc-by-nc-4.0Nov 2021View details →
zenodo32/100

Dataset of the study "Factor Structure, Validity, and Reliability of the STarT Back Screening Tool in Italian Obese and Non-obese Patients With Low Back Pain"

<p>Dataset of the study &quot;Factor Structure, Validity, and Reliability of the STarT Back Screening Tool in Italian Obese and Non-obese Patients With Low Back Pain&quot; published in Frontiers in Psychology 20 October 2021 12:740851, doi: 10.3389/fpsyg.2021.740851</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Training and validation sample datasets for Qinghai-Tibet Plateau Forest Cover Map 2021

<p>The training and validation sample datasets are label by the value of &quot;0&quot; and &quot;1&quot;, when &quot;0&quot; means the non-forest samples and &quot;1&quot; means the forest samples.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Dataset: Multi-Stakeholder Validation of Entrustable Professional Activities for a Family Medicine Care of the Elderly Residency Program: A Focus Group Study

<p>ZOOM Meeting transcripts from 5 stakeholder group meetings used in the study.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Dynamic deformation calculation of articular cartilage and cells using resonance-driven laser scanning microscopy - Osmotic Challenge and Validation Testing Dataset

<p>This archive contains 3-D image stacks obtained over time of articular cartilage undergoing osmotic swelling. These were acquired with a resonance scanning protocol, which allows for fast scanning but with&nbsp;decreased image quality. The Python package, resonant_lsm, was developed to segment and analyze the deformation of cells in such images. The archive also contains validation testing data of this software. Included README files document the archive contents and how to reproduce the analyses.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

Datasets and model outputs used for validating WASP for SURFEX v8.0

<p>These datasets contain the model outputs used for validating the WASP SURFEX v8.0 parameterization of turbulent fluxes at sea. The file LION_BUOY.tar contains (ascii format) the tubrulent fluxes computed from the surface parameters recorded at the M&eacute;t&eacute;o-France LION buoy (in the centre of the Gulf of Lion, NW Mediterranean Sea) at an hourly frequency between 2001 and 2014, using the bulk softwares COARE 3.0 with various wave representation and WASP with waves. The file AROME_MED.tar contains the outputs of the AROME operational model with ECUME or WASP representation of the turbulent fluxes, on the case study of October 2016, that were used to validate the parameterization as presented in the paper. The file AROME_OI.tar contains the outputs of AROME OI on the case of the tropical cyclone Batsirai in the Indian Ocean, in February 2022. The file ARPEGE_CLIM_AMIP contains the outputs of the low-resolution, AMIP runs performed using the ECUME and WASP parameterization.</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Trojan Detection Challenge Datasets (New Detection Track Validation Set)

<p>This upload contains the new version of the detection track validation set for the Trojan Detection Challenge NeurIPS 2022 competition.&nbsp;This is meant for participants in the detection who do not want to redownload the full dataset. For the full dataset including data for other tracks, see&nbsp;<a href="https://zenodo.org/record/6894041">https://zenodo.org/record/6894041</a>.</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

Validation of Peak Ground Velocities Recorded on Very-high rate GNSS Against NGA-West2 Ground Motion Models: Dataset

<p>Contains all the data files for the manuscript of the same name.</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record