Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,038
datasets available to search
ShareScore release 0.7.1
Dataset results
8,038 results for “validation”
Validation of Emission Spectroscopy Gas Temperature Measurements Using a Standard Flame Traceable to the International Temperature Scale of 1990 (ITS-90)
<p>Data underpinning the associated publication (https://doi.org/10.1007/s10765-019-2557-6) on accurate traceable measurement of post-flame temperatures.</p>
A Dataset of Global Land Cover Validation Samples
<p>A dataset of global land cover validation samples in 2015. In order to guarantee the confidence and objective of the validation samples, several existing reference datasets such as GLCNMO2008 training dataset, VIIRS reference dataset, STEP reference dataset, Global cropland reference data and so on, high resolution imagery in the Google earth and time-series NDVI,NDSI values of each related point are integrated to derive the global validation datasets. The dataset is provided in .shp format.</p>
HydroGeoSphere Model Input Files and Results for Validation of Pesticide Leaching to Groundwater for Nine EU FOCUS Scenarios
<p>This dataset includes HydroGeoSphere (HGS) model (Aquanty, 2024) input and output files for nine Forum for the Co-ordination of Pesticide Models and their Use (FOCUS) scenarios (EC, 2014) for simulation of leaching of four test contaminants to groundwater. It is recommended that users are familiar with HGS software in order to best make use of the available files. Scenarios are included in separate subfolders named using the first four letters of the FOCUS scenario location name, e.g. folder "chat" contains the model run for the "Chateaudun" scenario. It is recommended that users familiarize themselves with the (EC, 2014) groundwater scenarios. HGS model inputs are specified in the *.grok ASCII text file for each scenario in each subfolder. Soil material properties and evapotranspiration properties are included in HGS input files in ASCII text format in the "material_properties" subfolder. Solute application timing for each scenario are include in the "solute_app" subfolder. And climate times series inputs are included in the "weather" subfolder.</p>
QUADCOIL Validation Dataset
<p>This dataset contains the a prototype and the validation data for the stellarator coil optimization code QUADCOIL. Please extract and see <code>readme.md</code> for instructions to reproduce.</p>
Validated TRANSP simulations of Alcator C-Mod Experimental Plasma
<p>This dataset contains the inputs and outputs from TRANSP plasma transport simulations initialized using experimental data from the Alcator C-Mod tokamak plasma experimental device. The simulations include sawtooth instabilities according to the Kadomtsev model and the outputs are time-averaged over these crashes to emulate steady-state plasma scenarios.</p> <p>The data is stored within a NETCDF4 file written using the Python xarray package, which can be opened via the xarray package with the `open_dataset()` function. Only a subset of the output fields deemed relevant for plasma transport simulations are included, in order to reduce data bloat. For more information on the TRANSP code, including its input and output fields, please see https://transp.pppl.gov/.</p>
Experimental data for validation of a variational RANS level III flow model: water waves over an array of obstacles and Ogee weir flows
<p>Experimental dataset for the validation of a variational RANS level III flow model. The experimental data correspond to experiments on unsteady of water waves over an array of obstacles and steady curved flows over an Ogee weir. The experiments were conducted at the Hydraulics Laboratory at the Univeristy of Córdoba. </p>
Data from: External validation of prognostic and predictive gene signatures in 1097 European head and neck squamous cell carcinoma patients
<p><span>Anonymized data containing survival endpoints and gene signature scores for head and neck cancer patients.</span></p> <p><span>File <strong>data_os_gs.csv</strong> : data linking overall survival and gene signature scores</span></p> <p><span>File <strong>data_dfs_gs.csv</strong> : data linking disease-free survival and gene signature scores</span></p> <p><span><strong>Variables</strong>:</span></p> <ul> <li><span><em>supertreat_id</em>: patient ID</span></li> <li><span><em>GS_score_172GS</em>: gene signature score for the <em>172-GS</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_3clustersHPV</em>: gene signature score for the <em>3 clusters HPV</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_RSI</em>: gene signature score for the <em>radiosenstivity index (RSI) </em>signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_pancancerCisplatin</em>: gene signature score for the <em>pancancer-cisplatin</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_cl3Hypoxia</em>: gene signature score for the <em>Cl3-hypoxia</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span>Variables only available in <strong>data_os_gs.csv: </strong></span> <ul> <li><span><em>overall_survival_days_2years</em>: Overall survival censored at 2 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_days_5years</em>: Overall survival censored at 5 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_status_2years</em>: Overall survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> <li><span><em>overall_survival_status_5years</em>: Overall survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> </ul> </li> </ul> <ul> <li><span>Variables only available in <strong>data_dfs_gs.csv:</strong></span> <ul> <li><span><em>disease_free_survival_days_2years</em>: Disease-free survival censored at 2 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_days_5years</em>: Disease-free survival censored at 5 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_status_2years</em>: Disease-free survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> <li><span><em>disease_free_survival_status_5years</em>: Disease-free survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> </ul> </li> </ul>
Artifacts for ISSTA 2021 Paper : Validating Static Warnings via Testing Code Fragments
<p><strong>This data set is for ISSTA 2021 Paper: Validating Static Warnings via Testing Code Fragments</strong></p> <p>Static analysis is an important approach for finding bugs and vulnerabilities in software. However, inspecting and confirming static warnings are challenging and time-consuming. In this paper, we present a novel solution that automatically generates test cases based on static warnings to validate true and false positives. We designed a syntactic patching algorithm that can generate syntactically valid, semantic preserving executable code fragments from static warnings. We developed a build and testing system to automatically test code fragments using fuzzers, KLEE and Valgrind. We evaluated our techniques using 12 real-world C projects and 1955 warnings from two commercial static analysis tools. We successfully built 68.5% code fragments and generated 1003 test cases. Through automatic testing, we identified 48 true positives and 27 false positives, and 205 likely false positives. We matched 4 CVE and real-world bugs using Helium, and they are only triggered by our tool but not other baseline tools. We found that testing code fragments is scalable and useful; it can trigger bugs that testing entire programs or testing procedures failed to trigger.</p>
Decomposition Tool Validation Dataset 1
<p>This dataset provides artifacts, benchmarks and models used to evaluate the accuracy of the performance model as well as the corresponding results:</p> <ul> <li>Deployment package of the thumbnail generation function</li> <li>50 images collected to benchmark the actual deployment</li> <li>TOSCA model with the specified open workload and performance requirement</li> <li>Predictions given by the performance model and measurements taken from the cloud platform</li> </ul>
A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models
<p><strong>A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models.</strong></p> <p>Please refer to the following scientific paper for a description of the dataset.</p> <blockquote> <p>Holtrop, T.; Van Der Woerd, H.J. (accepted) HYDROPT: An Open-Source Framework for Fast Inverse Modelling of Multi- and Hyperspectral Observations from Oceans, Coastal and Inland Waters. <em>Remote Sens. </em><strong>2021</strong>, 13, 0.</p> </blockquote>
Artifacts supplementing the EuroUSEC '22 paper "Assessing Real-World Applicability of Redesigned Developer Documentation for Certificate Validation Errors"
<p>This upload supplements the conference EuroUSEC 2022 submission by providing the full questionnaire, anonymized dataset and all performed analyses presented in the paper specified below.</p> <ul> <li>Title: <strong>Assessing Real-World Applicability of Redesigned Developer Documentation for Certificate Validation Errors</strong></li> <li>Authors: Martin Ukrop, Michaela Balážová, Pavol Žáčik, Eric Vincent Valčík, Vashek Matyas</li> <li>Paper details: https://crocs.fi.muni.cz/public/papers/eurousec2022</li> <li>Paper abstract: <em>We face certificate validation errors commonly, yet the related tools and documentation had been shown to have very poor usability. Previous research suggests that just improving the error messages and corresponding documentation can have significantly positive effects. Our work aims at increasing the usability of certificate validation by 1) redesigning the API error messages and the corresponding documentation, and 2) validating the real-world applicability of the redesign by investigating the opinions of 180 IT professionals. We focus on the perceived obstacles, desired ideal form and overall satisfaction. The redesigned documentation exhibits a reliable significant decrease in perceived incompleteness, with a small amount of perceived bloat and tangle. The redesigned documentation, now published on a dedicated website, is preferred by 89% of our study participants.</em></li> </ul> <p>The artifacts accompanying this paper contain three major parts:</p> <ul> <li>The questionnaire used in the main study (described in Sections 3.1 and 3.2 of the paper and mostly present in Appendices A and B of the paper).</li> <li>The anonymized dataset (multiple formats) of all valid questionnaire answers and qualitative coding performed. Analyses of this dataset are the core of the paper and are present in subsection 3.4 and all parts of Sections 4 and 5.</li> <li>The set of analyses files (IBM SPSS scripts and outputs) producing all statistical results presented in the paper are included.</li> </ul> <p>More details about the artifacts can be found in the README file in the artifacts archive.</p>
Synthetic bus validation of Rennes public transportation network focus on Beaulieu campus
<p>Synthetic data based on the STAR/Keolis open data version GTFS_2021.1.0.3_20210830_20211017 for regular weeks and GTFS_2020.10.1_20210705_20210829 for holidays. This dataset contains the generation of bus validations for 3500 individuals during 14 weeks with an average of 15 travels per week. It also contains GPS information of the bus stops associated to the validation.</p>
RRI2SCALE_D2.3_Scenario Validation_2021.10.15_v1
<p>This dataset contains the results of a short survey to assess the realisation probability and the desirability of alternative techno-moral scenarios in the domains of intelligent cities, intelligent transport, and intelligent energy.</p> <p>More info:</p> <p>RRRI2SCALE partners designed and implemented a validation process to assess whether the six techno-moral scenarios developed previously (two per domain, i.e., intelligent cities, intelligent transport and intelligent energy) meet high-quality scenario criteria (are probable, desirable, different from one another, complete and internally consistent).</p> <p>More specifically, for the deliberation with the citizens, videos presenting the scenarios were created and uploaded on the RRI2SCALE website and social media (available here: https://rri2scale.eu/index.php/regional-validation-of-scenarios/). Citizens from the four regions (i.e., Kriti, Galicia, Overijssel and Vestland) were prompted to participate in a short survey, assigning scores to scenarios’ desirability and probability. The dataset uploaded includes all the replies received.</p> <p>Facebook was the primary platform selected to facilitate and promote the scenario deliberation process. According to statistics provided by the Facebook “Ads Manager” tool, the RRI2SCALE campaign reached more than 31,000 people. Almost half of them (about 13,000) watched the videos, more than 2,000 people clicked on the survey link, and finally, we collected 279 valid survey replies from citizens of the four regions.</p> <p>Among the people who clicked on the survey link, 60% were male, while most (about 62%) fall within the age of 45-64. Among 279 valid replies received, 114 replies were received from the citizens of Galicia, 93 from Kriti, 37 from Vestland and 35 from Overijssel. From another perspective, among the 279 valid replies received, 92 replies refer to the Intelligent Cities scenarios, 104 replies refer to the Intelligent Transport scenarios, and 83 replies refer to the Intelligent Energy scenarios.</p>
List of validated primers of gilthead sea bream (Sparus aurata) and European seabass (DIcentrarchus labrax) developed in PerformFISH project (D2.3)
<p>The document contains all the primers identified for the screening of genes tested for their potential as biomarkers to predict quality performance in gilthead sea bream and European sea bass larvae and juveniles in the context of PERFORMFISH (WP2). The spreadsheet has the following information: Pathway, phisiologic process in which the gene is involved; name of protein that gene produces; gene code; acession nº, code given in the consulted databases and the sequence extracted for primer design; FW and RV primer, forward and reverse primer sequence specific for target gene; melt temperature, optimized temperature that primers work at ; amplicon size, size in base pairs of the product produced with the primers; eff%, efficency of primers; r2; source, the origin of the primers, "in house" or "literature" (including available DOI. Each pair of primers are classified using a "traffic light" system indicating their validation status.</p>
Automated Minirhizotron Validation Data
<p>This dataset is contains the validation data (raw images and binary masks generated by hand) for the automated minirhizotrons in our paper 'HIGH FREQUENCY ROOT DYNAMICS: SAMPLING AND INTERPRETATION USING REPLICATE ROBOTIC MINIRHIZOTRONS' published in Journal of Experimental Botany. <a href="https://doi.org/10.1093/jxb/erac427">https://doi.org/10.1093/jxb/erac427</a></p> <p>The dataset consists of one readme, four .7z files and one python script all contained in one .7z file.</p> <p>The files are described in the readme file.</p> <p>Some of these data were collected and all processed as part of the Marie Sklodowska-Curie project 748893 'MrPARTS' awarded to Richard Nair. We also acknowledge the generous support of Markus Reichstein at MPI-BGC Jena including to the MaNiP project through the Max Planck research prize 2013</p>
The DESI Survey Validation: Results from Visual Inspection of the Quasar Survey Spectra
<p>Data files to reproduce the published figures from Alexander et al. (2022), AJ, in press (<a href="https://ui.adsabs.harvard.edu/link_gateway/2022arXiv220808517A/arxiv:2208.08517">arXiv:2208.08517</a>), titled "The DESI Survey Validation: Results from Visual Inspection of the Quasar Survey Spectra". The figure number is encoded in the file name to make it easy to relate the published figures to the corresponding data file. Please see the published paper for detailed information on what is plotted for each figure.</p>
Datasets from BUBBLES validation exercises
<p>This dataset contains the telemetry data sent by the drones during the validation exercises and the response from the BUBBLES Separation Management Environment Platform. The data was gathered during test flights, in which 14 drones performed different representative operations, including agricultural tasks, surveillance, deliveries and lifeguard operations.</p> <p>For more information about the test flights, see D5.1 Validation plan and D5.4 Validation report from BUBBLES project.</p>
Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images
<p>Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images. A README file with the contents of the dataset is included. </p>
Validation of MODIS11A2 LST and glacier surface heatwave during 2001-2020 over Tibetan Plateau
<p>1,Validation of MODIS11A2 LST in 2019 using AWS temperature on the glacier</p> <p>2,Validation of MODIS11A2 LST during 2001-2020 using CMA station temperature over the Tibetan Plateau</p> <p>3,Glacier surface heatwave during 2001-2020 over the Tibetan Plateau glacier </p>
DUDE competition train - validation - test splits ground truth
<p>This JSON file contains the ground truth annotations for the train and validation set of the DUDE competition (https://rrc.cvc.uab.es/?ch=23&com=tasks) of ICDAR 2023 (https://icdar2023.org/).</p> <p> </p> <p><strong>V1.0.7 release</strong>: 41454 annotations for 4974 documents (train-validation-test)</p> <pre>DatasetDict({ train: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 23728 }) val: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 6315 }) test: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 11402 }) }) ++update on answer_type +++formatting change to answers_variants ++++stricter check on answer_variants & rename annotations file <strong>+ blind test set (no ground truth answers provided) </strong>++ removed duplicates from test set: </pre> <blockquote> <p> "92bd5c758bda9bdceb5f67c17009207b_ac6964cbdf483e765b6668e27b3d0bc4",</p> <p> "6ee71a16d4e4d1dbd7c1f569a92d4e08_549f2a163f8ff3e9f0293cf59fdd98bc",</p> <p> "e6f3855472231a7ca6aada2f8e85fe5a_827c03a72f2552c722f2c872fd7f74c3",</p> <p> "e3eecd7cca5de11f1d17cd94ae6a8d77_6300df64e4cf6ba0600ac81278f68de2",</p> <p> "107b4037df8127a92ee4b6ae9b5df8fb_d7a60e7a9fc0b27487ea39cd7f56f98e",</p> <p> "300cc3900080064d308983f958141232_6a7cf1aad908d58a75ab8e02ddc856f4",</p> <p> "fdd3308efacddb88d4aa6e2073f481d4_138cb868ecc804a63cc7a4502c0009b2",</p> <p> "1f7de256ff1743d329a8402ba0d132e7_95b6e8758533a9817b9f20a958e7b776",</p> <p> "4f399b8c526ffb6a2fd585a18d4ed5ec_51097231bc327c26c59a4fd8d3ff3069",</p> </blockquote> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.