Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,038
datasets available to search
ShareScore release 0.7.1
Dataset results
8,038 results for “validation”
Validating additive correction schemes against gradient-based extrapolations
<p>Supporting information for the paper.<br> <br> Python files are used to generate figures, which will be placed in the "figures" folder. The table include files in "si" are also automatically generated.</p> <p>Each of the other folders includes all data (inputs, shell scripts and outputs) of the Psi4 calculations for a given species, with the exception of the NCDT16 folder, which is a benchmark code developed previously.</p>
Dataset with manually validated version histories of Stack Overflow posts
<p>We used this dataset to evaluate different string similarity metrics for SOTorrent (http://sotorrent.org/). For the versions published 2018-11-01 and 2018-12-14, we double-checked and updated the ground truth files.</p> <p>The dataset has been created with this tool: https://github.com/sotorrent/posthistory-gt</p> <p>The dataset has been validated with this tool: https://github.com/sotorrent/posthistory-comparator-gt-cs</p> <p>The dataset has been used in this project: https://github.com/sotorrent/metric-evaluation</p> <p>The most recent version of the files can always be found here: https://github.com/sotorrent/metric-evaluation/tree/master/testdata/samples_comparison</p>
Building Brain Invaders: EEG data of an experimental validation
<p><strong>Summary:</strong></p> <p>This dataset contains electroencephalographic (EEG) recordings of 25 subjects testing the <em>Brain Invaders </em>(Congedo, 2011), a visual P300 Brain-Computer Interface inspired by the famous vintage video game <em>Space Invaders</em> (Taito, Tokyo, Japan). The visual P300 is an event-related potential elicited by a visual stimulation, peaking 240-600 ms after stimulus onset. EEG data were recorded by 16 electrodes in an experiment that took place in the GIPSA-lab, Grenoble, France, in 2012 (Van Veen, 2013 and Congedo, 2013). A full description of the experiment is available <a href="https://hal.archives-ouvertes.fr/hal-02126068">https://hal.archives-ouvertes.fr/hal-02126068</a>. Python code for manipulating the data is available at <a href="https://github.com/plcrodrigues/py.BI.EEG.2012-GIPSA">https://github.com/plcrodrigues/py.BI.EEG.2012-GIPSA</a>. The ID of this dataset is <em>BI.EEG.2012-GIPSA</em>.</p> <p> </p> <p><strong>Full description of the experiment and dataset: </strong><a href="https://hal.archives-ouvertes.fr/hal-02126068">https://hal.archives-ouvertes.fr/hal-02126068</a></p> <p> </p> <p><strong><em>Principal Investigator</em>:</strong> B.Sc. Gijsbrecht Franciscus Petrus van Veen</p> <p> </p> <p><strong><em>Technical Supervisors</em>: </strong>Ph.D. Alexandre Barachant, Eng. Anton Andreev, Eng. Grégoire Cattan, Eng. Pedro. L. C. Rodrigues</p> <p> </p> <p><strong><em>Scientific Supervisor:</em></strong> Ph.D. Marco Congedo</p> <p> </p> <p><strong>ID of the dataset: </strong><em>BI.EEG.2012-GIPSA</em></p>
Validity and reliability of the Kinovea program in obtaining angles and distances using coordinates in 4 perspectives
<p>Dataset used to perform the statistical analysis for the article "Validity and reliability of the Kinovea program in obtaining angles and distances using coordinates in 4 perspectives"</p>
Raw data on the validation of the Polish language version of THE BRIEF HEPATITIS C KNOWLEDGE SCALE
<p>Raw data on the validation of the Polish language version of THE BRIEF HEPATITIS C KNOWLEDGE SCALE.</p>
Global River Ice Dataset - validation dataset
<p><strong>Documentation for nws_breakup_nogeo.csv and nws_freezeup_nogeo.csv</strong></p> <p>Alaskan river ice records from National Weather Service (NWS), including <strong>nws_breakup_nogeo.csv</strong> containing location (description) and dates of ice breakup and related conditions and <strong>nws_freezeup_nogeo.csv </strong>containing location (description) and dates of ice freeze-up and related conditions. Note that both dataset do not contain exact geolocations of the observation. We thank Dr. Scott Lindsey at the Alaska-Pacific River Forecast Center for providing these datasets.</p> <p><strong>Documentation for landsat_river_ice_validation.csv</strong></p> <p>This file contains 20,687 same-day river ice condition from Landsat and from in situ, and consists of the following associated properties for each comparison:</p> <ol> <li>date: The date on which both the Landsat river ice (length) fraction and in situ river ice condition were observed (data type: string; format: "YYYY-MM-DD").</li> <li>ice_in_situ: The ice condition on rivers observed in situ. For records from NWS, we assumed river has been ice covered between the date of "first_ice" to the date of "breakup" in the following year and ice-free between the date of "breakup" and the following "first_ice" date. For records from Water Survey of Canada, river was treated as ice-covered whenever the daily "Flow" data were flagged with "B"–meaning backwater effect (data type: integer; range: 0 (ice-free) or 1 (ice-covered)).</li> <li>ice_landsat: The river ice length fraction derived from Landsat image (data type: float; range: [0, 1]).</li> <li>cloud_landsat: The cloud fraction derived from Landsat image (data type: float; range: [0, 0.25]).</li> <li>LANDSAT_SCENE_ID: The unique Landsat TOA image identifier (data type: string).</li> <li>site_id: The ID of the site in its original dataset.</li> <li>dat_source: The source of the in situ river ice record (data type: string; values: ("National Weather Service (Alaska)", "Water Survey of Canada").</li> <li>longitude: The longitude of the site (data type: float, format: decimal degree).</li> <li>latitude: The latitude of the site (data type: float, format: decimal degree).</li> </ol> <p>A subset (N = 18,930) of this dataset was used in the evaluation of the river ice classification. This subset was calculated by applying the following two constraints on the full dataset in the <strong>landsat_river_ice_validation.csv</strong>:</p> <ol> <li><span class="math-tex">\(cloud\_landsat ≤ 0.05\)</span></li> <li><span class="math-tex">\(site\_id \neq 10BE013\)</span> & <span class="math-tex">\(site\_id \neq 08KE016\)</span></li> </ol> <p>The second constraint exclude two Canadian sites from the evaluation as via manual inspection, we found that the Landsat-derived ice fraction for this two sites came from river reaches that were different from where the in situ records were observed.</p>
Data file with manuscript titled 'A Structurally Validated Sequence Alignment of 497 Human Protein Kinase Domains'
<p>The files used in different analysis reported in the manuscript titled - 'A Structurally-Validated Multiple Sequence Alignment of 497 Human Protein Kinase Domains' are shared at two locations. Following is a brief description of these files.</p> <p>Location - https://github.com/DunbrackLab/Kinases<br> 1. HMM profile files - HMM files for each of the nine groups computed separately labeled as Groupname.hmm, like AGC.hmm<br> 2. HMM profile file - HMM file computed from the full alignment including all the sequences - Human-PK.hmm<br> 3. Score files - HMM scores of each kinase sequence against all the groupwise HMMs both for iteration1 (HMM-iter1-scores-tables.txt) and iteration2 (HMM-iter1-scores-tables.txt)<br> 4. Jalview session file - Kinase alignment with sequences colored by secondary structure information from PDB file if the structure is known; or predicted secondary structure if the experimental structure is not known. The file could be opened in Jalview - kinases-PDB-SSPred.jvp</p> <p>Location - https://zenodo.org/record/3445533<br> 1. The file contains list of residue pairs aligned in pairwise structural alignments of 272 human protein kinases which were used as a benchmark in the study. The alignments were created by FATCAT and optimized by SE program.</p>
NEECK Validation: Acoustic Measurements and BEM Simulations
<p>This repository contains the supporting data for the paper entitled "Acoustic Validation of a BEM-Suitable 3D Mesh Model of KEMAR'', K. Young, G. Kearney, and A. I. Tew, at the 2018 AES International Conference on Spatial Reproduction - Aesthetics and Science, Tokyo. Available at: http://www.aes.org/e-lib/browse.cfm?elib=19662. Please cite both the paper and dataset if used.</p> <p>Note: the azimuth angle system used in this work increments positively in the left direction, such that 90° is on the left and 270° is on the right. In elevation, -90° is below, +90° is above. </p> <p>---</p> <p>The data is organised as follows:</p> <p>- NEECK_HRIR_measured.sofa<br> (SOFA file (SimpleFreeFieldHRIR) containing the 185 acoustically measured HRIRs for the Neck-Extended Easily Computable KEMAR (NEECK))<br> - NEECK_HRTF_simulated.sofa<br> (SOFA file (SimpleFreeFieldTF) containing the 10,205 simulated HRTFs for the Neck-Extended Easily Computable KEMAR (NEECK))<br> - AdditionalData<br> (Zip folder containing data processed during the analysis stages)<br> - averageResponse_measured.mat<br> (mat file containing the average IR responses, corresponding inverse filters and inverse filter generation parameters)<br> - averageResponse_simulated.mat<br> (mat file containing the average TF responses in linear scale)<br> - measuredData.mat<br> (mat file containing the following data:)<br> - IRs<br> (Measured impulse responses as in SOFA file. Dimensions: M1xRxN1)<br> - IRs_DTF<br> (Impulse responses after application of average response inverse filter. Dimensions: M1xRxN1)<br> - HRTFs<br> (HRTF magnitudes in linear scale. Dimensions: M1xRxN1)<br> - HRTFs_dB<br> (As above in decibel scale. Dimensions: M1xRxN1)<br> - DTFs<br> (DTF magnitudes in linear scale - after application of average response inverse filter. Dimensions: M1xRxN1)<br> - DTFs_dB<br> (As above in decibel scale. Dimensions: M1xRxN1)<br> - measFs<br> (sampling rate of measured responses: Dimensions: 1x1)<br> - allSourcePositions_measured<br> (measured source positions in spherical coordinates (azimuth, elevation, radius). Units: degrees, degrees, metres. Dimensions: M1x3)<br> - simulatedData.mat<br> (mat file containing the following data:<br> - IRs<br> (Impluse responses generated from the simulated HRTF data. Dimensions: M2xRxN2)<br> - HRTFs_complex<br> (Complex simulated HRTF data. Dimensions: M2xRxN3)<br> - HRTFs_mag_dB<br> (Magnitudes of simulated HRTF data in decibel scale. Dimensions: M2xRxN3)<br> - DTFs_mag_lin<br> (Magnitudes of directional transfer function (DTF) data in linear scale. Dimensions: M2xRxN3)<br> - DTFs_mag_dB<br> (As above in decibel scale. Dimensions: M2xRxN3)<br> - simFs<br> (sampling rate of generated impulse responses. Dimensions: 1x1)<br> - allSourcePositions_simulated<br> (simulated source positions in spherical coordinates (azimuth, elevation, radius). units: degrees, degrees, metres. Dimensions: M2x3)<br> - frequencies<br> (frequencies used in the simulation. Dimensions: N3x1)<br> - license.mat<br> (mat file containing licensing information)<br> - License.txt<br> (Text file detailing the license under which this data is published.)</p> <p>For enquiries regarding the data in a different format, please email kaey500@york.ac.uk. <br> ---</p> <p>Data Dimensions:</p> <p>M1 = number of measured source positions, in this case 185<br> M2 = number of simulated source positions, in this case 10,205<br> R = number of channels, in this case 2, where 1 and 2 correspond to left and right respectively<br> N1 = number of samples in measured impulse responses, in this case 1024<br> N2 = number of samples in generated impulse responses, in this case (number of samples in HRTF*2)+2 = 400<br> N3 = number of samples in simulated transfer functions, in this case, the number of frequency points, 199</p> <p>---</p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License (http://creativecommons.org/licenses/by-nc/4.0/), with no warranty; or the implied warranty of merchantability or fitness for a particular problem.</p> <p>---</p> <p>Data produced by Kat Young at the AudioLab, Dept. of Electronic Engineering, University of York.<br> Contact: kaey500@york.ac.uk</p>
Dataset 1. Validated natural and cryptic mRNA splicing mutations
<p>Source data computed by the Shannon pipeline and Veridical, displayed on the ValidSpliceMut (<a href="https://validsplicemut.cytognomix.com/">https://validsplicemut.cytognomix.com/</a>) website.</p>
Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors - apo and validation MD
<p>Supplementary data of "Ligand-induced Conformational Selection Predicts the Selectivity of Cysteine Protease Inhibitors" paper.</p> <p>This dataset consists of molecular dynamics simulations trajectories and topology of Cruzain, Cathepsin K and Cathepsin L enzymes in it apo form, together with validation simulations. We ran five replicate 100ns simulations on each complex, with randomized initial velocities.</p>
An exploratory study (with and without time pressure) using mouse dynamics to detect faking-good behavior in the MMPI-2 and PPI-R validity scales
<p>Please find here the dataset generated and analyzed during the study entitled "Can mouse dynamics detect faking-good behavior in personality questionnaires? An exploratory study (with and without time pressure) using the MMPI-2 and PPI-R validity scales". Moreover, here you can find the source code of the experiment to execute the task using MouseTracker software, the code of the statistical analysis and a file containing the instructions to replicate ML model results reported in the original paper.</p>
Figure 1 in Validation of reference genes for quantitative expression analysis by qPCR in various tissues of date mussel (Lithophaga lithophaga)
Figure 1. Distribution of Cq values of candidate reference genes in date mussel (L. lithophaga).
Figure 2 in Validation of reference genes for quantitative expression analysis by qPCR in various tissues of date mussel (Lithophaga lithophaga)
Figure 2. Average expression stability (M-value) of reference genes evaluated by geNorm.
Fig. 1 in New data on the ichthyosaur Platypterygius hercynicus and its implications for the validity of the genus
Fig. 1. Location of MHNH 2010.4 in Saint−Jouin, France.
Dissimilarity-adaptive cross-validation experiments and datasets
<p>This data includes all datasets and codes for implementing dissimilarity-adaptive cross-validation experiments, datasets and code, Reademe.txt explains each file's meaning. Appendix includes the descriptions of datasets. </p>
Python functions -- cross-validation methods from a data-driven perspective
<p>This is the organized python functions of proposed methods in Yanwen Wang PhD research. Researchers can directly use these functions to conduct spatial+ cross-validation (SP-CV), dissimilarity quantification by adversarial validation (AVD), and dissimilarity-adaptive cross-validation (DA-CV). The description of how to run codes is in Readme.txt. The descriptions of functions are in functions.docx.</p>
Model and Data for the T&C-CROP Validation Paper: T&C-CROP: Representing mechanistic crop growth with a terrestrial biosphere model (T&C,v1.5): Model formulation and validation.
<p>Here included is the code used to run T&C-CROP as used for the GMD paper submission alongside with the necessary weather data and raw field data used as part of the validation exercise. </p> <p> </p>
Troodontid specimens from the Cretaceous Two Medicine Formation of Montana (USA) and the validity of Troodon formosus
<p>Supplementary analysis files for Varricchio et al. 'Troodontid specimens from the Cretaceous Two Medicine Formation of Montana (USA) and the validity of <em>Troodon formosus</em>'.</p>
Data/Code for: Sediment dynamics in the energetic nearshore zone: Acoustic remote sensing and model validation
<p>This archive contains data and postprocessed results used in the article "Sediment dynamics in the energetic nearshore zone: Acoustic remote sensing and model validation" by G. Wilson, P. Dickhudt & J. Aldrich. All data and code in this archive is copyright of the authors. Please contact the authors prior to publishing new results or derivative works based on data/code from this archive.</p>
Script and data from: The best of two worlds: toward large-scale monitoring of biodiversity combining metabarcoding and optimised parataxonomic validation.
<h2>Description</h2> <div> <p>Zenodo linked to : Penel, B., Meynard, C.N., Benoit, L., Bourdonné, A., Clamens, A., Soldati, L., Migeon, A., Chapuis, M.-P., Piry, S., Kergoat, G. and Haran, J. (2025), The best of two worlds: toward large-scale monitoring of biodiversity combining COI metabarcoding and optimized parataxonomic validation. Ecography, 2025: e07699. <a href="https://doi.org/10.1111/ecog.07699">https://doi.org/10.1111/ecog.07699</a></p> <div> <div> <div> <div> <p><strong>Publication abstract </strong></p> </div> </div> </div> <p>In a context of unprecedented insect decline, it is critical to have reliable monitoring tools to measure species diversity and their dynamic at large-scales. High-throughput DNA-based identification methods, and particularly metabarcoding, were proposed as an effective way to reach this aim. However, these identification methods are subject to multiple technical limitations, resulting in unavoidable false-positive and false-negative species detection. Moreover, metabarcoding does not allow a reliable estimation of species abundance in a given sample, which is key to document and detect population declines or range shifts at large scales. To overcome these obstacles, we propose here a Human-Assisted Molecular Identification (HAMI) approach, a framework based on a combination of metabarcoding and image-based parataxonomic validation of outputs and recording of abundance. We assessed the advantages of using HAMI over the exclusive use of a metabarcoding approach by examining 492 mixed beetle samples from a biodiversity monitoring initiative conducted throughout France. On average, 23% of the species are missed when relying exclusively on metabarcoding, this percent being consistently higher in species-rich samples. Importantly, on average, 20% of the species identified by molecular-only approaches correspond to false positives linked to cross-sample contaminations or mis-identified barcode sequences in databases. The combination of molecular methodologies and parataxonomic validation in HAMI significantly reduces the intrinsic biases of metabarcoding and recovers reliable abundance data. This approach also enables users to engage in a virtuous circle of database improvement through the identification of specimens associated with missing or incorrectly assigned barcodes. As such, HAMI fills an important gap in the toolbox available for fast and reliable biodiversity monitoring at large scales.</p> <div> <h4><strong>File description: </strong></h4> <h4>MiSeq raw sequences of the COI barcode from 492 Coleoptera field samples :</h4> <div>The Raw_sequencage_data ZIP directory contains the FASTQ files of the paired-end reads (R1: reads 1; R2: reads 2) produced for each Coleoptera field samples in duplicate using the MiSeq platform GenSeq (ISEM - University of Montpellier)</div> <div> </div> <div>The HAMI_data_script_results_R zip directory contains Rmarkdown script files (.Rmd and .html) and associated data used to analyse the systemic errors of the metabarcoding approach (N= 492 Coleoptera field samples).</div> <div> </div> <div>The HAMI_pipeline zip directory contains all the codes associated with the HAMI pipeline, as well as a ReadMe file and a test data set.</div> <div> </div> <div>The Residual_chimera.zip directory contains lists of MOTUs associated to residual chimeric sequences that were not filtered using FROGS pipeline but secondarily detected with the <em>de novo</em> approach implemented in HAMI pipeline with ‘isBimeraDenovo’ R function from DADA2 v1.28.0. It contains two distinct files according to the two sequencing runs.</div> <div> </div> <div>The NUMTS_filtered.zip directory contains lists of MOTUs that were excluded of the final dataset according to the NUMTS filtering. File xxx_pseudogene_f1_deteled.csv corresponds to MOTUs that were excluded according to the first filtrering step based on DNA sequencing. File xxx_pseudogene_f2_deteled.csv corresponds to merged MOTUs that were excluded according to the second filter based on occurrence and percentage of identity. This folder contains files for the two sequencing runs.</div> </div> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.