Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,956

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,956 results for “test data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data for: Testing the mettle of METAL: A comparison of phylogenomic methods using a challenging but well-resolved phylogeny

<p>Sequence alignments and gene trees for: Braun et al. "Testing the mettle of METAL: A comparison of phylogenomic methods using a challenging but well-resolved phylogeny." See README file for detailed description of all files.</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

Data from: Use of the lung flute ECO to assist in sputum collection for tuberculosis testing: a randomized crossover trial

<p>The Lung Flute ECO, a self-powered, low cost, oscillatory positive expiratory pressure (OPEP) device, assisted people with presumptive tuberculosis to produce an adequate sputum volume for diagnostic testing and was well-tolerated.</p>

opencc-zeroMar 2024View details →
zenodo36/100

Training and test data for antibody humanness evaluation

<p>### Training and test data for humanness evaluation</p> <p>This data was collected in conjunction with and used for<br>training and testing for Parkinson / Wang et al 2024. The<br>data is organized as follows:</p> <p>- Heavy chain training and multispecies test data (under the heavy chain folder)<br>&nbsp; &nbsp; - The conslidated cAb rep file contains training human sequences<br>&nbsp; &nbsp; - The test sample sequences folder contains fasta files with test sequences for each species<br>- Light chain training and multispecies test data (under the light chain folder)<br>&nbsp; &nbsp; - The conslidated cAb rep file contains training human sequences<br>&nbsp; &nbsp; - The test sample sequences folder contains fasta files with test sequences for each species<br>- Abybank data (under the abybank compiled data folder)<br>&nbsp; &nbsp; - This folder contains separate folders for heavy and light chain<br>&nbsp; &nbsp; - Each subfolder contains test data for a more diverse species set under fasta files for each species<br>- Humanization test data (under the humanization test data folder)<br>&nbsp; &nbsp; - The sequences in the parental.fa file were originally humanized as part of drug discovery programs<br>&nbsp; &nbsp; - The experimental.fa file contains the humanization results<br>- IMGT and ADA data (under the imgt test data folder)<br>&nbsp; &nbsp; - The imgt mab db fa and tsv files contain sequences and species assignments for IMGT mAb DB<br>&nbsp; &nbsp; - The thera ada fa file contains sequences evaluated in the clinic<br>&nbsp; &nbsp; - The Therapeutic ADA txt file contains anti drug antibody results for those antibodies<br>- VDJ statistics (under the vdj_statistics_eval folder)</p> <p>The data was retrieved from the following sources.</p> <p>1. All heavy and light chain training data is from the cAb-Rep database from [Guo et al.](https://pubmed.ncbi.nlm.nih.gov/31649674/)<br>2. All testing data is from the Observed Antibody Space [(OAS) database](https://opig.stats.ox.ac.uk/webapps/oas/)</p> <p>The training and test data show is after filtering for quality. The testing data was additionally randomly sampled to yield a set of 50,000 sequences for each species, then filtered to remove duplicates. The human test data was checked to ensure no overlap with the human training set.</p> <p><br>The IMGT, ADA and humanization test data was retrieved from Prihoda et al. and<br>the associated [Github repo](https://github.com/Merck/BioPhi-2021-publication).</p> <p>See Parkinson et al. 2024 and the associated github repos for more details on how models other than<br>SAM / AntPack were evaluated on this data.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Data to "Cardiopulmonary exercise testing complements both spirometry and nuclear imaging for assessing sarcoidosis disease stage and for monitoring disease activity"

<p>This record contains analysis scripts (written in Matlab) as well as raw and processed data to reproduce the results shown in:</p> <p>Torregiani, C., Reale, M., Confalonieri, M., Dore, F., Crisafulli, C., Baratella, E., ... &amp; Maiello, G. (2024). Cardiopulmonary exercise testing complements both spirometry and nuclear imaging for assessing sarcoidosis stage and for monitoring disease activity.&nbsp;<em>Sarcoidosis, Vasculitis, and Diffuse Lung Diseases</em>, 41(1), e2024017-e2024017. https://doi.org/10.36141/svdld.v41i1.15125</p>

opencc-by-4.0Jan 2024View details →
zenodo36/100

Raw data of surface texture parameter values (incl. output of statistical tests).

<p>This is the raw data (supplement 3) of the publication with the title "Prey size reflected in tooth wear &ndash; a comparison of two wolf populations from Sweden and Alaska" written by: Ellen Schulz-Kornas, Mirella H. Skiba and Thomas M. Kaiser and accepted for publication as manuscript RSFS-2023-0070.R1 in the journal Interface Focus on 2-April 2024.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

ccAFv2 test data

<p>This is a test dataset for the ccAFv2 classifier. It is originally from the manuscript O'Connor et al., 2021 in Molecular Systems Biology (<a href="https://pubmed.ncbi.nlm.nih.gov/34101353/">https://pubmed.ncbi.nlm.nih.gov/34101353/</a>).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

TRPV1 calcium imaging assay data from testing of capsaicinoids

<p>irTRPV1 calcium imaging assay data collected in HTS format included legend to files, raw data, feature files with description, aligned data and final GraphPad file. Correspoinding paper's DOI will be added once generated.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

hydroflows-test-data

<p>test datasets for hydroflows</p> <p>for individual data licenses see the data_catalog.yml file</p> <p>v0.1.10 Update sfincs-model files to be compatible with hydromt_fiat v0.5.3 and fiat-toolbox v0.1.16</p> <p>v0.1.9 Update CMIP6 derived data</p> <p>v0.1.8 Add Aggregation areas in Delft-FIAT model</p> <p>v0.1.7 Add CMIP6 derived stats</p> <p>v0.1.6 Add CMIP6 data</p> <p>v0.1.5 Add Coast-RP data</p> <p>v0.1.4 Update wflow test data</p> <p>v0.1.3: Added Merit-Basins and updated GPEX</p> <p>v0.1.2: Added model simulations (missing global-data.tar.gz)</p> <p>v0.1.1: Added GPEX &amp; GTSM data</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Data: Another Simple Test for the Presence of Multidomain Behaviour During Palaeointensity Experiments

<p>SD_thresh_Out.csv - Data generated from simulations of IZZI protocol palaeointensity experiments.&nbsp; The simulations were run with single domain type behaviour.&nbsp; Naming conventions for parameters are similar to those found on the Standard Paleointensity Definitions website (https://earthref.org/PmagPy/SPD/home.html).</p> <p>MD_thresh_Out.csv - Data generated from simulations of IZZI protocol palaeointensity experiments.&nbsp; The simulations were run with multi-domain type behaviour.&nbsp; Naming conventions for parameters are similar to those found on the Standard Palaeointensity Definitions website (https://earthref.org/PmagPy/SPD/home.html).</p> <p>AraiPlot_XY_HighMD.txt - The X and Y co-ordinate points for a simulated Arai plot with a high degree of non-ideal behaviour.</p> <p>parameter_noise_data.csv - Zig-Zag parameters calculated for a single Arai Plot with varying amounts of noise added.&nbsp; The noise is sampled from a gaussian distribution with variances taken from <a href="https://doi.org/10.1029/2012GC004046">https://doi.org/10.1029/2012GC004046</a>.&nbsp; The parameters broad_noise_TRM/NRM and narrow_noise_TRM/NRM give the noise applied to the x and y points,&nbsp; broad and narrow denote the type of unblocking behaviour. The x and y points for the original Arai plot are given in x/y_noise_free.&nbsp; The zig-zag statistics for the original Arai plot are given by parameters {statistic}_orig.&nbsp; The suffix is then changed to denote whether the statistic was calculated with broad or narrow blocking noise applied.&nbsp; The additional suffix _norm indicates that the value has been normalised by {statistic}_orig.&nbsp;</p> <p>zigzag_delta.csv - Contains zig-zag statistics calculated for a straigth Arai plot, which is then incrementally given a zig-zag by displacing points perpendicularly to the oringal Arai plot.&nbsp; X and Y points are given along with the amount of displacement (delta) and the zig-zag statistics.</p> <p>CurvatureData.csv - Contains zig-zag statistics calculated for an increasingly curved Arai plot with no zig-zag present.&nbsp; The Arai plot points, zig-zag statistics and the curvature are included.</p> <p>ExperimentalData.csv - Selected statistics for experimental data where the expected field is known.&nbsp; Includes the accuracy, natural log of the accuray (Acc), basic selection statistics and zig-zag selection statistics.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

IWC : Test Data For VGP Decontamination Workflow

<p>Dataset used to test the decontamination workflow published with the VGP assembly pipeline in Galaxy.&nbsp;</p>

openmit-licenseAug 2022View details →
zenodo36/100

Data about upper limb kinematics to test the efficacy of Action Observation Treatment

<p>The dataset contains the upper-limb kinematics collected in 40 participants before and after a short-term immobilization of the right arm. Three reach-to-grasp movements were required, targeting an object placed a) frontally at the level of the shoulders (A_low), b) frontally above the head (A_high), or c) laterally at the level of the shoulders (L_low). These movements were intended to test different degrees of freedom of the shoulder joint.</p> <p>Half of the participants underwent a virtual-reality action observation treatment (see Rizzolatti et al, Neuroscience and Biobehabioral Reviews, 2021), while the control group underwent a non-motor VR stimulation.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Source data for testing results of the ADS software

<p>This is the source data for the testing results of the ADS software at&nbsp;https://zenodo.org/record/5579390</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
dryad36/100

Data from: N-mixture models estimate abundance reliably: a field test on Marsh Tit using time-for-space substitution

<p>Imperfect detection in field studies on animal abundance, including birds, is common and can be corrected for in various ways. The binomial N-mixture (hereafter binmix) model developed for this task is widely used in ecological studies owing to its simplicity: it requires replicated count results as the input. However, it may overestimate abundance and be sensitive to even small violations of its assumptions. We used a 33-year dataset on the Marsh Tit, Poecile palustris, a sedentary forest passerine, from Białowieża Forest, Poland to validate inference from binmix models by comparing model-estimated abundances to the true number of breeding pairs within the plots, determined by exhaustive population study. The abundance estimates, derived from six springtime (April-May) counts of males on each plot in each year, were highly reliable: 116 out of 132 year-plot estimates (88%) included the true number of pairs within the 95% confidence intervals. Over- and underestimations were thus rare and similarly frequent (9 and 12 cases, respectively), with a tendency to overestimate at low densities and underestimate at high densities. Marsh Tits sing rarely but the frequency of countersinging increases with abundance, leading to non-independence in detections. When accounted for in a submodel for detection, the per-survey number of countersinging events positively affected detection probability but only weakly affected abundance estimates. Simulations further demonstrate that this property, overestimation at low densities and underestimation at high densities, may be a systematic bias of binmix model even if density-dependent detection is absent. While the behaviour of binmix models in specific situations requires more study, we conclude that these models are a valid tool to estimate abundance reliably when intensive population monitoring is not feasible.</p>

opencc-zeroNov 2021View details →
zenodo36/100

Irregular data for testing non-tabular data reading and parsing

<p>key1#value1,value2,...,valueN</p> <p>key2#value1,value2,...,valueM</p> <p>Generated with gen_confus.jl from&nbsp;https://github.com/abelsiqueira/call-julia-from-python-experiments</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Wilcoxon Rank Sum Test and Keyphrase Extraction Data Cited in "What Everyone Says: Public Perceptions of the Humanities in the Media"

<p>This repository contains Wilcoxon rank sum test and keyphrase extraction data cited in the WhatEvery1Says (WE1S) Project&#39;s article &nbsp;&quot;What Everyone Says: Public Perceptions of the Humanities in the Media&quot;. The organization of the materials is discussed below.</p> <p><strong>Wilcoxon Rank Sum Test</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>wilcoxon-tests</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. The Wilcoxon rank sum test identifies specific words that appear significantly more in one group of documents as compared to another, thus providing researchers with an understanding of what words are &ldquo;distinctive&rdquo; to each group. Further information on WE1S&#39;s use of Wilcoxon rank sum testing can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf</a>.</p> <p>Each subdirectory in the <code>wilcoxon-test</code> folder contains the data and results of a particular comparison experiment based on a metadata category such as whether the data contained articles published by public or private institutions. Each data file is a <code>.txt</code> file representing a sample of the overall data from the collection. The <code>README</code> file provides information on the collection used, the sample size, and the nature of the comparison. The results for the test are in a file called <code>results.csv</code>.</p> <p>The <code>results.csv</code> file for each test includes a row for each term included in the test. Each row displays the term, the term&#39;s raw count in each category compared (count 1 and count 2), the difference between the 2 counts (count 1 minus count 2), the percentage change in the counts, the Wilcoxon statistic, and the Wilcoxon p-value. Sorting the csv by the Wilcoxon stat from greatest to least will cause the terms most strongly associated with category 1 to come to the top (category 1 is the category listed first in the title field of the README.md file for each test), while sorting it by the Wilcoxon stat from least to greatest will cause the terms most strongly associated with category 2 to come to the top (category 2 is the category listed second). The p-value column provides you with information about how confident you can be about each comparison&#39;s significance.</p> <p><strong>Keyphrase Extraction</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>keyphrase-extraction</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. Keyphrase extraction generates a list of the most significant words or phrases (1-6 words long) within individual documents. WE1S takes the top ten keyphrases in each document and ranks them according to their frequency across the collection. WE1S uses the SGRank algorithm for keyphrase extraction, and because this algorithm is computationally intensive, WE1S limits keyphrases to lemmatized nouns and proper nouns within a window of 70 words to either side of candidate keyphrases. Further information on WE1S&#39;s use of keyphrase extaction can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf</a>.</p> <p>Each subdirectory in the <code>keyphrase-extraction</code> folder contains the data and results of keyphrase extraction on a particular collection. Details of the collection and resulting files can be found in each subdirectory. Each list of keyphrases is in a file called <code>SGRank.csv</code>, which lists the keyphrases and their number of occurrences in the collection. The article additionally cites keyphrases that are shared with the terms in the public topic model produced by Andrew Goldstone and Ted Underwood, &ldquo;The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,&rdquo; <em>New Literary History</em> 45, no. 3 (2014): 359&ndash;84, <a href="https://doi.org/10.1353/nlh.2014.0025">https://doi.org/10.1353/nlh.2014.0025</a>. The list of terms is derived from the public visualization at <a href="https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words">https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words</a>. Keyphrases extracted from WE1S data were split into single-word terms and compared with the list of vocabulary in Goldstone and Underwood&#39;s word list (<code>quiet_transformations_wordlist.txt</code>) to compile lists of shared vocabulary. These lists are given in files called <code>shared_terms.txt</code>.</p> <p>Note that keyphrases were extracted for corpora produced using the Python <a href="https://textacy.readthedocs.io/en/latest/index.html">Textacy</a> library. Because these corpora contain the full text of articles with intellectual property restrictions they cannot be reproduced here.</p>

opencc-by-sa-4.0Jul 2021View details →
zenodo36/100

Data on plant health stakeholder priorities for tests and general prioritisation framework

<p>Data collected in the framework of work package 4 of the Valitest project. They correspond to a qualitative assessment of plant health stakeholder requirements, have been collected using online surveys supplemented by desk-based research, as well as impact assessments</p>

opencc-by-4.0Dec 2021View details →
dryad36/100

Data from: Nutritional challenges of feeding a mutualist: testing for a nutrient-toxin tradeoff in fungus-farming leafcutter ants

<p>The biochemical heterogeneity of food items often yields tradeoffs as each bite of food tends to contain some nutrients in surplus and others in deficit, as well as other less palatable or even toxic compounds. These multidimensional nutritional challenges are likely compounded when foraged foods are used to provision others (<i>e.g</i>. offspring or symbionts) with different physiological needs and tolerances. We explored these challenges in free-ranging colonies of leafcutter ants that navigate a diverse tropical forest to collect plant fragments they use to provision a co-evolved fungal cultivar. We tested the prediction that leafcutter farmers face provisioning tradeoffs between the nutritional quality and concentration of toxic tannins in foraged plant fragments. Chemical analyses of plant fragments sampled from the mandibles of Panamanian <i>Atta colombica </i>leafcutter ants provided little support for a nutrient-tannin foraging tradeoff. First, colonies foraged for plant fragments ranging widely in tannin concentration. Second, high tannin levels did not appear to restrict colonies from selecting plant fragments with blends of protein and carbohydrates that maximized cultivar performance when measured with <i>in vitro </i>experiments. We also tested whether tannins expand the realized nutritional niche selected by leafcutter ants into high-protein dimensions since: 1) tannins can bind proteins and reduce their accessibility during digestion, and 2) <i>in vitro</i> experiments have shown that excess protein provisioning reduces cultivar performance. Contrary to this hypothesis, the most protein-rich plant fragments did not have highest tannin levels. More generally, the approach developed here can be used to test how multidimensional interactions between nutrients and toxins shape the costs and benefits of providing care to offspring or symbionts.</p>

opencc-zeroJan 2022View details →
zenodo36/100

GCM Filters Test Data

<p>Test datasets for the GCM Filters test suite.</p> <p>https://github.com/ocean-eddy-cpt/gcm-filters/</p> <p>POP data were extracted from a single timestep of the simulations described in the following paper</p> <p>Small, R. J., et al. (2014),&nbsp;A new synoptic scale resolving global climate simulation using the Community Earth System Model,&nbsp;<em>J. Adv. Model. Earth Syst.</em>,&nbsp;6,&nbsp;1065&ndash;&nbsp;1094, doi:<a href="https://doi.org/10.1002/2014MS000363">10.1002/2014MS000363</a>.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Data from Test Performance Study Euphresco project 2019-A-327

<p>Data of the test performance study organised in the framework of the Euphresco project 2019-A-327 &#39;Validation of molecular tests for the detection of tomato brown rugose fruit virus<em> </em>(ToBRFV) in seed of tomato and pepper&#39;</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Test data for fake_surface

<p>Test data for a fake atmospheric model consisting of an atmosphere module and a land-surface module. The atmosphere module reads the test data and feeds it to the land-surface module. The land surface module calculates fluxes of radiation, sensible and latent heat into the atmosphere.</p> <p>The data was generated using the MERRA-2 reanalysis dataset: https://gmao.gsfc.nasa.gov/reanalysis/MERRA-2/</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record