Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,956

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,956 results for “test data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data for testing the SIAMCAT R package

<p>Datasets needed for the vignettes of the <a href="https://siamcat.embl.de/">SIAMCAT</a> R package</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Test data in spreadsheet format

<p>Synthetic test data for a tutorial that explains how to convert spreadsheet data to tidy data.</p> <p> </p>

opencc-by-4.0Oct 2017View details →
zenodo36/100

Data for: "Different dry-wet pulses favor different functional strategies: a test using tropical dry forest tree species" by Vega-Ramos, Flor, Cifuentes Gómez, Lucas, Pineda-García, Fernando, Dawson, Todd, Paz, Horacio

<p>These data represent those published in "Different dry-wet pulses favor different functional strategies: a test using tropical dry forest tree species" by Vega-Ramos, Flor, Cifuentes G&oacute;mez, Lucas, Pineda-Garc&iacute;a, Fernando, Dawson, Todd, Paz, Horacio</p>

opencc-by-sa-4.0Apr 2024View details →
dryad36/100

Data from: Testing frameworks for early life effects: The developmental constraints and adaptive response hypotheses do not explain key fertility outcomes in wild female baboons

<p>In evolutionary ecology, two classes of explanations are frequently invoked to explain "early life effects" on adult outcomes. Developmental constraints (DC) explanations contend that costs of early adversity arise from limitations adversity places on optimal development. Adaptive response (AR) hypotheses propose that later life outcomes will be worse when early and adult environments are poorly "matched." Here, we use recently proposed mathematical definitions for these hypotheses and a quadratic-regression based approach to test the long-term consequences of variation in developmental environments on fertility in wild baboons. We evaluate whether low rainfall and/or dominance rank during development predict three female fertility measures in adulthood, and whether any observed relationships are consistent with DC and/or AR. Neither rainfall during development nor the difference between rainfall in development and adulthood predicted any fertility measures. Females who were low-ranking during development had an elevated risk of losing infants later in life, and greater change in rank between development and adulthood predicted greater risk of infant loss. However, both effects were statistically marginal and consistent with alternative explanations, including adult environmental quality effects. Consequently, our data do not provide compelling support for either of these common explanations for the evolution of early life effects.</p>

opencc-zeroApr 2024View details →
zenodo36/100

Pilot 1 Model-based decision support for testing drought-related adaptation strategies in the Aa of Weerijs river basin, the Netherlands: Hydrological model description, input data sources and model results

<p>This dataset contains:&nbsp; the report with the description of the model structure, the input data sources and the spatial locations within the catchment for which surface&nbsp; and groundwater results data are provided.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Final data for paper "A nonperturbative test of nucleation calculations for strong phase transitions"

<p>In this Zenodo deposit we include the data used to make three key figures in our paper, showing our final nucleation rate results. The columns of the raw data files are labelled. In each case, the quantity being tabulated is the logarithm of the nucleation rate, typically labelled&nbsp;<em>logGamma</em>&nbsp;in our data files. We also include plotting scripts to generate the figures from the data.</p> <ul> <li> <p>The directory&nbsp;<code>a_limit</code>&nbsp;contains the extrapolation to zero lattice spacing, with the raw data in the file&nbsp;<code>a_limit_data</code>. The temperature&nbsp;<em>T</em>&nbsp;is fixed to our benchmark point of 93.121 GeV. The linear size&nbsp;<em>Nx</em>&nbsp;and the lattice spacing&nbsp;<em>a</em>&nbsp;are varied to give a constant total volume. The file&nbsp;<code>a_limit_lin_op_data</code>&nbsp;contains our comparison measurement using the linear order parameter. The script&nbsp;<code>plot_a_limit.py</code>&nbsp;performs a nonlinear least squares fit to a cubic, writes the continuum extrapolated nuclation rate result to stdout, and plots the data and the fit.</p> </li> <li> <p>The directory&nbsp;<code>vol_limit</code>&nbsp;contains the extrapolation to infinite volume, with the raw data in the file&nbsp;<code>vol_limit_data</code>. Here the temperature&nbsp;<em>T</em>=93.121 GeV and lattice spacing&nbsp;<em>a</em>=1.5 are fixed, while the linear size&nbsp;<em>Nx</em>&nbsp;is varied. The script&nbsp;<code>plot_vol_limit.py</code>&nbsp;performs a nonlinear least squares fit to an exponential function, writes the infinite volume limit nucleation rate result to stdout, and again plots the data and the fit.</p> </li> <li> <p>The directory&nbsp;<code>rate_reweighted</code>&nbsp;contains the data needed to plot the nucleation rate as a function of temperature. This plot combines lattice data and various perturbative estimates of the nucleation rate. The lattice data are in three files:</p> <ul> <li><code>rate_lattice_BM2_vol_limit</code>&nbsp;contains the infinite volume limit result at the simulated lattice temperature&nbsp;<em>T</em>=93.121 GeV and with lattice spacing&nbsp;<em>a</em>=1.5.</li> <li><code>rate_lattice_BM2</code>&nbsp;reweights the lattice data to different temperatures and then takes the infinite volume limit.</li> <li><code>rate_lattice_BM2_Nx40_a1.5</code>&nbsp;contains data reweighted from the simulation at fixed volume&nbsp;<em>Nx</em>=40. In each lattice data file, the temperature, continuum potential parameters and nucleation rate are given.</li> </ul> <p>The perturbative results are in three files,&nbsp;<code>rate_perturbative_BM2_rg_low</code>,&nbsp;<code>rate_perturbative_BM2</code>&nbsp;and&nbsp;<code>rate_perturbative_BM2_rg_high</code>, corresponding to the three RG scales mentioned in the text. These files again contain the temperature and continuum lattice parameters. The column&nbsp;<em>eps</em>&nbsp;gives the dimensionless loop expansion parameter around the metastable phase within the EFT. The column&nbsp;<em>logGamma_0</em>&nbsp;gives the tree level result,&nbsp;<em>logGamma_A</em>&nbsp;and&nbsp;<em>logGamma_B</em>&nbsp;the LPA results with the two options for dealing with the imaginary parts described in the text, and&nbsp;<em>logGamma_1</em>&nbsp;is the one loop result. The script&nbsp;<code>plot_rate_reweighted.py</code>&nbsp;plots all of these data together. For the tree level and one loop cases, the error bands plotted correspond to the minimum and maximum values of the nucleation rate across the three RG scales. For the LPA results, the bands include the extreme values for all six options including the two approaches to handling the imaginary parts.</p> </li> </ul> <p>In each case, the plots are saved to PDF and PNG plot files.</p> <p>The figures used in this deposit and in the paper used the following package versions (obtained with the&nbsp;<code>pipreqs</code>&nbsp;package):</p> <pre><code>matplotlib==3.5.1 numpy==1.21.5 scipy==1.8.0 seaborn==0.11.2</code></pre> <p>Note that in <a href="https://doi.org/10.5281/zenodo.10891524">Version v1</a> there was an error in the normalisation of our perturbative tree-level and LPA results which has now been fixed.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Data for "Testing for variation in photoperiodic plasticity in a butterfly: inconsistent effects of circadian genes between geographic scales" (Ecology and Evolution, accepted manuscript)

<p>Raw data and analysis scripts belonging to <em>Testing for variation in photoperiodic plasticity in a butterfly: inconsistent effects of circadian genes between geographic scales</em> (Ecology and Evolution, accepted manuscript).</p> <p>Includes data from a photoperiodic assay with larvae of the speckled wood butterfly,&nbsp;<em>Pararge aegeria</em>, as well as a small bioinformatic analysis of SNP variation in two circadian candidate genes.</p> <p>List of files:</p> <table> <tbody> <tr> <td>statistics_and_figures.R</td> <td>Statistical analysis of phenotyping experiment; drawing figures</td> </tr> <tr> <td>variant_calling.sh</td> <td>Shell script for mapping sequencing reads; calling and tabulating SNPs</td> </tr> <tr> <td>experiment_diapause.txt</td> <td>Data from phenotyping experiment, for analysis of diapause induction</td> </tr> <tr> <td>experiment_larval</td> <td>Data from phenotyping experiment, for analysis of larval development</td> </tr> <tr> <td>snptable_timeless</td> <td>Tabulated allele frequencies for exonic SNPs in timeless</td> </tr> <tr> <td>snptable_period</td> <td>Tabulated allele frequencies for exonic SNPs in period</td> </tr> </tbody> </table>

opencc-by-4.0May 2024View details →
zenodo36/100

Benchmark datasets for testing AIRR-seq data processing with pyIgMap pipeline

<p>This is a set of raw FASTQ files produced by various AIRR-seq protocols, stored here for benchmark conveinience and for reference purposes.</p> <p>Currently it contains the following datasets:</p> <table> <tbody> <tr> <td><strong>id</strong></td> <td><strong>fastq</strong></td> <td><strong>reference</strong></td> <td><strong>method</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>allergy</td> <td> <p>ERR7425614_1.fastq.gz,</p> <p>ERR7425614_2.fastq.gz</p> </td> <td>https://doi.org/10.7554/eLife.79254</td> <td>5'RACE, UMI, long MiSeq reads, IGH w/ isotype</td> <td>Longitudinal full-length IGH repertoire profiling and clonal lineage dynamics in memory B cells, plasmablasts and plasma cells of human peripheral blood</td> </tr> <tr> <td>covid</td> <td> <p>fmba_TRAB_R1.fastq.gz,</p> <p>fmba_TRAB_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1101/2023.11.08.566227</p> </td> <td>DNA multiplex, UMI, TRA+TRB mix, NextSeq</td> <td>TCR sequencing in COVID-19 convalescent and healthy donors</td> </tr> <tr> <td>natprot</td> <td> <p>PMID27490633_R1.fastq.gz,</p> <p>PMID27490633_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/nprot.2016.093</p> </td> <td>5'RACE, UMI, long MiSeq reads, high-quality overlap, IGH no isotype</td> <td>High-quality full-length immunoglobulin profiling with unique molecular barcoding</td> </tr> <tr> <td>brnaseq</td> <td> <p>SRR3743469_R1.fastq.gz,</p> <p>SRR3743469_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1016/j.immuni.2016.08.012</p> </td> <td>RNA-Seq, all chains, B-cells</td> <td>Primary human mature na&iuml;ve B-cells (IgD+CD38lo; NB) and GCB-cells (CD77+CD38hi; GCB) were purified from tonsils of healthy individuals. RNA-seq libraries were prepared using the Illumina TruSeq RNA sample kits according to the manufacturer.</td> </tr> <tr> <td>uhrr</td> <td> <p>UHRR_full_R1.fastq.gz,</p> <p>UHRR_full_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/s41598-021-04583-z</p> </td> <td>RNA-Seq, bulk</td> <td>Universal Human Reference RNA</td> </tr> <tr> <td>10x</td> <td> <p>10x_bcr_R1.fastq.gz,</p> <p>10x_bcr_R2.fastq.gz,</p> <p>10x_tcr_R1.fastq.gz,</p> <p>10x_tcr_R2.fastq.gz</p> </td> <td> <p>https://www.10xgenomics.com/datasets/human-pbmc-from-a-healthy-donor-10-k-cells-v-2-2-standard-5-0-0</p> </td> <td>10x Genomics vdj</td> <td>See reference. AIRR-seq is split into TCR and BCR parts</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Additional raw video and pose estimation data of top view mouse behavior recordings (marble burying test, light-dark box, fear conditioning box) of acute and chronic stress models

<p>This repository contains raw data for 296 different behavioral recordings of mice (marble burying test, light-dark box, fear conditioning box). These include top view raw video .mp4 files (Videos.zip) and the corresponding .csv pose estimation data (data.zip) obtained with DeepLabCut. The data is from multiple different experiments. The METADATA.csv or METADATA.xlsx files contain all grouping variables and help linking the pose estimation files (located in multiple subfolders of /data) to the video files. Visit https://github.com/ETHZ-INS/BehaviorFlow to find out more about how this data has be used by us.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Data from: Understanding Test Convention Consistency as a Dimension of Test Quality

<div> <div>This archive provides additional data for the article "Understanding Test Convention Consistency</div> <div>as a Dimension of Test Quality" by Martin P. Robillard, Mathieu Nassif, and Muhammad Sohail,</div> <div>published in ACM Transactions on Software Engineering and Methodology.</div> </div>

opencc-by-4.0May 2024View details →
zenodo36/100

geNomad minimal test data

<p>Very small reference database for geNomad derived from version 1.7 as described here&nbsp;https://github.com/apcamargo/genomad/issues/104</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Simulated Cucurbit Tobamovirus Testing Dataset for Illustrating Possible Data Format Only

<p>This dataset contains wholly simulated and not particularly realistic data. It is intended to demonstrate the format of a possible dataset of tobamovirus testing results for cucurbit imports into Australia.</p> <p>If permission is obtained to publish real data, the real data will contain de-identified data on testing for virus presence in cucurbit seed lots where importation into Australia was attempted, over a period between 2019 and 2021. Each row represents a seed lot. For each lot, a number of groups of 400 seeds were tested for tobamovirus contamination, returning a presence or absence result for each group.There is no identifier or name representing the lot, consignment, importer or source country. The ID variable on the file was randomly assigned.</p> <p>Other authors will be added to the real data page subject to their permission and preference.</p> <p>For more information on the data, see&nbsp;</p> <p>Dall DJ, Lovelock DA, Penrose LDJ, Constable FE. Prevalences of Tobamovirus Contamination in Seed Lots of Tomato and Capsicum.&nbsp;<em>Viruses</em>. 2023; 15(4):883. https://doi.org/10.3390/v15040883</p> <p>&nbsp;</p> <p><strong>Variable Definitions</strong></p> <table> <tbody> <tr> <td> <p>Variable Name</p> </td> <td> <p>Values</p> </td> <td> <p>Description</p> </td> </tr> <tr> <td> <p>ID</p> </td> <td> <p>1, &hellip;, 800 (will be 793 in the real data)</p> </td> <td> <p>Randomised lot identifier number. The data set contains one row per lot.</p> </td> </tr> <tr> <td> <p>species</p> </td> <td> <p>tomato, capsicum</p> </td> <td> <p>species of cucurbit being imported</p> </td> </tr> <tr> <td> <p>year</p> </td> <td>2019, 2020, 2021</td> <td> <p>year in which importation was attempted</p> </td> </tr> <tr> <td> <p>month</p> </td> <td> <p>1, &hellip;, 12</p> </td> <td> <p>month in which importation was attempted</p> </td> </tr> <tr> <td> <p>NUM_TESTED</p> </td> <td> <p>1, ...., 50</p> </td> <td> <p>Number of groups tested. More groups were sampled from larger seed lots.</p> </td> </tr> <tr> <td>NUM_POS</td> <td>0, ..., 50</td> <td>Number of groups where virus contamination was detected.</td> </tr> <tr> <td> <p>vspecies</p> </td> <td> <p>PMMoV, PSTVd, PSTVd1, TASVd, TMV, ToBRFV, ToMMV, ToMV, blank</p> </td> <td> <p>Species of tobamovirus detected (equal to "blank", i.e., empty string, when no virus was detected.</p> </td> </tr> </tbody> </table> <p><strong>Author Field</strong></p> <p>If permission is obtained, this section will state that I am distributing this data with the permission of the data owner. (Other creators/contributers may be added to the "Creator" or "Contributer" section if they give permission for their names to appear.)</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

AMIA Pipeline Test Data

<p>The first datasets in the HIV-1C_ZA folder contains the initial Human Immuniodefeciency Virus 1C (HIV-1C) Integrase (IN) amino acid mutations associated with patient sequence data extrapoloated in a previous research project conducted by R. Chitongo (DOI: 10.1371/journal.pone.0223464). This dataset is curated in a standard Comma Separated Variable (.csv) file with each of the mutations associated with a particular patient listed benath a random identifier with no direct or indirect patient information present. These random identifiers do not correlate to a specific patient either and have just been assigned for differential purposes between which mutations were present in a particular patient. This also includes the protein homology model generated in our study present in the standar Protein Data Bank (PDB) format of the HIV-1C IN. The second variant_ouputs folder contains the various datasets and files generated from the initial phase of the AMIA pipeline which is further elaborated upon in the published literature.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Test data for Galaxy IUC `muon` tool

<p>&nbsp;Test data for Galaxy IUC `muon`&nbsp; tool.&nbsp; The data is based on published 10x human PBMC 3k multiomics data. The data was filtered for chromosome 21 only.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

stark_test_data

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
dryad36/100

Data from: Assessing temporal transition between microgranular and hyaline tests of calcareous microplankton during the Late Jurassic

<p>Calcareous microplankton increased in abundance during the latest Jurassic, coinciding with the increase in abundance of calcareous nannofossils and with the onset of deposition of pelagic calcareous oozes. However, the timing and causes of the shift from microgranular tests of the earliest microplankton (chitinoidellids) to hyaline tests of calpionellids are obscured because the ultrastructure of two-layered praecalpionellids that occur during the Tithonian is poorly documented. Here, we investigate the ultrastructure of chitinoidellids and praecalpionellids from Upper Tithonian deposits in the Western Carpathians. We show that (1) the chitinoidellid microgranular layer is formed by elongated, euhedral, densely-packed, nanometric needles rather than by fragments of calcareous nannofossils, (2) two-layered chitinoidellids (<em>Semichitinoidella</em>) are formed by an internal microgranular layer (identical to that of <em>Chitinoidella</em>) and by an external hyaline prismatic layer, and (3) two-layered <em>Praetintinnopsella</em> exhibits an internal hyaline layer (with densely-packed, equant microcrystals) and an external layer formed by a dark organic rim. The external layer in <em>Praetintinnopsella</em> thus does not have any relation to the microgranular layer in chitinoidellids and the external hyaline layer of <em>Semichitinoidella</em> is not equivalent in structure to the hyaline layer of <em>Praetintinnopsella</em>. As both single-layer and two-layered chitinoidellids appear prior to the first appearance of <em>Praetintinnopsella</em> but still co-occur with this genus in the lowermost Upper Tithonian deposits, the origin of two-layered <em>Praetintinnopsella</em> either reflects a major transformation in biomineralization towards larger and more packed crystals during their earlier divergence from the chitinoidellid lineage or an origination of two-layered tests with a hyaline layer from an independent non-chitinoidellid ancestor.</p>

opencc-zeroJul 2024View details →
zenodo36/100

Fig. 1 in Evolution and systematics of Green Bush-crickets (Orthoptera: Tettigoniidae: Tettigonia) in the Western Palaearctic: testing concordance between molecular, acoustic, and morphological data

Fig. 1 Map showing the sampling sites for Tettigonia

opencc-by-4.0Dec 2016View details →
zenodo36/100

Test data for "Imaging of cellular dynamics in vitro and in situ: from a whole organism to sub-cellular imaging with self-driving, multi-scale microscopy"

<p>This repository contains test data associated with analysis code of low- and high-resolution self-driving multi-scale data from our manuscript &nbsp;"Imaging of cellular dynamics <em>in vitro</em> and <em>in situ</em>: from a whole organism to sub-cellular imaging with self-driving, multi-scale microscopy"</p> <p>by Stephan Daetwyler, Hanieh Mazloom-Farsibaf, Felix Y. Zhou, Dagan Segal, Etai Sapoznik, Bingying Chen, Jill M. Westcott, Rolf A. Brekken, Gaudenz Danuser and Reto Fiolka</p> <p>&nbsp;</p> <p>Code repository: <a href="https://github.com/DaetwylerStephan/multi-scale-image-analysis">https://github.com/DaetwylerStephan/multi-scale-image-analysis</a></p> <p>Documentation: <a href="https://daetwylerstephan.github.io/multi-scale-image-analysis/">https://daetwylerstephan.github.io/multi-scale-image-analysis/</a></p> <p>Preprint: <a href="https://www.biorxiv.org/content/10.1101/2024.02.28.582579v1.full">https://www.biorxiv.org/content/10.1101/2024.02.28.582579v1.full</a></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Test data, do not use

<p>Chinese regional 2001-2021 annual data.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

FRS Test 2 data

<p>Before diving into the data please refer to the description of the <a href="https://mbd.pages.rwth-aachen.de/dlr_rdm/data/FRS/intro.html">FRS</a> system and the <a href="https://mbd.pages.rwth-aachen.de/dlr_rdm/">TRIPLE</a> project.</p> <p>The data represented here describes the permittivity measured with respect to depth. The measurement was carried out on 10.09.2024 in the Neumayer station, located in Antarctica, under the campaign name TRIPLE-Phase-2. The surface ice temperature was 240K, atmospheric pressure 1 bar, and 60% humidity. The readings were taken on ice as the IceCraft moved vertically downwards by melting the ice. The data obtained can also be theoretically assessed as here(link).</p> <p>Column: Depth, Vertical Distance, Spatially Varying, Permittivity</p>

opencc-zeroJul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record