Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

113

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

113 results for “VCF”

Learn how ShareScore rates datasets ↗
zenodo48/100

Genomic vcf file for D.melanogaster, Sussex LHM population

<p>Data, logs and code for genomic vcf genotypes file for the Morrow lab, D.melanogaster LHM sequencing and genotyping project.</p>

opencc-by-4.0Dec 2016View details →
zenodo44/100

Simulated Ancient Genomic Kinship Dataset: VCF and BAM (1x) Files for Related (including inbred) Pairs

<p>Simulated Ancient Genomic Kinship Dataset: VCF and BAM (1x and 5x) Files for Related (including inbred) Pairs</p><p><strong>Description:</strong></p><p>This dataset comprises simulated pedigrees (VCF files containing 8,677,101 autosomal biallelic and 298,625 X chromosomal SNP positions) generated using Ped-sim (v1.3) and comprising pairs of diverse familial relationship types up to third-degree. The first-degree relationships are parent-offspring and siblings; the second-degree relationships are half-siblings, grandparent-grandchild, and avuncular pairs; and third-degree relationships are first cousins, great-grandparent-great-grandchild, and grand avuncular pairs. For each of these 8 relationship types, our dataset includes 48 pairs of individuals. It also contains unrelated pairs. Additionally, the dataset includes first- and second-degree relatives, with inbreeding (parent-offspring pairs where the parents of the offspring are the first cousins and grandparent-grandchild pairs where the grandchild is the offspring of first cousins). Our simulations encompass all combinations of kinship types regarding sex. The dataset was further enriched by simulating ancient DNA-like sequencing data (5x and 1x BAM files) of Ped-sim simulated individuals using the gargammel tool, employing procedures akin to standard paleogenomic sequencing libraries. Note that the BAM files contain only randomly chosen 200K autosomal SNP positions. Positions can be found in the "200K_positions" file. Details can be found in Aktürk, Mapelli and Güler et al. 2023.</p><p><strong>Data Sources and Generation:</strong></p><p>Founder genotypes for pedigree simulation were created from the Tuscany (TSI) population SNPs within the 1000 Genomes Dataset v3. Notably, the founder genotypes lack background relatedness or runs of homozygosity (ROH).</p><p><strong>Description of File Naming Conventions:</strong></p><p>The naming conventions of the BAM files in this dataset are designed to convey key information regarding the specifics of each file.</p><p><strong>cov1x or cov5x:</strong> This segment denotes the coverage level of the BAM files, indicating whether the sequencing coverage for the individuals in the files is 1x or 5x.</p><p><strong>run_*:</strong> Signifies the particular batch from which the pedigree and individuals are derived. This name segment also applies to VCF files.</p><p><strong>parent-offspring_* or similar identifiers:</strong> Reflects the origin of the individual from the corresponding VCF file. For instance, "parent-offspring_1" corresponds to the individuals present in the "run_*_parent-offspring_1.vcf" file.</p><p><strong>parent-offspring* or similar identifiers:</strong></p><p>&nbsp;Indicates the origin of the individual from the sets within the VCF files. For example, "parent-offspring1" signifies the first set of parent-offspring pedigrees within the VCF file. Note that parent-offspring, grandparent-grandchild, and great-grandparent-great-grandchild and the inbreeding VCFs contain only one set, so this identifier is always 1. This convention can be 1 or 2 for the rest of the pedigrees, as the VCF files contain two sets of related pairs.</p><p><strong>_g*-b*-: </strong>Provides information about the individual's generational level within the VCF. This follows the Ped-sim syntax. For example, for parent-offspring type, "_g1-b1-" indicates the first parent (generation 1) within a specific pedigree, and "_g1-b2-" indicates the second parent (generation 1) while "_g2-b1-" represents the offspring (generation 2).</p><p><strong>Example Naming Structure:</strong></p><p>For instance, the file "cov1x_run1_parent-offspring_1_parent-offspring1_g1-b1-i1.all.hs37d5.cons.90perc.trimBAM.bam" signifies a BAM file with 1x coverage, originating from "run1," containing individuals from the "run_*_parent-offspring_1.vcf" file (first set of parent-offspring pairs) where "_g1-b1-" designates the first parent in the first generation. The latter half of the name "hs37d5.cons.90perc.trimBAM.bam" is the same across all files.&nbsp;&nbsp;</p><p><strong>Note1:</strong> Segments such as <strong>parent-offspring*_g*-b*- </strong>can also be tracked in the naming of the genotype columns in the VCF.</p><p><strong>Note2: </strong>Sexual information within the VCF files is discernible from the genetic data present at X chromosome positions. Individuals carrying two genotypes on the X chromosome are female, while those with a single genotype are male.</p><p><strong>Note3: Some of the individuals from distinct pedigrees may</strong>,<strong> in fact</strong>,<strong> be related due to shared ancestry through common founders. To suit specific research objectives, researchers may need to identify and exclude such relatives if the full dataset is used for kinship estimation.</strong></p><p>For more details about the dataset's generation process, unique characteristics, or any specific inquiries, our team is available for further information. We welcome and encourage inquiries, aiming to provide comprehensive support and additional details that might aid researchers in utilizing this dataset effectively. Please don't hesitate to contact us for any specific information you may need.</p><p>This repository contains only VCFs and cov1x BAM and 200K_positions files. The rest of the files can be found at <strong>10.5281/zenodo.10079625 </strong>and<strong> 10.5281/zenodo.10079685.</strong></p><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study VCF files

<p>VCF files used and cited in the article: Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study</p>

opencc-by-4.0May 2025View details →
zenodo44/100

Harmonized Vegetation Continuous Fields (VCF)

<p><strong>Motivation</strong></p> <p><a href="https://www.nature.com/articles/s41586-018-0411-9">Song&rsquo;s Vegetation Continuous fields (VCF) product</a>, based on AVHRR satellite data, is the longest time-series of its type, but lacks updates past 2016 due to the extensive degradation of the sensor. We used machine learning to extend this time-series using data from the <a href="https://land.copernicus.eu/global/products/lc">Copernicus Land Cover dataset</a>, which provides per-pixel proportions of different land cover classes between 2015 and 2019. In addition, we included <a href="https://modis.gsfc.nasa.gov/data/dataprod/mod44.php">MODIS VCF data</a>.</p> <p><strong>Content</strong></p> <p>This repository contains the infrastructure used to model Song-like VCF data past 2016. This infrastructure contains a yaml file that configures the modelling framework (e.g. variables, directories, hyper-parameter tuning), and that interacts with a standardized folder structure.</p> <p><strong>Modelling approach</strong></p> <p>Song&#39;s VCF dataset includes data on generic categories, namely &ldquo;tree cover&rdquo;, &ldquo;non-tree vegetation&rdquo;, and &ldquo;non vegetated&rdquo;. Given the Copernicus dataset has a higher thematic detail, we first aggregated these data into comparable classes. We created a &ldquo;Non-tree vegetation&rdquo; layer (i.e. total per-pixel proportion of crops, grasses, shrubs, and mosses), and a &ldquo;Non Vegetated&rdquo; layer (i.e. total per-pixel proportion of bare land, permanent water, urban, and snow). Independent data on &ldquo;Tree cover&rdquo; was already present.</p> <p>We then constructed a Random Forest Regression (RFReg) model to predict Song-like VCF layers between 2016 and 2019. The predictions were informed by variables on topography, climate, and fires (which limit the density of vegetation), and by variables on differences between the Copernicus VCF and MODIS-based VCF data. Because MODIS data is available past 2016, its inclusion informs our models on how MODIS data, and their differences compared to Copernicus data, relate to the values reported in Song&#39;s data.</p> <p><strong>Sampling scheme</strong></p> <p>For each VCF category, we collected samples on a country-by-country basis. Within each country, we estimated the difference in percent cover between the Song&#39;s and Copernicus VCF data, and sampled across a gradient of differences, from -100% (no cover in AVHRR and full cover in Copernicus) to +100% (full cover in AVHRR and no cover in Copernicus). We iterated through this range in intervals of 10% and sampled across a gradient of &ldquo;tree cover&rdquo;, &ldquo;non-tree vegetation&rdquo;, and &ldquo;non vegetated&rdquo;, in intervals of 10% from 0% to 100%. We collected at least one sample per 50 km<sup>2</sup> in 2016, the last year where all VCF-related variables (Song&#39;s, Copernicus, MODIS) are available simultaneously. The amount of samples attributed to each range of differences is proportional to the area covered by this range within the country of reference. The sampling approach was repeated for each VCF class, and the outputs were later combined into a single set of samples that exclude duplicates, resulting in 238,052 samples.</p> <p><strong>Validation</strong></p> <p>The model outputs were validated using leave-one-out cross-validation. For each VCF class, the validation framework iterates through each country where samples were collected, excluding it for validation and using the remaining samples to train a RFReg models.This resulted in R<sup>2</sup> values of 0.91, 0.87 and 0.91 for &ldquo;tree cover&rdquo;, &ldquo;non-tree vegetation&rdquo;, and &ldquo;non vegetated&rdquo;. respectively. The RMSE values were of 2.31%, 3.05%, and 2.25%.</p> <p>The model was applied to data from 2015, which was not used to neither predict nor validate our models. A comparison between the 2015 Song data against our predictions, which consist of 8,764,232 pixels, yielded R<sup>2</sup> values of 0.94, 0.91, and 0.97. The RMSE were 6.65%, 8.92%, and 5.96%. Additionally, we compared changes between 2015 and 2016, resulting in RMSE values of 2.83%, 3.69%, and 2.57%.</p> <p><strong>Post-processing</strong></p> <p>When observing annual VCF time-series based on Song&#39;s data, we noted that our predictions were the most plausible for &ldquo;tree cover&rdquo; and &ldquo;non-tree vegetation&rdquo;. In turn, our &ldquo;non vegetated&rdquo; are seemingly underestimated (see &quot;temporal_trend_check.png&quot;), reporting large year-to-year decreases om cover (-3.05% between 2016 and 2017, compared to -0.14% for &quot;tree cover&quot; and -0.26% for &ldquo;non-tree vegetation&rdquo;). To address this issue, we recommend deriving data on &ldquo;non-vegetated&rdquo; cover by computing the difference between 100% and the sum of &quot;tree cover&rdquo; and &ldquo;non-tree vegetation&rdquo;.</p> <div class="notranslate">&nbsp;</div>

opencc-by-4.0Aug 2023View details →
dryad40/100

Dsuite - fast D-statistics and related admixture evidence from VCF files

<p>Patterson's D, also known as the ABBA-BABA statistic, and related statistics such as the f4-ratio, are commonly used to assess evidence of gene flow between populations or closely related species. Currently available implementations often require custom file formats, implement only small subsets of the available statistics, and are impractical to evaluate all gene flow hypotheses across datasets with many populations or species due to computational inefficiencies. Here we present a new software package Dsuite, an efficient implementation allowing genome scale calculations of the D and f4-ratio statistics across all combinations of tens or hundreds of populations or species directly from a variant call format (VCF) file. Our program also implements statistics suited for application to genomic windows, providing evidence of whether introgression is confined to specific loci and it can also aid in interpretation of a system of f4-ratio results with the use of the 'f-branch' method. Dsuite is available at https://github.com/millanek/Dsuite, is straightforward to use, substantially more computationally efficient than comparable programs, and provides a convenient suite of tools and statistics, including some not previously available in any software package. Thus, Dsuite facilitates the assessment of evidence for gene flow, especially across larger genomic datasets.</p>

opencc-zeroDec 2019View details →
zenodo40/100

VCF file of Variants in a wide onion cross segregating for bolting

<p>VCF (variant call format) file of&nbsp;variants from bulked segregant RNA PoolSeq (BSR-seq) of F2 progeny pools of a wide onion cross segregating for bolting (precocious flowering). The reference assembly is&nbsp;GBGJ00000000.1&nbsp;http://www.ncbi.nlm.nih.gov/nuccore/656904698</p> <p>The pools were taken from materials sampled for validation of the <em>AcBlt1</em> locus as described in&nbsp;http://www.ncbi.nlm.nih.gov/pubmed/24247236</p> <p>Pools were from bolting or non-bolting plants homozygous at the most closely linked marker to&nbsp;<em>AcBlt1 &nbsp;</em>either for the bolt-associated genotype (AA) or the non-bolt genotype (BB).&nbsp;</p>

opencc-zeroAug 2014View details →
zenodo40/100

VCF, pileup, and other files for Betula analyses using ebg & GATK

<p>These are the VCF, pileup, and input files files produced by the Genome Analysis ToolKit, SAMtools, and our own scripts, respectively, used for comparing the genotypes estimated by the GATK UnifiedGenotyper and the new model for allopolyploids that we introduced in our paper. Additional processing of the files was conducting using Python and R for filtering and extracting error values and read counts (scripts available on GitHub: https://github.com/pblischak/polyploid-genotyping).</p> <p><strong>Main files:</strong></p> <ul> <li>pendula-ug-filtered30.vcf: VCF file from GATK with all called variants for <em>Betula</em> <em>pendula</em> (diploid).</li> <li>filtered30-pendula.pileup: SAMtools pileup file for variant sites identified by GATK in <em>B</em>. <em>pendula</em>.</li> <li>pubescens-ug-filtered30.vcf: VCF file from GATK with all called variants for <em>Betula</em> <em>pubescens</em> (allotetraploid).</li> <li>filtered30-pubescens.pileup: SAMtools pileup file for variant sites identified by GATK in <em>B</em>. <em>pubescens</em>.</li> </ul> <p><strong>Processed files:</strong></p> <ul> <li>filtered30-variants.txt: tab delimited file of shared variants between B. pendula and B. pubescens.</li> <li>filtered30-vcf1.vcf: VCF file for B. pendula with variants extracted from filtered30-variants.txt.</li> <li>filtered30-vcf2.vcf: VCF file for B. pubescens with variants extracted from filtered30-variants.txt.</li> </ul> <p><strong>Input files</strong>:</p> <ul> <li>filtered30-pubescens-tot.txt, filtered30-pubescents-alt.txt, filtered30-pubescens-err.txt: input files for running the alloSNP model in ebg.</li> <li>filtered30-pubescens-alloSNP-freqs2.txt, filtered30-pubescens-alloSNP-g1.txt, filtered30-pubescens-alloSNP-g2.txt: output files from the alloSNP model.</li> <li>filtered30-pendula-tot.txt, filtered30-pendula-alt.txt, filtered30-pendula-err.txt: input files for running the hwe model in ebg.</li> <li>filtered30-pendula-hwe-freqs.txt, filtered30-pendula-hwe-genos.txt: output files from the hwe model. The allele frequency estimates here were also used as the reference panel for the alloSNP model.</li> </ul>

opencc-by-4.0Jul 2017View details →
dryad40/100

VCF files of common grassland plants from wild collected seeds of 19 common European grassland species with up to 4 consecutive generations grown in monoculture for seed production for restoration

<p>A growing number of restoration projects require large amounts of seeds. As harvesting natural populations cannot cover the demand, wild plants are often propagated in large-scale monocultures. There are concerns that this cultivation process may cause genetic drift and unintended selection, which would alter the genetic properties of the cultivated populations and reduce their genetic diversity. Such changes could reduce the pre-existing adaptation of restored populations, and limit their adaptability to environmental change.</p> <p>We used single nucleotide polymorphism (SNP) markers and a pool-sequencing approach to test for genetic differentiation and changes in gene diversity during cultivation in 19 wild grassland species, comparing the source populations and up to four consecutive cultivation generations. We then linked the magnitudes of genetic changes to the species' breeding systems and seed dormancy, to understand the roles of these traits in genetic change.</p> <p>The propagation changed the genetic composition of the cultivated generations only moderately. The genetic differentiation we observed as a consequence of cultivation was much lower than the natural genetic differentiation between different source regions. The propagated generations harbored even higher gene diversity than wild-collected seeds. Genetic change was stronger in self-compatible than in self-incompatible species, probably as a result of increased outcrossing in the monocultures.</p> <p><em>Synthesis and applications</em>: Our study indicates that large-scale seed production maintains the genetic integrity of natural populations. Increased genetic diversity may be indicative of increased adaptive potential of propagated seeds, which would make them especially suitable for ecological restoration. Yet, it remains to be tested whether these patterns observed on the level of molecular markers will be mirrored also in plant phenotypes. Further, we used seeds produced in Germany and Austria, where the seed production is regulated and certified. Whether other seed production systems perform equally well remains to be tested.</p>

opencc-zeroMar 2022View details →
zenodo40/100

VCF data for genetic polyploid phasing

<p>Used data for evaluation in pending Recomb 2022 submission &quot;Genetic Polyploid Phasing using Marker Signals from Low-Depth Progeny Samples&quot;. Data was used in and partly created by the scripts from here:&nbsp;https://github.com/AlBi-HHU/genetic-phasing-scripts</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

WES cropbioBonn WGGC CN 200M vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 200M reads</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Development of a high-density 665 K SNP array for rainbow trout genome-wide genotyping. Supplemental VCF file

<p>Single nucleotide polymorphism (SNP) arrays, also named &laquo; SNP chips &raquo;, enable very large numbers of individuals to be genotyped at a targeted set of thousands of genome-wide identified markers. We used preexisting variant datasets from USDA, a French commercial line and 30X-coverage whole genome sequencing of INRAE isogenic lines to develop an Affymetrix 665 K SNP array (HD chip) for rainbow trout. In total, we identified 32,372,492 SNPs that were polymorphic in the USDA or INRAE databases. A subset of identified SNPs were selected for inclusion on the chip, prioritizing SNPs whose flanking sequence uniquely aligned to the Swanson reference genome, with homogenous repartition over the genome and the highest Minimum Allele Frequency in both USDA and French databases. Of the 664,531 SNPs which passed the Affymetrix quality filters and were manufactured on the HD chip, 65.3% and 60.9% passed filtering metrics and were polymorphic in two other distinct French commercial populations in which, respectively, 288 and 175 sampled fish were genotyped. Only 576,118 SNPs mapped uniquely on both Swanson and Arlee reference genomes, and 12,071 SNPs did not map at all on the Arlee reference genome. Among those 576,118 SNPs, 38,948 SNPs were kept from the&nbsp; commercially available medium-density 57K SNP chip. We demonstrate the utility of the HD chip by describing the high rates of&nbsp; linkage disequilibrium at 2 kb to 10 kb in the rainbow trout genome in comparison to the linkage disequilibrium observed at 50 kb to&nbsp; 100 kb which are usual distances between markers of the medium-density chip.</p> <p>&nbsp;</p> <p>File submitted correspond to the supplementary data 1 of the publication (under submission) : INRAE_USDA_MAF1.vcf.gz</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

VCF File containing genotype calls for 136 Populus alba x Populus tremula hybrids obtained through both RAD-seq and GBS

<p>VCF file used to compare genotype calls obtained through RAD-seq and GBS for 126 common garden seedlings of Populus tremula and Populus alba hybrids. See Bresadola et al. (2019) for more details.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

somalier files for thousand genomes high-coverage VCF for 2504 samples

<p>somalier files extracted from the thousand genomes VCF to be used for ancestry prediction with somalier.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Baboon and Gelada SNP Calls VCF

<p>Bgzipped vcf and tabix index files of baboon and gelada SNP calls of 16&nbsp;individuals on the papAnu2&nbsp;assembly as published in Rogers et al. (2019). The comparative genomics and complex population history of Papio baboons. Science Advances. &nbsp;30 Jan 2019: Vol. 5, no. 1, eaau6947 DOI: 10.1126/sciadv.aau6947.</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

VCF file from: Signatures of local adaptation to current and future climate in phenology-related genes in natural populations of Quercus robur

<p>Qrob_seqcap_87_18799.vcf - containing 18,799 single nucleotide variants (SNPs) in phenology-related genes in 87 individuals of Quercus robur from 6 populations,&nbsp;described in Meger&nbsp;et al.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

WES cropbioBonn WGGC CN Twist Exome vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Twist Exome, NovaSeq 6000</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

WES cropbioBonn WGGC CN 75M vcf results

<p>Whole Exome Sequencing Data Analysis with <a href="https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof">https://github.com/tgstoecker/WES_WGGC_2021_AG_Schoof</a></p> <p>HG001/NA12878 Agilent v6 2x101bp PE, NovaSeq 6000, 75M reads</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

VCF files of ddRAD seq data of Bermuda petrel population

<p>Two VCF files resulted from two different filtering and analyses of ddRAD sequencing data of the endangered Bermuda petrel. The one_snp VCF contains only one snp per RAD locus, while the all_snps contained all SNPs in a RAD locus.</p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Genomic datasets of Laminaria digitata: Paired-end reads from dd-RADseq, reference genome assembly and filtered VCF

<p>The long-term persistence of species in the face of climate change can be evaluated by examining the interplay between selection and genetic drift in the contemporary evolution of populations. In this study, we focused on spatial and temporal genetic variation in four populations of the cold-water kelp Laminaria digitata using thousands of SNPs (ddRAD-seq). These populations were sampled from the center to the south margin in the North Atlantic at two different time points, spanning at least two generations. By conducting genome scans for local adaptation from a single time point, we successfully identified candidate loci that exhibited clinal variation, closely aligned with the latitudinal changes in temperature. This finding suggests that temperature may drive the adaptive response of kelp populations, although other factors, such as the species' demographic history should be considered. Furthermore, we provided compelling evidence of selection through the examination of allele frequency changes over time, by taking into the impact of genetic drift. Specifically, we detected candidate loci exhibiting temporal differentiation that surpassed the levels typically attributed to genetic drift at the south margin, confirmed through simulations. This finding was in sharp contrast with the lack of detection of outlier loci based on temporal differentiation in a population from the North Sea, exhibiting low and decreasing levels of genetic diversity. These contrasting evolutionary scenarios among populations can be primarily attributed to the differential prevalence of selection relative to genetic drift. In conclusion, our study highlights the potential of temporal genomics to gain deeper insights into the contemporary evolution of marine foundation species in response to rapid environmental changes.</p>

opencc-zeroJun 2023View details →
dryad40/100

VCF files of common grassland plants from wild collected seeds of 19 common European grassland species with up to 4 consecutive generations grown in monoculture for seed production for restoration

Open the record for dataset details and reuse information.

publicMar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record