Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Tuatara (Sphenodon punctatus) ab initio interspersed repeat consensus sequences from the Tuatara genome assembly.
<p>These repeat consensus sequences are part of the genome analysis of the Tuatara genome. </p>
An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.
<p>This archive is associated with the article “An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.”. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin & Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>
Sanger sequencing of target and off-target genomic regions for gene-edited iPSC clones with SETBP1 genetic variants
<p>This data set includes chromatograms generated using sanger sequencing of targeted regions of genomic DNA from clonal iPSC lines. The iPSC lines include clones generated using CRISPR/Cas9 homology directed repair to introduce genetic variants into <em>SETBP1,</em> and their wild-type controls. Additional files have been included in the data set to link chromatogram (ab1) files to specific iPSC clones for genomic regions across the variant in <em>SETBP1 (</em>SETBP1 clones genetic variant sanger sequencing.xslx)<em> </em>and top<em> </em>off-target sites (SETBP1 clones off-target sanger sequencing.xlsx). </p>
AntiSMASH results of the genome sequences of four novel Endozoicomonas strains associated with the octocoral Litophyton in a long-term aquarium facility
<p>Secondary Metabolite-Encoding Biosynthetic Gene Cluster (SM-BGC) annotation files from antiSMASH bacterial version 7.1.0 for four <em>Endozoicomonas</em> strains associated with the tropical octocoral Litophyton in a long-term aquarium facility. Data corresponds to the assemblies of NE35, NE40, NE41 and NE43, available under the BioProject accession numbers <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075803">PRJNA1075803</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075804">PRJNA1075804</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075805">PRJNA1075805</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075806">PRJNA1075806</a>, respectively.<u> </u></p> <p>To easily and interactively review the results, please download the genome you wish to examine, extract the entire contents of the folder, and open the HTML file named "Results".</p> <p> </p> <p>This dataset is part of the following study:</p> <p>Marques M, da Silva DMG, Santos E, Baylina N, Peixoto R, Kyrpides NC, Woyke T, Whitman WB, Keller-Costa T, Costa R. 2024. Genome sequences of four novel <em>Endozoicomonas </em>strains associated with a tropical octocoral in a long-term aquarium facility. Microbiology Resource Announcements</p>
Whole genome RNA-sequencing reveals modulation of genes related to brain disorders by Withania somnifera in human neuroblastoma SK-N-SH cells
<p>Table S1: Human reference genome based differential gene expression; Figure S1: Reactome Pathway (50 μg/mL_3h vs C_3h); Figure S2: Reactome Pathway (50 μg/mL_9h vs C_9h); Figure S3: Reactome pathway (100 μg/mL_3h vs C_3h); Figure S4: GO Dose comparison; Reactome pathway (100 μg/mL_3h vs 50 μg/mL_3h); Figure S5: GO Dose comparison; Reactome pathway (100 μg/mL_9h vs 50 μg/mL_9h); Figure S6: GO Time comparison; Reactome pathway (50 μg/mL_9h vs 50 μg/mL_3h); Figure S7: GO Time comparison; Reactome pathway (100 μg/mL_9h vs 100 μg/mL_3h); Table S2: Disease ontology analysis of 100 μg/mL_3h vs 50 μg/mL_3h WS-treated SK-N-SH cells; Table S3: Disease ontology analysis of 100 μg/mL_9h vs 50 μg/mL_9h WS-treated SK-N-SH cells; Table S4: Disease ontology analysis of 100 μg/mL_9h vs 100 μg/mL_3h WS-treated SK-N-SH cells.</p>
COG and Pfam annotation results of the genome sequences of four novel Endozoicomonas strains associated with the octocoral Litophyton in a long-term aquarium facility
<p>COG and Pfam annotation files from DOE-JGI Microbial Genome Annotation Pipeline (MGAP) version 4(1), for four <em>Endozoicomonas</em> strains associated with the tropical octocoral Litophyton in a long-term aquarium facility. Data correspond to the assemblies of NE35, NE40, NE41, and NE43, available under the BioProject accession numbers <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075803">PRJNA1075803</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075804">PRJNA1075804</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075805">PRJNA1075805</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075806">PRJNA1075806</a>, respectively. Results were submitted to the Integrated Microbial Genomes and Microbiomes system v7 (IMG/M) (2) for comparative analysis. The genome annotations can be interactively accessed on IMG/M (https://img.jgi.doe.gov/cgi-bin/m/main.cgi) using the following identifiers: 8036267134 (strain NE35), 8036277142 (strain NE40), 8045494135 (strain NE41); 8036272146 (strain NE43). </p> <p>This dataset is part of the following study:</p> <p>Marques M, da Silva DMG, Santos E, Baylina N, Peixoto R, Kyrpides NC, Woyke T, Whitman WB, Keller-Costa T, Costa R. 2024. Genome sequences of four novel <em>Endozoicomonas </em>strains associated with a tropical octocoral in a long-term aquarium facility. Microbiology Resource Announcements</p> <p> </p> <p>Other reference sources:</p> <p>(1) Huntemann M, Ivanova NN, Mavromatis K, James Tripp H, Paez-Espino D, Palaniappan K, Szeto E, Pillay M, Chen IMA, Pati A, Nielsen T, Markowitz VM, Kyrpides NC. 2015. The standard operating procedure of the DOE-JGI Microbial Genome Annotation Pipeline (MGAP v.4). Stand Genomic Sci 10:1–6.</p> <p>(2) Chen IMA, Chu K, Palaniappan K, Ratner A, Huang J, Huntemann M, Hajek P, Ritter SJ, Webb C, Wu D, Varghese NJ, Reddy TBK, Mukherjee S, Ovchinnikova G, Nolan M, Seshadri R, Roux S, Visel A, Woyke T, Eloe-Fadrosh EA, Kyrpides NC, Ivanova NN. 2023. The IMG/M data management and analysis system v.7: content updates and new features. Nucleic Acids Res 51:D723–D732.</p>
genomic sequencies of phages (prokaryotic vriuses) and their hosts
<p>This repository contains data relevant to the reserach <a href="https://github.com/yoheyokubo/MCL4PHI">Multi-Instance Contrastive Learning with Binomial K-Mers for Phage Host Interaction Prediction</a>. It is comprised of the preprocessed data (<em>data_preprocessed.zip</em>) to quickly reproduce our results and the raw genomic sequencies (<em>data_raw.zip</em>) of phages and hosts, originally from <a href="https://github.com/KennthShang/CHERRY">CHERRY's dataset</a>.</p>
Data from: "Identification of SNP markers for the endangered Ugandan red colobus (Procolobus rufomitratus tephrosceles) using RAD sequencing" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015
Despite dramatic growth in the field of primate genomics over the past decade, studies of primate population and conservation genomics in the wild have been hampered due to the difficulties inherent in studying non-model organisms and endangered species, such as lack of a reference genome and current challenges in de novo primate genome assembly. Here, we used Restriction-site Associated DNA (RAD) sequencing to develop a population-based SNP panel for the Ugandan red colobus (P. rufomitratus tephrosceles), which is a highly threatened monkey due to habitat loss. We analyzed blood samples from 24 individuals from Kibale National Park (Uganda) using single-end RAD sequencing. We obtained 70,773,857 reads, of which 58,814,906 passed the filtering steps. Using the program STACKS v. 1.11 we identified 113,376 loci, of which 50,558 were polymorphic and had a mean observed heterozygosity of 0.25. These data will be used to study the effects of habitat fragmentation on genomic diversity, dispersal, and disease transmission in this species. Our approach provides a good example of the potential of RAD sequencing in studies of wild primate populations.
Data from: Genome-wide RAD sequence data provide unprecedented resolution of species boundaries and relationships in the Lake Victoria cichlid adaptive radiation
Although population genomic studies using next generation sequencing (NGS) data are becoming increasingly common, studies focusing on phylogenetic inference using these data are in their infancy. Here, we use NGS data generated from reduced representation genomic libraries of restriction-site-associated DNA (RAD) markers to infer phylogenetic relationships among 16 species of cichlid fishes from a single rocky island community within Lake Victoria's cichlid adaptive radiation. Previous attempts at sequence-based phylogenetic analyses in Victoria cichlids have shown extensive sharing of genetic variation among species and no resolution of species or higher-level relationships. These patterns have generally been attributed to the very recent origin (<15 000 years) of the radiation, and ongoing hybridization between species. We show that as we increase the amount of sequence data used in phylogenetic analyses, we produce phylogenetic trees with unprecedented resolution for this group. In trees derived from our largest data supermatrices (3 to >5.8 million base pairs in width), species are reciprocally monophyletic with high bootstrap support, and the majority of internal branches on the tree have high support. Given the difficulty of the phylogenetic problem that the Lake Victoria cichlid adaptive radiation represents, these results are striking. The strict interpretation of the topologies we present here warrants caution because many questions remain about phylogenetic inference with very large genomic data set and because we can with the current analysis not distinguish between effects of shared ancestry and post-speciation gene flow. However, these results provide the first conclusive evidence for the monophyly of species in the Lake Victoria cichlid radiation and demonstrate the power that NGS data sets hold to resolve even the most difficult of phylogenetic challenges.
Data from: Chironomus riparius (Diptera) genome sequencing reveals the impact of minisatellite transposable elements on population divergence
Active transposable elements (TEs) may result in divergent genomic insertion and abundance patterns among conspecific populations. Upon secondary contact, such divergent genetic backgrounds can theoretically give rise to classical Dobzhansky-Muller incompatibilities (DMI), thus contributing to the evolution of endogenous genetic barriers and eventually cause population divergence. We investigated differential TE abundance among conspecific populations of the non-biting midge Chironomus riparius and evaluated their potential role in causing endogenous genetic incompatibilities between these populations. We focussed on a Chironomus-specific TE, the minisatellite-like Cla-element, whose activity is associated with speciation in the genus. Using a newly generated and annotated draft genome for a genomic study with five natural C. riparius populations, we found highly population-specific TE insertion patterns with many private insertions. A significant correlation of the pairwise FST estimated from genome-wide single nucleotide polymorphisms (SNPs) and the FST estimated from TEs, is consistent with drift as the major force driving TE population differentiation. However, the significantly higher Cla-element FST level due to a high proportion of differentially fixed Cla-element insertions also indicates selection against segregating (i.e. heterozygous) insertions. With reciprocal crossing experiments and fluorescent in-situ hybridisation of Cla-elements to polytene chromosomes, we documented phenotypic effects on female fertility and chromosomal mispairings. We propose that the inferred negative selection on heterozygous Cla-element insertions may cause endogenous genetic barriers and therefore acts as DMI among C. riparius populations. The intrinsic genomic turnover exerted by TEs may thus have a direct impact on population divergence that is operationally different from drift and local adaptation.
Data from: HyRAD-X, a versatile method combining exome capture and RAD sequencing to extract genomic information from ancient DNA
Over the last decade, protocols aimed at reproducibly sequencing reduced-genome subsets in non-model organisms have been widely developed. Their use is however limited to DNA of relatively high molecular weight. During the last year, several methods exploiting hybridization capture using probes based on RAD-sequencing loci have circumvented this limitation and opened avenues to the study of samples characterized by degraded DNA, such as historical specimens. Here, we present a major update to those methods, namely Hybridization capture from RAD-derived probes obtained from a reduced eXome template (hyRAD-X), a technique applying RAD-sequencing to messenger RNA from one or few fresh specimens to elaborate bench-top produced probes, i.e., a reduced representation of the exome, further used to capture homologous DNA from a samples set. In contrast to previous hybridization-capture methods, the reference catalog on which reads are aligned does not rely on de novo assembly of anonymous RAD-sequencing loci, but on an assembled transcriptome obtained from RNAseq data, thus increasing the accuracy of loci definition and Single-Nucleotide-Polmorphisms (SNP) call, and targeting, specifically, expressed genes. Finally, the capture step of hyRAD-X relies on RNA probes, increasing stringency of hybridization, making it well suited for low-content DNA samples. As a proof of concept, we applied hyRAD-X to subfossil needles from the coniferous tree Abies alba, collected in lake sediments (Origlio, Switzerland) and dating back from 7200-5800 years before present (BP). More specifically we investigated genetic variation before, during, and after an anthropogenic perturbation that caused an abrupt decrease in Abies alba population size, 6500-6200 years BP. HyRAD-X produced a matrix encompassing 524 exome-derived SNPs. Despite a lower observed heterozygosity was observed during the 6.500-6.200 years BP time slice, genetic composition was nearly identical before and after the perturbation, indicating that re-expansion of the population after the decline was driven by autochthonous specimens. To the best of our knowledge, this is the first time a population genomic study incorporating ancient DNA samples of tree subfossils is conducted at a moderate cost using reproducible exome-reduced complexity.
Data from: Whole genome sequencing shows sleeping sickness relapse is due to parasite regrowth and not reinfection
The trypanosome Trypanosoma brucei gambiense (Tbg) is a cause of human African trypanosomiasis (HAT) endemic to many parts of sub-Saharan Africa. The disease is almost invariably fatal if untreated and there is no vaccine, which makes monitoring and managing drug resistance highly relevant. A recent study of HAT cases from the Democratic Republic of the Congo reported a high incidence of relapses in patients treated with melarsoprol. Of the 19 Tbg strains isolated from patients enrolled in this study, four pairs were obtained from the same patient before treatment and after relapse. We used whole genome sequencing to investigate whether these patients were infected with a new strain, or if the original strain had regrown to pathogenic levels. Clustering analysis of 5938 single nucleotide polymorphisms supports the hypothesis of regrowth of the original strain, as we found that strains isolated before and after treatment from the same patient were more similar to each other than to other isolates. We also identified 23 novel genes that could affect melarsoprol sensitivity, representing a promising new set of targets for future functional studies. This work exemplifies the utility of using evolutionary approaches to provide novel insights and tools for disease control.
Data from: Genotyping-by-sequencing reveals genomic homogeneity among overwintering Pacific Dunlin (Calidris alpina pacifica) aggregations along the Pacific coast of North America
Information on how migratory populations are genetically structured during the overwintering season of the annual cycle can improve our understanding of the strength of migratory connectivity and help identify populations as units for management. Here, we use a genotype-by-sequencing approach to investigate whether population genetic structure exists among overwintering aggregations of the Pacific Dunlin subspecies (Calidris alpina pacifica) sampled at two spatial scales (i.e. within and among overwintering sites) in the eastern Pacific Flyway. Genome-wide analyses of 874 single nucleotide polymorphisms across 80 sampled individuals revealed no evidence for genetic differentiation among aggregations overwintering at three locations within the Fraser River Estuary (FRE) of British Columbia. Similarly, comparisons of aggregations in the FRE and those overwintering in southern sites in California and Mexico indicated no genetic segregation between northern and southern overwintering areas. These results suggest that Pacific Dunlin residing within the FRE, Sacramento Valley (California) and Guerrero Negro (Mexico) are genetically homogeneous, with no evident genetic structure between sampled sites or regions across the overwintering range. Despite no evidence for differentiation among aggregations, we identified a significant effect of geographical distance between sites on the distribution of individual genotypes in a redundancy analysis; however, a small proportion of the total genotypic variance (R2 = 0.036, P = 0.011) was explained by the combined effect of latitude and longitude, suggesting weak genomic patterns of isolation-by-distance that are consistent with chain-like migratory connectivity between breeding and overwintering areas. Our study represents the first genome-scale investigation of population structure for a Dunlin subspecies and provides essential baseline estimates of genomic diversity and differentiation within the Pacific Dunlin.
Data from: Revisiting comparisons of genetic diversity in stable and declining species: assessing genome-wide polymorphism in North American bumble bees using RAD sequencing
Genetic variation is of key importance for a species' evolutionary potential, and its estimation is a major component of conservation studies. New DNA sequencing technologies have enabled the analysis of large portions of the genome in nonmodel species, promising highly accurate estimates of such population genetic parameters. Restriction site-associated DNA sequencing (RADseq) is used to analyse thousands of variants in the bumble bee species Bombus impatiens, which is common, and Bombus pensylvanicus, which is in decline. Previous microsatellite-based analyses have shown that gene diversity is lower in the declining B. pensylvanicus than in B. impatiens. RADseq nucleotide diversities appear much more similar in the two species. Both species exhibit allele frequencies consistent with historical population expansions. Differences in diversity observed at microsatellites thus do not appear to have arisen from long-term differences in population size and are either recent in origin or may result from mutational processes. Additional research is needed to explain these discrepancies and to investigate the best ways to integrate next-generation sequencing data and more traditional molecular markers in studies of genetic diversity.
Data from: Whole genome sequencing of two North American Drosophila melanogaster populations reveals genetic differentiation and positive selection
The prevailing demographic model for Drosophila melanogaster suggests that the colonization of North America occurred very recently from a subset of European flies that rapidly expanded across the continent. This model implies a sudden population growth and range expansion consistent with very low or no population subdivision. As flies adapt to new environments, local adaptation events may be expected. To describe demographic and selective events during North American colonization, we have generated a data set of 35 individual whole-genome sequences from inbred lines of D. melanogaster from a west coast US population (Winters, California, USA) and compared them with a public genome data set from Raleigh (Raleigh, North Carolina, USA). We analysed nuclear and mitochondrial genomes and described levels of variation and divergence within and between these two North American D. melanogaster populations. Both populations exhibit negative values of Tajima's D across the genome, a common signature of demographic expansion. We also detected a low but significant level of genome-wide differentiation between the two populations, as well as multiple allele surfing events, which can be the result of gene drift in local subpopulations on the edge of an expansion wave. In contrast to this genome-wide pattern, we uncovered a 50-kilobase segment in chromosome arm 3L that showed all the hallmarks of a soft selective sweep in both populations. A comparison of allele frequencies within this divergent region among six populations from three continents allowed us to cluster these populations in two differentiated groups, providing evidence for the action of natural selection on a global scale.
Data from: Whole genome sequencing and rare variant analysis in essential tremor families
Essential tremor (ET) is one of the most common movement disorders. The etiology of ET remains largely unexplained. Whole genome sequencing (WGS) is likely to be of value in understanding a large proportion of ET with Mendelian and complex disease inheritance patterns. In ET families with Mendelian inheritance patterns, WGS may lead to gene identification where WES analysis failed to identify the causative single nucleotide variant (SNV) or indel due to incomplete coverage of the entire coding region of the genome, in addition to accurate detection of larger structural variants (SVs) and copy number variants (CNVs). Alternatively, in ET families with complex disease inheritance patterns with gene x gene and gene x environment interactions enrichment of functional rare coding and non-coding variants may explain the heritability of ET. We performed WGS in eight ET families (n=40 individuals) enrolled in the Family Study of Essential Tremor. The analysis included filtering WGS data based on allele frequency in population databases, rare SNV and indel classification and association testing using the Mixed-Model Kernel Based Adaptive Cluster (MM-KBAC) test. A separate analysis of rare SV and CNVs segregating within ET families was also performed. Prioritization of candidate genes identified within families was performed using phenolyzer. WGS analysis identified candidate genes for ET in 5/8 (62.5%) of the families analyzed. WES analysis in a subset of these families in our previously published study failed to identify candidate genes. In one family, we identified a deleterious and damaging variant (c.1367G>A, p.(Arg456Gln)) in the candidate gene, CACNA1G, which encodes the pore forming subunit of T-type Ca(2+) channels, CaV3.1, and is expressed in various motor pathways and has been previously implicated in neuronal autorhythmicity and ET. Other candidate genes identified include SLIT3 which encodes an axon guidance molecule and in three families, phenolyzer prioritized genes that are associated with hereditary neuropathies (family A, KARS, family B, KIF5A and family F, NTRK1). Functional studies of CACNA1G and SLIT3 suggest a role for these genes in ET disease pathogenesis.
Phylogeography of the Rough Greensnake, Opheodrys aestivus (Squamata: Colubridae), using multilocus Sanger sequence and genomic ddRADseq data
<p>The Rough Greensnake, <i>Opheodrys aestivus,</i> is a moderately-sized, semi-arboreal snake broadly distributed throughout eastern North America. While numerous taxa with similar distributions have been shown to be comprised of multiple species, <i>O. aestivus</i> has yet to be examined in a detailed phylogeographic context. Here, we use Sanger-sequence data of one mitochondrial and three nuclear loci for samples from throughout the distribution of <i>O. aestivus</i> to elucidate phylogeographic patterns in this species. We combine this with ddRADseq data for a subset of samples to test patterns on a more genomically comprehensive scale. In both datasets, we find strong support for three deeply divergent clades within <i>O. aestivus</i>: peninsular Florida, central Texas, and a main clade comprising the rest of the distribution, with the Florida clade the earliest diverging lineage of the three. Estimates of divergence time suggest that the central Texas and main clades diverged approximately 1.34 million years ago (Mya), while the peninsular Florida clade diverged from other lineages approximately 2.94 Mya, and these lineages diverged from the sister taxon, <i>O. vernalis</i>, approximately 6.43 Mya.<i> </i>These results also suggest that the historically recognized Florida subspecies, <i>O. a. carinatus</i>, could be elevated to species status. While the divergence of peninsular Florida or central Texas populations is not unique among squamates, nor is low levels of divergence from the Atlantic coast to eastern Texas, this combination of patterns is unusual, and yields important insight into the biogeography of North American biota. Further, our approach helps illustrate how dense geographic sampling with limited genomic sequencing can be used as a guide for the selection of samples to test phylogeographic patterns comprehensively.</p>
Comparative Analysis of Complete Chloroplast Genomes and Multiple DNA Sequences Reveals Interspecific Relationships of C. bretschneideri and Related Species in China
<p><strong> ITS, and <em>LEAFY</em> intron 1 sequencing of 36 Crataegus accessions.</strong></p>
Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing
<p>Studies of structural variation (SV) have been challenging due to technological contraints. With the advent of third generation (long-read) sequencing technology, exploration of longer stretches of DNA not easily examined previously has been made possible. In the present study, we utilized third generation (long-read) sequencing techniques to examime SV in the <em>EGFR </em>landscape of four haplotypes derived from two human samples. We analyzed the <em>EGFR</em> gene and its landscape (+/- 500,000 base pairs) using this sequencing approach and were able to identify regions of non-coding DNA which had relatively high similarity to the most common activating <em>EGFR</em> mutation in non-small cell lung cancer. We discovered that reverse complements to the exon 19 deletion mutation which had at least 60% homology to the <em>EGFR</em> exon 19 canonical deletion and were within ± 421,000 bp of the deletion varied across the five haploid genomes examined (4 patient landscapes and hg38). Although the sample size is limited in this study, the estimated variation observed in genomic stability between the five <em>EGFR</em> haplotypes examined is novel and encourages further work to examine structural variation in larger cohorts.</p>
Figure 1 in Molecular phylogeny of major lineages of the avian family Phasianidae inferred from complete mitochondrial genome sequences
Figure 1. Molecular phylogenetic tree derived from complete DNA sequences of the 12 mitochondrial protein-coding genes using Bayesian inference, maximum parsimony and maximum likelihood analysis. The numbers beside the nodes are Bayesian posterior probabilities (≥ 0.95 retained) and bootstrap proportions (≥ 50% retained). Anas platyrhynchos was set as outgroup. ∗demonstrates that MP analysis does not support this branch.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.