Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
zenodo36/100

Genome and annotation files (v4_23) for Blumeria graminis f. sp. tritici isolate CHE_96224 (genome assembly: Bgt_genome_v3_16)

<p>Genome and annotation (v4_23) files for Blumeria graminis f. sp. tritici isolate CHE_96224 (genome assembly: Bgt_genome_v3_16)</p>

opencc-by-4.0Sep 2022View details →
dryad36/100

A chromosome-level genome assembly of the highly heterozygous sea urchin Echinometra sp. EZ reveals adaptation in the regulatory regions of stress response genes

<p><em>Echinometra</em> is the most widespread genus of sea urchin and has been the focus of a wide range of studies in ecology, speciation, and reproduction. However, available genetic data for this genus are generally limited to a few select loci. Here, we present a chromosome-level genome assembly based on 10x Genomics, PacBio, and Hi-C sequencing for <em>Echinometra</em> sp. EZ from the Persian/Arabian Gulf. The genome is assembled into 210 scaffolds totaling 817.8 Mb with an N50 of 39.5 Mb. From this assembly we determined that the <em>E</em>. sp. EZ genome consists of 2n = 42 chromosomes. BUSCO analysis showed that 95.3% of BUSCO genes were complete. ab initio and transcript-informed gene modeling and annotation identified 29,<span>405</span> genes, including a conserved Hox cluster. <em>E.</em> sp. EZ can be found in high-temperature and high-salinity environments, and we therefore compared gene families and transcription factors associated with environmental stress response ("defensome") with other echinoid species with similar high-quality genomic resources. While the number of defensome genes was broadly similar for all species, we identified strong signatures of positive selection in non-coding elements near genes involved in environmental response pathways as well as losses of transcriptions factors important for environmental response. These data provide key insights into the biology of <em>E</em>. sp. EZ as well as the diversification of <em>Echinometra</em> more widely and will serve as a useful tool for the community to explore questions in this taxonomic group and beyond.</p>

opencc-zeroSep 2022View details →
dryad36/100

De novo genome assembly of Kallima inachus

<p><span>Oakleaf butterflies in the genus <em>Kallima</em> have a polymorphic wing phenotype, enabling these insects to masquerade as dead leaves. By studying mechanisms that shape the genetic and species diversity of these butterflies, a new perspective can be provided to understand the evolutionary innovation driven by geographic changes and natural selection.</span></p> <p><span>We found that leaf wing polymorphism in <em>Kallima</em> butterflies is controlled by the wing patterning gene cortex. We hypothesized that multiple mechanisms may independently lead to the reduction or suppression of recombination among different cortex haplotypes. To test this hypothesis, w</span>e performed Nanopore re-sequencing and <em>de novo</em> genome assembly for 4 <em>Kallima inachus</em> individuals and obtained 4 individual genomes. We identified two chromosomal inversions spanning these haplotypes.</p>

opencc-zeroAug 2022View details →
zenodo36/100

Test data for the Large Genome Assembly tutorial

<p>A set of test data to use for the Galaxy Training Network tutorial, Large genome assembly.&nbsp;This data is publicly available in other sources, but has been combined here and subsampled for easier use in the tutorial.&nbsp;We do not claim ownership of this data - please see the full attribution to each of the data sources explained below.</p> <p><strong>Sequencing reads:</strong></p> <p>From the Snow gum: <em>Eucalyptus pauciflora</em>. From NCBI BioProject number: PRJNA450887; Paper: Wang W, Das A, Kainer D, Schalamun M, Morales-Suarez A, Schwessinger B, Lanfear R; 2020, doi: 10.1093/gigascience/giz160.</p> <p>From NCBI, three read files were imported into Galaxy for this tutorial: nanopore reads (SRR7153076), and paired Illumina reads (SRR7153045). For the test data set: these were randomly subsampled to 10% of the original file size, and reads mapping to related chloroplast gene sequences (rbcL sequence: accession KM360776.1; matK sequence: accession KT632904.1) were excluded.&nbsp;</p> <p>Files:&nbsp; Nanopore reads; Illumina reads, R1 and R2</p> <p><strong>Reference genome:&nbsp;</strong></p> <p><em>Arabidopsis thaliana.&nbsp;</em>Although this is not the same species as above, we can use it as an example for a comparison step in the tutorial.&nbsp;This has been downloaded from The Arabidopsis Information Resource at <a href="https://www.arabidopsis.org/index.jsp">https://www.arabidopsis.org/index.jsp</a>&nbsp;from Genes: Download: TAIR10 genome release: TAIR10 chromosome files: file TAIR10_chr_all.fas.gz. Then unzipped into a fasta file.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Genome assemblies and respective cgMLST profiles of a diverse dataset comprising 1,874 Listeria monocytogenes isolates

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 1,748-loci core-genome (cg) Multiple Locus Sequence Type (MLST) profiles [Pasteur schema (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>) available in <a href="https://chewbbaca.online/species/6/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)]&nbsp; of a final set of 1,874 <em>Listeria monocytogenes</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of Sequence Type [ST]). In total, 204 different STs are represented in this dataset, with ST121, ST6, ST9, ST1 and ST155 being in the top 5 and, together, corresponding to 37.9% of the dataset.</p> <p>File &ldquo;Lm_metadata.xlsx&rdquo; contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST.</p> <p>The directory &ldquo;assemblies/&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.&nbsp;</p> <p>The file &ldquo;profiles/Lm_profile.tsv&rdquo; corresponds to a tab separated file with the 1,748-loci cgMLST profile of each isolate presented in the metadata file. These profiles were determined as explained below.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>L. monocytogenes </em>genome assemblies, we collected information about the genetic diversity (STs) of the isolates available at <a href="https://bigsdb.pasteur.fr/listeria/">BIGSdb-Lm</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,957 samples associated with three previous studies (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>; <a href="https://pubmed.ncbi.nlm.nih.gov/28827366/">Maury et al. 2017</a>; <a href="https://pubmed.ncbi.nlm.nih.gov/30775964/">Painset et al. 2019</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,874 isolates passed the dataset curation step and were included in the final dataset. cgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 1,748-loci Pasteur schema (<a href="https://pubmed.ncbi.nlm.nih.gov/27723724/">Moura et al. 2016</a>) available in <a href="https://chewbbaca.online/species/6/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on June 23<sup>rd</sup>, 2022.</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Twin colony isolate genome assemblies: Chlamydomonas 3112/3222 and WS3/WS7 and Scenedesmus ARA/ARA3

<p>Genomes and assemblies for twin isolate pairs from:</p> <p>David R. Nelson, Amphun Chaiboonchoe, Weiqi Fu, Khaled M. Hazzouri, Ziyuan Huang, Ashish Jaiswal, Sarah Daakour, Alexandra Mystikou, Marc Arnoux, Mehar Sultana, Kourosh Salehi-Ashtiani,<br>Potential for Heightened Sulfur-Metabolic Capacity in Coastal Subtropical Microalgae,<br>iScience,<br>Volume 11,<br>2019,<br>Pages 450-465,<br>ISSN 2589-0042,<br>https://doi.org/10.1016/j.isci.2018.12.035.<br>(https://www.sciencedirect.com/science/article/pii/S2589004218302657)</p> <p>Abstract: Summary<br>The activities of microalgae support nutrient cycling that helps to sustain aquatic and terrestrial ecosystems. Most microalgal species, especially those from the subtropics, are genomically uncharacterized. Here we report the isolation and genomic characterization of 22 microalgal species from subtropical coastal regions belonging to multiple clades and three from temperate areas. Halotolerant strains including Halamphora, Dunaliella, Nannochloris, and Chloroidium comprised the majority of these isolates. The subtropical-based microalgae contained arrays of methyltransferase, pyridine nucleotide-disulfide oxidoreductase, abhydrolase, cystathionine synthase, and small-molecule transporter domains present at high relative abundance. We found that genes for sulfate transport, sulfotransferase, and glutathione S-transferase activities were especially abundant in subtropical, coastal microalgal species and halophytic species in general. Our metabolomics analyses indicate lineage- and habitat-specific sets of biomolecules implicated in niche-specific biological processes. This work effectively expands the collection of available microalgal genomes by &sim;50%, and the generated resources provide perspectives for studying halophyte adaptive traits.<br>Keywords: Global Nutrient Cycle; Phycology; Genomics; Metabolomics</p> <p>https://www.sciencedirect.com/science/article/pii/S2589004218302657</p> <p>&nbsp;</p> <p>Files are as follows:</p> <p>&nbsp;</p> <p>.fa = assembly</p> <p>.tr.fa = coding sequences</p> <p>.aa.fa = predicted proteins</p> <p>.gff = annotation files</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Haematococcus lacustris genome assembly and functional annotation

<p>Nuclear, chloroplast and mitochondrial genome assembly and functional annotation of Haematococcus lacustris (formely Haematococcus pluvialis) K-0084</p>

opencc-by-4.0May 2024View details →
dryad36/100

Genome assembly variation and its implications for gene discovery in nematode species

<p>Genome assemblers are a critical component of genome science, but the choice of assembly software and protocols can be daunting. Here, we investigate genome assembly variation and its implications for gene discovery across three nematode species—<em>Caenorhabditis bovis</em>, <em>Haemonchus contortus</em>, and <em>Heligmosomoides bakeri</em>—highlighting the critical interplay between assembly choice and downstream genomic analysis. Selecting popular genome assemblers, we generated multiple assemblies for each species, analyzing their structure, completeness, and effect on gene family analysis. Our findings demonstrate that assembly variations can significantly affect gene family composition, with notable differences in critical gene families like <em>cyp</em>, <em><u>gst</u></em>, <em>ugt</em>, and <em>nhr</em>. Despite broadly similar performance using various assembly metrics, comparisons of assemblies with a single species revealed underlying structural rearrangements and inconsistencies in gene content. This emphasizes the imperative for continuous refinement of genomic resources. Our findings advocate for a cautious and informed approach to genome assembly and annotation to ensure reliable and insightful genomic interpretations.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Genome assemblies of BMCs from the study "Unlocking the genomic potential of Red Sea coral probiotics"

<p>CheckM evaluation report and collection of all BMC isolate genome assemblies (in fasta format) from "Unlocking the genomic potential of Red Sea coral probiotics" study</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Genome assembly of Triticum turgidum subsp. durum cv. Langdon

<p><strong>Summary of the datasets</strong></p> <p>Durum wheat (<em>Triticum turgidum</em> L. subsp. <em>durum</em> (Desf.) Husnot) cultivar Langdon is an experimental standard strain that has been used as a parental strain to produce chromosome substitution lines and synthetic hexaploid wheat. We maintain 'Langdon' pure line (strain No.: LPGKU2272) under National BioResource Project Wheat (NBRP-Wheat) by self-pollination.</p> <p>We constructed a genome assembly of Langdon from about 252 Gbp of HiFi reads using Hifiasm v0.19.8-r603 with additional options '-l 0 -f 39'. The assembly consists of 4,391 contigs (total size: 10,497,834,563 bp, N50: 27,495,971 bp).</p> <p>Assembly: Triticum_durum.Langdon.hifiasm_assembly.v0.1.fa.gz</p> <p>We further performed reference-guided scaffolding to assign the contigs into 14 chromosomes of tetraploid wheat using RagTag v2.1.0 software with aligner option 'unimap'. In this scaffolding process, the public sequence of durum wheat cv. Svevo (Svevo.v1; Maccaferri et al., Nat. Genet., 2019) is used as reference genome.</p> <p>Scaffolded sequence: Triticum_durum.Langdon.ragtag_scaffold.v0.1.fa.gz</p> <p><strong>Acknowledgement</strong></p> <p>This work has been conducted under National BioResource Project (NBRP), Ministry of Education, Culture, Sports, Science and Technology, Japan.</p>

openMay 2024View details →
zenodo36/100

Fathi Camel Microbiome Project (FCMP) Fecal Metagenome-assembled Genomes (MAGs)

<p>The Fathi Camel Microbiome Project (FCMP) aims to characterize the diversity and phenotypic associations of the dromedary camel microbiome. The gut microbiome of N = 55 camels was deeply sequenced via dropped stool. The raw reads, after QC, were assembled and binned into metagenome-assembled genomes (MAGs). We include here a collection of 3165 medium-quality or higher prokaryotic MAGs by MiMAG-like criteria (completeness &gt;= 50%, contamination &lt;= 5%).&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Genome and annotation files for Blumeria graminis f. sp. tritici isolate CHVD042201 (genome assembly: Bgt_CHVD042201 _genome_v1)

<p>Genome and annotation files for Blumeria graminis f. sp. tritici isolate CHVD042201 (genome assembly: Bgt_CHVD042201 _genome_v1)</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Chromosome-level genome assembly of Cyamophila willieti (Hemiptera: Psyllidae)

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo36/100

Chromosome-level genome assemblies of sunflower oilseed and confectionery cultivars

<p>In this study, we obtained high-quality genomes of two cultivar representatives of oil and confectionery common sunflower (Helianthus annuus L.) lineages in China at the chromosome level using the PacBio Revio system Circular Consensus Sequencing and high-throughput chromatin conformation capture (Hi-C) scaffolding sequencing technologies. OXS is an inbred oil-type sunflower line with high kernel rate (74.5% kernels), contains 42.9% oil and fat. It is highly susceptible to Verticillium wilt and moderately susceptible to several other diseases. YDS is an inbred non-oil line with high plant height (180-220 cm) and large plump seeds (19.23 grams per 100 seeds). The genome assembly of OXS, spans 3.03 Gb, with 99.58% of sequences anchored to 17 chromosomes and a contig N50 length of 154 Mb. Similarly, the assembly size of YDS, is 3.02 Gb, with 99.40% of sequences mapped to 17 chromosomes and a contig N50 length of 153 Mb. The gene completeness of BUSCO reached 98.2% for OXS and 98.4% for YDS, while the LTR Assembly Index (LAI) stood at 24.73 and 25.85 for OXS and YDS, respectively. A comparative genomics approach identified 6,535 (OXS) and 6,498 (YDS) gene families that have evolved rapidly and are associated with substance synthesis, cell growth, grain weight, and protective mechanisms against biotic and abiotic stresses. We discovered that the YDS genome assembly shows high collinearity with the OXS assembly, apart from three significant inversions on chromosomes 7 and 17. We also identified 15,056 large deletions and insertions between the OXS and YDS assemblies. The publication of these genomes has greatly contributed to the improvement of genetic breeding by integrating internal genetic and external environmental factors in Helianthus annuus L. crops.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Genome assembly of E. coli C : Final Assembly

<p>This record contains the following file:</p> <ol> <li><code>`Ecoli_C_assembly.fna`</code>&nbsp;- finished assembly in FASTA format. Contains two sequences: nuclear genome of&nbsp;<em>E. coli&nbsp;</em>C and genomes of bacteriophage phiX174, which was used as a spike-in</li> </ol> <p>&nbsp;Complete data including the raw data can be found <a href="https://doi.org/10.5281/zenodo.1257429">here</a>.</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo36/100

Supplemental files for the manuscript "Chromosome assembly of large and complex genomes using multiple references"

<p>This archive contains supplemental&nbsp;files for the manuscipt: &quot;Chromosome assembly of large and complex genomes using multiple references&quot;.</p> <p>It contains assemblies generated by Ragout and RACA as well as evaluation scripts that were used in our analysis.</p> <p>Each subdirectory contains an additional README file with details.</p> <p>Please note that some intermediate files were deleted&nbsp;in the interest&nbsp;of saving space. If you need access to&nbsp;those files or having&nbsp;issues with reproducing our results, don&#39;t hesitate to contact Mihkail Kolmogorov: fenderglass@gmail.com</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

Annotations of sapSuaA1 genome assembly

<p>Annotations accompanying paper &quot;Genome Sequence of Flavor-Producing Yeast Saprochaete suaveolens NRRL Y-17571&quot;</p>

openother-openFeb 2019View details →
zenodo36/100

Genomes assemblies presented in Zhou et al., 2018, Current Biology 28, 2420–2428

<p>Snapshot of genome assemblies published as part of Zhou et al., 2018, Current Biology 28, 2420&ndash;2428. Please see <a href="http://dx.doi.org/10.1016/j.cub.2018.05.058">the paper</a> for full description and list of authors.</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Annotations of sapFunA1 genome assembly

<p>Annotations accompanying paper &quot;Genome Sequence of an Arthroconidial Yeast Saprochaete fungicola CBS 625.85&quot;</p>

openother-openFeb 2019View details →
zenodo36/100

Dataset for "Nanopore-based genome assembly and the evolutionary genomics of basmati rice"

<p><strong>Description of uploaded files:</strong></p> <p>Basmati334.basmati.not_scaffolded.fa</p> <p>- Polished genome assembly for Basmati 334 but not scaffolded.</p> <p>&nbsp;</p> <p>Basmati334.basmati.not_scaffolded.sorted.gff</p> <p>- Gene annotation for the assembly&nbsp;Basmati334.basmati.not_scaffolded.fa</p> <p>&nbsp;</p> <p>Basmati334.basmati.ragoo_scaffold.fa</p> <p>- Polished genome assembly for Basmati 334 and scaffolding with RaGOO using the Nipponbare RAPDB1.0 as reference genome.</p> <p>&nbsp;</p> <p>Basmati334.basmati.ragoo_scaffold.sorted.gff</p> <p>-&nbsp;Gene annotation for the assembly&nbsp;Basmati334.basmati.ragoo_scaffold.fa</p> <p>&nbsp;</p> <p>Basmati334.basmati.ragoo_scaffold.repeatmasker.bed</p> <p>- Repetitive DNA coordinates for the&nbsp;assembly&nbsp;Basmati334.basmati.ragoo_scaffold.fa</p> <p>&nbsp;</p> <p>CONSEL.tar.gz</p> <p>- Files used for CONSEL analysis</p> <p>&nbsp; &nbsp;- Folder&nbsp;CONSEL/PHYLOGENY_TEST/ contains the input files for CONSEL</p> <p>&nbsp; &nbsp;- Folder&nbsp;CONSEL/CONSEL_RESULT/ contains the CONSEL test results</p> <p>&nbsp;</p> <p>DADI_ANALYSIS.tar.gz</p> <p>- Input file for dadi analysis and scripts used for dadi modeling</p> <p>&nbsp;</p> <p>DomSufid.sadri.not_scaffolded.fa</p> <p>- Polished genome assembly for Dom Sufid but not scaffolded.</p> <p>&nbsp;</p> <p>DomSufid.sadri.not_scaffolded.sorted.gff</p> <p>- Gene annotation for the assembly&nbsp;DomSufid.sadri.not_scaffolded.fa</p> <p>&nbsp;</p> <p>DomSufid.sadri.ragoo_scaffold.fa</p> <p>-&nbsp;Polished genome assembly for Dom Sufid and scaffolding with RaGOO using the Nipponbare RAPDB1.0 as reference genome.</p> <p>&nbsp;</p> <p>DomSufid.sadri.ragoo_scaffold.sorted.gff</p> <p>- Gene annotation for the assembly DomSufid.sadri.ragoo_scaffold.fa</p> <p>&nbsp;</p> <p>DomSufid.sadri.ragoo_scaffold.repeatmasker.bed</p> <p>- Repetitive DNA coordinates for the&nbsp;assembly&nbsp;DomSufid.sadri.ragoo_scaffold.fa</p> <p>&nbsp;</p> <p>Four_rice_population.vcf.gz</p> <p>- Filtered SNP VCF file used in the basmati population relationship with japonica and aus.</p> <p>&nbsp;</p> <p>MULTIZ_ALIGNMENT.tar.gz</p> <p>- Reference genome alignment using&nbsp;Nipponbare RAPDB1.0 as reference and aligning various Oryza de novo genome assemblies</p> <p>&nbsp;</p> <p>Multi_Oryza_gene_FASTAs.tar.gz</p> <p>- Using the alignments from MULTIZ_ALIGNMENT/ pulled out coding DNA sequences of each&nbsp;Nipponbare RAPDB1.0 gene</p> <p>&nbsp;</p> <p>Obarthii_outgroup_AlignedToBasmatiScaffolded_genome.fa</p> <p>-&nbsp;O. barthii reference genome sequence was aligned to scaffolded Basmati 334 reference genome. For every Basmati 334 genome coordinate was converted into a O. barthii&nbsp;sequence resulting in a basmati-ized&nbsp;O. barthii&nbsp;genome sequence. Not alignable regions were indicated as &#39;N&#39;.&nbsp;</p> <p>&nbsp;</p> <p>Only_basmati_rice_population.vcf.gz</p> <p>-&nbsp;Filtered SNP VCF file used in the basmati population analysis.</p> <p>&nbsp;</p> <p>Oryza_LTR_DivergenceTime.txt</p> <p>- LTR retrotransposon annotated in various Oryza reference genomes and their estimated insertion time (based on the divergence between the LTRs).</p> <p>&nbsp;</p> <p>TWISST.tar.gz&nbsp;</p> <p>- TWISST input and results file.</p> <p>&nbsp; &nbsp;-&nbsp;Four_rice_population.geno.gz, genotype file generated from the genomic_general from S. Martin and used as input for TWISST analysis.</p> <p>&nbsp; &nbsp;- *.trees.gz phylogenetic trees generated from sliding windows</p> <p>&nbsp; &nbsp;- *.data.tsv sliding window coordinates</p> <p>&nbsp; &nbsp;- *.weights.csv.gz topology weights</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record