Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo48/100

Effects of crown gall disease on natural microbiota of Vitis vinifera - genome annotations

<p>Young grapevines (Vitis vinifera) frequently die due to the crown gall (CG) disease induced by the plant pathogen Allorhizobium vitis (Rhizobiaceae). Virulent members of A. vitis harbour a tumor-inducing (Ti) plasmid and cause formation of CGs due to genes encoded on the T-DNA. Expression of the oncogenes by transformed host cells induce cell proliferation, metabolic and physiological changes. The CG produces opines uncommon to plants, which provide an important nutrient source for A. vitis harbouring opine catabolism enzymes. CGs host a defined bacterial community and the mechanisms establishing a CG-specific bacterial community are currently unknown. Thus, we were interested in whether genes homologous to those of the Ti-plasmid coexist in the genomes of the microbial species coexisting in CGs. We isolated eight bacterial strains from grapevine CGs, sequenced their genomes and tested their virulence and opine utilization ability in bioassays. In addition, the eight genome sequences were aligned to the sequences of a Ti-plasmid and seven published bacterial genomes, including closely related plant associated bacteria but not from CGs. Homologous genes for virulence and opine anabolism were only present in the virulent Rhizobiaceae. By contrast, homologs of the opine catabolism genes were present in all strains including the non-virulent members of the Rhizobiaceae and non-Rhizobiaceae, indicating horizontal gene transfer of the opine degradation cluster from virulent to non-virulent strains. These results along with those of the opine utilization assay support the important role of opine utilization for co-colonization of virulent and non-virulent bacteria in CGs, thereby shaping the CG community.</p> <p>This dataset contains the prokka annotations of the genomes as used in &quot;Opportunistic bacteria of grapevine crown galls are equipped with the genomic repertoire for opine utilization&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Gene Annotations of 49 Bacillariophyta Genome Assemblies

<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div>&nbsp;</div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div>&nbsp;</div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div>&nbsp;</div> <div>To extract the dataset, execute the following command:</div> <div>&nbsp;</div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div>&nbsp;</div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p>&nbsp;</p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p>&nbsp;</p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p>&nbsp;</p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^&gt;/ s/ .*//' genome.fasta &gt; genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div>&nbsp;</div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div>&nbsp;</div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)

<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a &ldquo;91 eutherian mammals&rdquo; multiple whole genome alignment in Ensembl.&nbsp;</p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells.&nbsp;This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p>&nbsp;</p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Timema genome sequences and annotations. Version 8.

<p>Genome sequence (fasta) files&nbsp;and annotation (gff) files for ten <em>Timema </em>species:&nbsp;<em>T. bartmani, T. cristinae, T. poppensis, T. californicum,&nbsp; T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and&nbsp; T. genevievae.</em><br> <br> Species are abbreviated as follows: Tbi =&nbsp;<em>T. bartmani</em>, Tce =&nbsp;<em>T. cristinae</em>, Tps =&nbsp;<em>T. poppensis</em>, Tcm =&nbsp;<em>T. californicum</em>, Tpa =&nbsp;<em>T. podura</em>, Tte =&nbsp;<em>T. tahoe</em>, Tms =&nbsp;<em>T. monikensis</em>, Tdi =&nbsp;<em>T. douglasi</em>, Tsi =&nbsp;<em>T. shepardi</em>, and Tge =&nbsp;<em>T. genevievae</em><br> &nbsp;</p> <p>For details of assembly and annotation see:&nbsp;<br> <br> Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, &nbsp;Z., Figuet, E., Fran&ccedil;ois, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, &nbsp;M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540</p> <p>&nbsp;</p> <p><strong>File list:</strong><br> <br> Tbi_b3v08.fasta = T. bartmani genome sequence file<br> Tbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation file<br> Tce_b3v08.fasta = T. cristinae genome sequence file<br> Tce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation file<br> Tcm_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. bartmani genome sequence file<br> Tcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation file<br> Tdi_b3v08.fasta = T. douglasi genome sequence file<br> Tdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation file<br> Tge_b3v08.fasta = T. genevievae genome sequence file<br> Tge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation file<br> Tms_b3v08.fasta = T. monikensis genome sequence file<br> Tms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation file<br> Tpa_b3v08.fasta = T. podura genome sequence file<br> Tpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation file<br> Tps_b3v08.fasta = T. poppensis genome sequence file<br> Tps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation file<br> Tsi_b3v08.fasta = T. shepardi genome sequence file<br> Tsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation file<br> Tte_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. tahoe genome sequence file<br> Tte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Enhanced genome annotation strategy provides novel insights on the phylogeny of 'Flaviviridae': Supplementary material

<p>SUPPLEMENTARY MATERIAL</p> <p><strong>Index</strong></p> <ul> <li> <p>Table S1 (tableS1.csv): genomic data.</p> </li> <li> <p>Table S2 (tableS2.csv): character categorization for selected nodes.</p> </li> <li> <p>Table S3 (tableS3.csv): programs and parameters.</p> </li> <li> <p>Table S4 (tableS4.csv): annotation efficiency.</p> </li> <li> <p>File S1 (fileS1.gff): gene annotation.</p> </li> <li> <p>File S2 (fileS2.xml): configuration file for BEAST 2 (configuration.xml).</p> </li> <li> <p>Figure S1 (figureS1.pdf): dendrogram depicting the hierarchical clusters of trees based on match-split distances.</p> </li> <li> <p>Figure S2 (figureS2.pdf): full version of the working phylogenetic hypothesis (tree No. 0 in table 1).</p> </li> </ul> <p><strong>Figure captions</strong></p> <ul> <li>Figure S1: A dendrogram depicting the hierarchical clusters of trees based on match-split distances.&nbsp;Outgroup sequences (<em>Hepacivirus</em>, <em>Pegivirus</em>, and <em>Pestivirus</em>) were removed to guarantee the compared tree topologies would have the same terminals. Tree numbers correspond to those in table 1 of the manuscript. I. No outgroup sequences; some matrices were partitioned. II. Outgroup sequences and partitioned matrices. *This tree was produced without outgroup sequences.</li> <li>Figure S2: Full version of the working phylogenetic hypothesis (tree No. 0 in table 1). Branch lengths represent an estimation of the number of substitutions per site. Node labels indicate SH-aLRT support / ultrafast bootstrap (only shown if one of there is a value&nbsp;below 90%). Clade names correlate to the character categorization analysis (see table S2). Branch labels represent the four genera: I = <em>Pestivirus</em>; II = <em>Pegivirus</em>; III = <em>Hepacivirus</em>; IV = <em>Flavivirus</em>. *&nbsp;The Ecuador Paraiso Escondido virus (EPEV) was isolated from sand flies (<em>Psathyromyia abonnenci</em>). The EPEV was the first sand fly-borne <em>Flavivirus</em> identified in the New World.</li> </ul> <p><strong>Manuscript title</strong></p> <p>FLAVi: an enhanced annotator for viral genomes of <em>Flaviviridae</em>.</p> <p><strong>Authors</strong></p> <ul> <li> <p>de Bernardi Schneider, Adriano. University of California San Diego. ORCID: 0000-0001-7487-266X.</p> </li> <li> <p>Jacob Machado, Denis. University of North Carolina at Charlotte. ORCID: 0000-0001-9858-4515. Corresponding author.</p> </li> <li> <p>Guirales, Sayal.&nbsp;University of North Carolina at Charlotte.</p> </li> <li> <p>Janies, Daniel. University of North Carolina at Charlotte.</p> </li> </ul> <p><em>First author</em>: Adriano de Bernardi Schneider and Denis Jacob Machado have contributed equally to the manuscript.</p> <p><strong>Contact information</strong></p> <ul> <li> <p>Corresponding author: Denis Jacob Machado, Ph.D.</p> </li> </ul> <ul> <li> <p>OrcID: 0000-0001-9858-4515.</p> </li> </ul> <ul> <li> <p>Email: dmachado [at] uncc.edu.</p> </li> </ul> <p><strong>Other additional material</strong></p> <ul> <li>In addition to the material listed above, all 31 tree topologies and 15 alignment matrices discussed in this manuscript will are available in TreeBASE (<a href="http://purl.org/phylo/treebase/phylows/study/TB2:S24096">http://purl.org/phylo/treebase/phylows/study/TB2:S24096</a>) after the publication of the manuscript.</li> <li>The FLAVi pipeline and all the original scripts are available at GitLab (<a href="https://gitlab.com/MachadoDJ/FLAVi">https://gitlab.com/MachadoDJ/FLAVi</a>).</li> <li>The web application can be accessed at <a href="http://flavi-web.com">http://flavi-web.com</a>.</li> </ul>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Genome annotation workflow for Effrenium voratum RCC1521

<p>Scripts of complete genome annotation workflow for Effrenium voratum RCC1521, associated with the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in&nbsp;<em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p>See <strong>README_EvRCC1521.txt</strong> for more detail.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Penicillium fuscoglaucum Pf_T2 Genome Assembly and Annotation

<p>During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum&nbsp;Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of 5 scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum, and contained a lineage specific, intact cylcopaizonic acid (CPA) gene cluster.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Genome and annotations of cotton rat (Sigmodon hispidus)

<p>The chromosome level reference genome of <em>Sigmodon hispidus</em> based on third-generation high fidelity (HiFi) reads, high-throughput chromosome conformation capture (Hi-C), and second-generation sequencing techniques.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Actinidia chinensis Red5 genome assembly (version 2) and annotation files

<p>We present version 2 of the genome assembly for <em>Actinidia chinensis</em> var. <em>chinensis</em> genotype Red5. The Red5 genome was originally assembled using short read Illumina data (Pilkington et al, 2018; <a href="https://doi.org/10.1186/s12864-018-4656-3">https://doi.org/10.1186/s12864-018-4656-3</a>). In version 2 we employed Pacific BioSciences Sequel Single Molecule Real Time (SMRT) sequencing technology in place of Illumina paired end read sequencing for the main assembly but leveraged that short read data (Pilkington et al, 2018) for post assembly base correction of long read assembly contigs. Additionally the Illumina long insert libraries from Pilkington et al (2018) were used for post assembly scaffolding of contigs. Scaffold assignment to linkage groups leveraged the genetic map described in Pilkington et al (2018) as well as consensus evidence from DNA synteny comparisons to existing whole genome sequences from <em>Actinidia</em>.</p> <p>To meet the file size restrictions some dataset components have been split into multiple parts.</p> <p><strong>Assembly</strong></p> <p>The assembly work flow used the FALCON/FALCON-unzip assembly suite is described in Red5_version_2_genome_assembly.md. The assembly yielded both primary and haplotig contig data sets, the metrics for which are documented in this file. The CDS and predicted peptide fasta and GFF3 gene annotation for the primary and haplotig sets are provided in separate files.</p> <p><strong>File Descriptions</strong></p> <ul> <li>Files named chr1.fasta to chr29.fasta represent the primary assembly linkage group level assembly units</li> <li>Files named haplotig_part_1.fasta to haplotig_part_10.fasta represent the haplotig contig sets split into 10 parts to meet upload file size restrictions</li> <li>Files named primary_assembly.primary.gff3 and haplotig.gff3 contain the gene model annotations for the primary and haplotig assembly datasets respectively</li> <li>primary_assembly.cds.fasta and primary_assembly.pep.fasta contain the CDS and peptide sequences for the annotations on the primary contigs</li> <li>haplotig.cds.fasta and haplotig.pep.fasta contain the CDS and peptide sequences for the annotations on the haplotig contigs</li> <li>haplotigs.placements.tsv and haplotigs.reassignments.tsv describe the placement of haplotigs relative to the primary contigs as derived from purge_haplotigs</li> <li>The file Red5_version_2_genome_assembly.md describes the assembly work flow and code steps used as well as assembly metrics</li> <li>Files&nbsp;HYV3_1.v.R5V2_1.png to&nbsp;HYV3_29.v.R5V2_29.png depict Circos plots of DNA:DNA synteny based on 1coords alignment filter of nucmer alignments using dnadiff</li> </ul> <p>See Red5_version_2_genome_assembly.md for description of assembly methods and assembly metrics.</p> <p><strong>Funding</strong></p> <p>This work was funded by Kiwifruit Royalty Investment Program by The New Zealand Institute for Plant &amp; Food Research Ltd. with support from Zespri, and the CORE grant Endeavour Smart Idea Fund (UOOX1801) from the New Zealand Ministry of Business, Innovation and Employment (MBIE). The funding bodies had no role in the design of the study, the collection, analysis, or interpretation of data or writing this manuscript.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Annotation of the the assembled genome of Fusarium oxysporum f. sp. albedinis strain 133, the causal agent of date palm dieback.

<p>Annotation of&nbsp;the the assembled genome of <em>Fusarium oxysporum f. sp. albedinis</em> strain 133 (Khayi et al., 2020). Gene prediction and annotation were carried out using funnotate pipeline v1.8.1 (Stajich, 2020), which&nbsp;includes masking, ab initio gene-prediction training, using Augustus and Genmark, with the EST dataset&nbsp;reported to the Ganoderma mycocosm repository, gene prediction, and the assignment of functional&nbsp;annotation to protein-coding gene models.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Gila monster (Heloderma suspectum) genome assembly and annotation

<p><em>De novo</em>&nbsp;genome assembly and annotation of a male Gila monster (<em>Heloderma suspectum</em>). We annotated the genome&nbsp;using the Comparative Annotation Toolkit (CAT), and we have also included GFF3 files of the consensus gene set,&nbsp;output for each taxon included in this process.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

The first annotated genome assembly of Macrophomina tecta associated with charcoal rot of sorghum

<p>Raw reads of Macrophomina tecta were obtained from Nanopore, Illumina,&nbsp;and NextSeq (RNA). Files with information about the genome annotation, functional prediction, repeats, effectors and orthologous genes are included.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Drosophila willistoni genome annotation

<p>Genome annotation of Drosophila willistoni de novo assembly using long reads. Protein coding genes were predicted with Funannotate. This software aligned proteins and transcripts from the Drosophila willistoni Flybase annotation against the assembly using minimap2, diamond and exonerate. It processed these hints to be included by augustus when predicting protein coding genes.</p> <p>LncRNAs were identified by FEELnc using stranded ribodepleted RNA-seq libraries from ovaries, testes, male accessory glands and whole body samples. Additional lncRNAs were identified by mapping lncRNAs sequences downloaded from RNA central identified for Drosophila willistoni.</p> <p>Ribosomal RNAs were predicted with RNAmmer; tRNAs, with tRNAscan2; and other miscellaneous ncRNA, with cmsearch and Rfam models.</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

"Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan" - Genomes and annotation

<p>File-dump associated with the manuscript:</p> <p>&quot;<strong>Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan&quot; (Not yet published)</strong></p> <p>MGVs refer to the viral genome database in the publication:&nbsp;https://www.nature.com/articles/s41564-021-00928-6&nbsp;</p> <p>&nbsp;</p> <p>Following uploaded:</p> <p>File 1: VOG Markers in vOTUs/vMAGs and MGV genomes</p> <p>File 2: Viral Tree Newick&nbsp;file with vOTUs/vMAGs and MGV genomes</p> <p>File 3: All vOTUs/vMAGs genomes</p> <p>File 4: Master table annotation of vOTUs/vMAGs</p> <p>File 5: Centenarian bacterial isolate proviruses</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5

<p>Gene annotation files for&nbsp;<em>Fraxinus excelsior</em>&nbsp;genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These&nbsp;files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB:&nbsp;The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>

opencc-by-4.0Dec 2016View details →
zenodo44/100

Annotation and Orthofinder results for three Mytilus species genomes.

<p>Annotations and Orthofinder results for three Mytilus species genomes: accessions JAKGDF000000000 (MgalMED), JAKGDG000000000 (MeduEUS), and JAKGDH000000000 (MeduEUN).</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

New annotation of the Lolium perenne genome described by Bryne et al, (2015)

<p>Annotation of the Lolium perenne genome described by Bryne <em>et al,</em>&nbsp;(2015, DOI: 10.1111/tpj.13037).</p> <p>To identify genic regions RNA sequencing data were aligned to the genome using Tophat (Tophat version: V2.0.11; Bowtie2 version: 2.1.0). Isoforms, of genes, were identified using Cufflinks (Version: 2.2.0). Open reading frames (ORF), were found using using program ORFpredictor (version: 3.0). Frame selection was assisted by BLASTX searching the proteomes of <em>Arabidopsis thaliana</em> (TAIR, version: 10), <em>Oryza sativa</em> (Ensembl)<em>, Gycine max </em>(Ensembl)<em>, Populus trichocarpa </em>(Ensembl)<em> and Manihot esculenta </em>(cassava, v4.1). The predicted CDS was back translated to annotate the GFF file created by Cufflinks for CDS using scripts kindly provided by Palmieri<em> et al.,</em> 2012 (doi: 10.1371/journal.pone.0046415). These results are included in the file LG_V2_full.gtf.</p> <p>Functional annoatation was using three sources. First, protein sequences were search against the <em>A. thaliana</em> proteome using BLASTP. Second, the proteins were search against the Swiss-Prot non-redundant protein database (<a href="http://www.uniprot.org/downloads">http://www.uniprot.org/downloads</a> downloaded 14/03/2016, UniProt Consortium, 2014), again using BLASTP. In the third step, the protein sequences were scanned against InterPro&#39;s signatures using InterProScan (Version: 5.16-55). These data are included in the file ALLXLOC.txt.</p>

opencc-by-4.0Mar 2018View details →
zenodo44/100

Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes

<p>&nbsp;</p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p>&nbsp;</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; RELEASE MAG-2018/01<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --------------------------------------</p> <p>&nbsp;</p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p>&nbsp;</p> <p>2. LOCATION</p> <p>&nbsp;</p> <p>&nbsp;&nbsp; MAGs sequences are available under ENA project PRJEB25176.</p> <p>&nbsp;&nbsp;</p> <p>&nbsp;&nbsp;&nbsp; ENA_accession&nbsp;&nbsp; Isolate<br> &nbsp;&nbsp; &nbsp;--------------------&nbsp;&nbsp; &nbsp;--------------<br> &nbsp;&nbsp; &nbsp;ERZ494218&nbsp;&nbsp; &nbsp;AM_0118<br> &nbsp;&nbsp; &nbsp;ERZ494219&nbsp;&nbsp; &nbsp;AM_0219<br> &nbsp;&nbsp; &nbsp;ERZ494220&nbsp;&nbsp; &nbsp;AM_0226<br> &nbsp;&nbsp; &nbsp;ERZ494221&nbsp;&nbsp; &nbsp;AM_0228<br> &nbsp;&nbsp; &nbsp;ERZ494222&nbsp;&nbsp; &nbsp;AM_0233<br> &nbsp;&nbsp; &nbsp;ERZ494223&nbsp;&nbsp; &nbsp;AM_0240<br> &nbsp;&nbsp; &nbsp;ERZ494224&nbsp;&nbsp; &nbsp;AM_0244<br> &nbsp;&nbsp; &nbsp;ERZ494225&nbsp;&nbsp; &nbsp;AM_0256<br> &nbsp;&nbsp; &nbsp;ERZ494226&nbsp;&nbsp; &nbsp;AM_0268<br> &nbsp;&nbsp; &nbsp;ERZ494227&nbsp;&nbsp; &nbsp;AM_0275<br> &nbsp;&nbsp; &nbsp;ERZ494228&nbsp;&nbsp; &nbsp;AM_0466<br> &nbsp;&nbsp; &nbsp;ERZ494229&nbsp;&nbsp; &nbsp;AM_0507<br> &nbsp;&nbsp; &nbsp;ERZ494230&nbsp;&nbsp; &nbsp;AM_0510<br> &nbsp;&nbsp; &nbsp;ERZ494231&nbsp;&nbsp; &nbsp;AM_0519<br> &nbsp;&nbsp; &nbsp;ERZ494232&nbsp;&nbsp; &nbsp;AM_0528<br> &nbsp;&nbsp; &nbsp;ERZ494233&nbsp;&nbsp; &nbsp;AM_0546<br> &nbsp;&nbsp; &nbsp;ERZ494234&nbsp;&nbsp; &nbsp;AM_0608<br> &nbsp;&nbsp; &nbsp;ERZ494235&nbsp;&nbsp; &nbsp;AM_0615<br> &nbsp;&nbsp; &nbsp;ERZ494236&nbsp;&nbsp; &nbsp;AM_0616<br> &nbsp;&nbsp; &nbsp;ERZ494237&nbsp;&nbsp; &nbsp;AM_0619<br> &nbsp;&nbsp; &nbsp;ERZ494238&nbsp;&nbsp; &nbsp;AM_0621<br> &nbsp;&nbsp; &nbsp;ERZ494239&nbsp;&nbsp; &nbsp;AM_0630<br> &nbsp;&nbsp; &nbsp;ERZ494240&nbsp;&nbsp; &nbsp;AM_0643<br> &nbsp;&nbsp; &nbsp;ERZ494241&nbsp;&nbsp; &nbsp;AM_0729<br> &nbsp;&nbsp; &nbsp;ERZ494242&nbsp;&nbsp; &nbsp;AM_0764<br> &nbsp;&nbsp; &nbsp;ERZ494243&nbsp;&nbsp; &nbsp;AM_0832<br> &nbsp;&nbsp; &nbsp;ERZ494244&nbsp;&nbsp; &nbsp;AM_0849<br> &nbsp;&nbsp; &nbsp;ERZ494245&nbsp;&nbsp; &nbsp;AM_0854<br> &nbsp;&nbsp; &nbsp;ERZ494246&nbsp;&nbsp; &nbsp;AM_0876<br> &nbsp;&nbsp; &nbsp;ERZ494247&nbsp;&nbsp; &nbsp;AM_0902<br> &nbsp;&nbsp; &nbsp;ERZ494248&nbsp;&nbsp; &nbsp;AM_0936<br> &nbsp;&nbsp; &nbsp;ERZ494249&nbsp;&nbsp; &nbsp;AM_1003<br> &nbsp;&nbsp; &nbsp;ERZ494250&nbsp;&nbsp; &nbsp;AM_1104<br> &nbsp;&nbsp; &nbsp;ERZ494251&nbsp;&nbsp; &nbsp;AM_1111<br> &nbsp;&nbsp; &nbsp;ERZ494252&nbsp;&nbsp; &nbsp;AM_1205<br> &nbsp;&nbsp; &nbsp;ERZ494253&nbsp;&nbsp; &nbsp;AM_1312<br> &nbsp;&nbsp; &nbsp;ERZ494254&nbsp;&nbsp; &nbsp;AM_1409<br> &nbsp;&nbsp; &nbsp;ERZ494255&nbsp;&nbsp; &nbsp;AM_1503<br> &nbsp;&nbsp; &nbsp;ERZ494256&nbsp;&nbsp; &nbsp;AM_1603<br> &nbsp;&nbsp; &nbsp;ERZ494257&nbsp;&nbsp; &nbsp;AM_1606<br> &nbsp;&nbsp; &nbsp;ERZ494258&nbsp;&nbsp; &nbsp;AM_1801<br> &nbsp;&nbsp; &nbsp;ERZ494259&nbsp;&nbsp; &nbsp;AM_1811<br> &nbsp;&nbsp; &nbsp;ERZ494260&nbsp;&nbsp; &nbsp;AM_2104<br> &nbsp;&nbsp; &nbsp;ERZ494261&nbsp;&nbsp; &nbsp;AM_2116<br> &nbsp;&nbsp; &nbsp;ERZ494262&nbsp;&nbsp; &nbsp;AM_2124<br> &nbsp;&nbsp; &nbsp;ERZ494263&nbsp;&nbsp; &nbsp;AM_2202<br> &nbsp;&nbsp; &nbsp;ERZ494264&nbsp;&nbsp; &nbsp;AM_2207<br> &nbsp;&nbsp; &nbsp;ERZ494265&nbsp;&nbsp; &nbsp;AM_2208<br> &nbsp;&nbsp; &nbsp;ERZ494266&nbsp;&nbsp; &nbsp;AM_2324<br> &nbsp;&nbsp; &nbsp;ERZ494267&nbsp;&nbsp; &nbsp;AM_2502<br> &nbsp;&nbsp; &nbsp;ERZ494268&nbsp;&nbsp; &nbsp;AM_2804<br> &nbsp; &nbsp;</p> <p>3. ACKNOWLEDGEMENTS<br> &nbsp; &nbsp;</p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of S&atilde;o Carlos, S&atilde;o Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Cient&iacute;fico e Tecnol&oacute;gico (CNPq), as well as, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; the spanish funding organ Consejo Superior de Investigaciones Cient&iacute;ficas (CSIC).</p> <p>This study was financed in part by the Coordena&ccedil;&atilde;o de Aperfei&ccedil;oamento de Pessoal de N&iacute;vel Superior - Brasil (CAPES) - Finance Code 001.</p> <p>&nbsp;</p> <p>4. CONTACT INFORMATION</p> <p>&nbsp;&nbsp; Current curators:</p> <p>&nbsp;&nbsp; - C&eacute;lio Dias Santos J&uacute;nior (celio.diasjunior@gmail.com)<br> &nbsp;&nbsp; - Flavio Henrique-Silva (dfhs@ufscar.br)<br> &nbsp;&nbsp; - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> &nbsp;</p> <p>5. COPYRIGHT NOTICE</p> <p>&nbsp;&nbsp; Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> &nbsp;&nbsp; Copyright (C) 2018 The AMnrGC consortium.</p> <p>&nbsp;&nbsp; This database is provided &ldquo;as is&rdquo; and without any warranty of any kind,<br> &nbsp;&nbsp; of openly available. You can redistribute and/or modify it<br> &nbsp;&nbsp; as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p>&nbsp;&nbsp; &nbsp;https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

ST131_4071_genome_assembly_annotation_files

<p>The genome assembly annotation files of 4,071 E. coli ST131 genomes (see Decano &amp; Downing 2019).</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

SFB genomes and annotations

<p>This dataset contains sequence files for a&nbsp;Metagenome Assembled Genome (MAG) from human metagenomes, as well as 5&nbsp;SFB reference genomes:</p> <ul> <li>GCF_000270205 Candidatus Arthromitus sp. SFB-mouse-Japan</li> <li>GCF_000283555 Candidatus Arthromitus sp. SFB-rat-Yit</li> <li>GCF_000284435 Candidatus Arthromitus sp. SFB-mouse-Yit</li> <li>GCF_000709435 Candidatus Arthromitus sp. SFB-mouse-NL</li> <li>GCF_001655775 Candidatus Arthromitus sp. SFB-turkey isolate UMNCA01</li> </ul> <p>The dataset consists of 8 gzipped tar archives. Here&#39;s brief summary of their contents:</p> <ul> <li><strong>sfb_abundance</strong>: Counts of mapped reads&nbsp;and normalized counts&nbsp;for each contig in 825 samples (see <strong>sfb_map</strong>)&nbsp;Files named &#39;raw_counts&#39; are number of reads assigned to each contig while files named &#39;tpm&#39; are counts normalized to Transcripts Per Million. The &#39;percontig&#39; files show numbers per contig while raw_counts.tab and tpm.tab files have counts summed for each genome.</li> <li><strong>sfb_abundance.cds</strong>: Counts of mapped reads and normalized counts as above but only for reads mapping to protein-coding regions.</li> <li><strong>sfb_annotations</strong>: Annotation files, from running the prokka pipeline on the genomes and subsequently eggnog-mapper, pfam_scan and&nbsp;dbCAN.</li> <li><strong>sfb_checkm</strong>: Results from running &#39;checkm lineage_wf&#39; on the genomes.</li> <li><strong>sfb_collated</strong>: Collated counts of annotations in each genome.</li> <li><strong>sfb_fastani</strong>: Results from running fastANI on the genomes, with subsequent clustering of genomes based on 75% overlap and 95% ANI.</li> <li><strong>sfb_gtdb</strong>: Results from the &#39;gtdbtk classify_wf&#39; on the genomes. This shows how the genomes are classified against the <a href="https://gtdb.ecogenomic.org/">Genome Taxonomy Database</a>&nbsp;(release86).</li> <li><strong>sfb_gtdb_denovo</strong>: Phylogeny as created using the following command on the genomes.</li> </ul> <pre><code class="language-bash">gtdbtk de_novo_wf --bac120_ms --outgroup_taxon p__Patescibacteria -x .fna --cpus 20 --rnd_seed 123</code></pre> <ul> <li><strong>sfb_map</strong>: Results from mapping reads from 825 samples to the 6 genomes. Reads were aligned using bowtie2 with &#39;--very-sensitive --no-unal&#39; settings and &#39;--score-min C,0,0&#39; to only report reads aligning without mismatches.Output was sorted by position and duplicates removed using MarkDuplicates of the picard tools suite. The archive contains a single merged bam file (&#39;sfb.bam&#39;) where each sample has been assigned a ReadGroup inferred from its file name.&nbsp;Note that this mapping step was performed to investigate the presence of the SFB MAG in other metagenomes and was not part of the actual binning step.</li> <li><strong>SFB.unoise.vsearch.tsv:&nbsp;</strong>Count table of amplified 16S sequence variants with one sample per column and one Amplicon Sequence Variant (ASV)&nbsp;per row. The sixth column shows the assigned taxonomy, and the seventh, the sequence. Total DNA was amplified with the universal &nbsp;bacterial 16S primer pair 341f-805r. Primer sequences and low quality bases were removed from the raw reads with Cutadapt. ASVs were picked using Unoise3 with standard parameters. Taxonomy was assigned by the SINA classifier, based on the SILVA database v132.</li> </ul>

opencc-by-4.0Jun 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record