Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Annotations of sapFunA1 genome assembly

<p>Annotations accompanying paper &quot;Genome Sequence of an Arthroconidial Yeast Saprochaete fungicola CBS 625.85&quot;</p>

openother-openFeb 2019View details →
zenodo36/100

Virulence and antibiotic resistance plasticity of Arcobacter butzleri: insights on the genomic diversity of an emerging human pathogen (genome assembly, annotation dataset, core- and pan-genome loci)

<p>This dataset refers to the analysis of 49 <em>Arcobacter butzleri</em> genomes and includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the predicted&nbsp;transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files), the respective amino acid sequences of the translated CDS sequences (.faa files), the nucleotide alignments of all the 1165 core-genome loci,&nbsp;the nucleotide alignments of the genes <em>hecA</em>, <em>tetR </em>and <em>porA</em>, the categorized amino acid sequences of the six hypervariable regions of PorA, and the nucleotide sequences of the first allele of each of the 7474 pan-genome loci with the respective complete allelic profile matrix.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB34441).</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Updated annotation for Aedes aegypti reference genome AaegL5 with extended 3' UTRs

<p>Updated annotation file for the AaegL5 genome generated and used in the Adavi et al. 2024 <em>bioRxiv </em>preprint: https://doi.org/10.1101/2024.08.21.608847</p> <p>Key updates (to VectorBase-55_AaegyptiLVP_AGWG.gff) include:</p> <ul> <li>Addition of several chemoreceptors that were annotated in previous work</li> <li>Automated extension of 3' UTRs (by up to 750bp) for all genes where supported by antennal neuron snRNAseq data</li> <li>Further manual extension of 3'UTRs for some chemoreceptors where supported by antennal neuron snRNAseq data</li> </ul> <p>For more information on this annotation and the way it was generated, please see the Methods section of the above preprint.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Fusarium Pan-Annotations: improved individual and collective genome annotations

<p>A 'pan-annotation' of genomes from 77 taxa (83 accessions) across the genus <em>Fusarium</em>, delivered as part of the Earlham Institute Strategic Programme Grant 'Decoding Biodiversity' (BBSRC).</p> <p>Citation:</p> <p><a href="https://www.doi.org/10.1101/2025.03.12.642647" target="_blank" rel="noopener">Leveraging existing data to maximise quality and consistency across gene model annotations: a <em>Fusarium</em> pan-annotation. bioRxiv doi:10.1101/2025.03.12.642647</a></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

genome_annotation_exercise

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

Melange (COGs and Pfams) and AntiSMASH (BGCs) annotations of the genome sequence of Lentilitoribacter sp. EG35

<p>Secondary Metabolite Encoding Biosynthetic Gene Cluster (BGC) annotation files from antiSMASH bacterial version 7.1.0 as well as Clusters of Orthologous Groups of proteins (COG) and Protein families (Pfam) annotations from the Melange pipeline (<a href="https://github.com/sandragodinhosilva/melange">https://github.com/sandragodinhosilva/melange</a>) of the genome assembly of <em>Lentilitoribacter </em>sp. strain EG35, isolated from the temperate octocoral <em>Eunicella gazella</em> sampled in the Northeast Atlantic Ocean, Portugal.&nbsp;</p> <p>The data correspond to the genome assembly of EG35 available under the BioProject accession number<a href="https://www.ncbi.nlm.nih.gov/bioproject/1075806">&nbsp;PRJNA1135483</a>.</p> <p>To interactively view the AntiSMASH results, please download and extract the entire content of the AntiSMASH folder, and open the HTML file named "index".</p> <p><strong>This dataset is part of the following study:</strong></p> <p><strong>Tina Keller-Costa, Selene Madureira, Ana S. Fernandes, Lydia Kozma, Jorge M.S. Gon&ccedil;alves, Cristina Barroso, Con&ccedil;eic&atilde;o Egas, &amp; Rodrigo Costa<sup>&nbsp;</sup>2024.&nbsp;Genome sequence of the marine alphaproteobacterium <em>Lentilitoribacter</em> sp. EG35 isolated from the temperate octocoral <em>Eunicella gazella</em>. Microbiology Resource Announcements. MRA00872-24.</strong></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

BIL20 and BIL24 Assembled, Annotated, and Mapped Genomes

<p>The dataset is featured in the article <em>Beyond Buro</em>, which investigates the probiotic potential of <em>Limosilactobacillus fermentum</em> BIL20 and BIL24 using genomic analysis and probiotic assays. These strains were isolated from <em>burong isda</em>, a traditional fermented fish product from Arayat, Pampanga. The study adds to the limited research from the Philippines that applies genomic analysis and probiotic testing on LAB isolates from fermented foods, revealing their promising probiotic attributes and potential health advantages. The dataset includes genome assemblies, RAST and Prokka annotations, and mapped assemblies locating the probiotic-related genes and their functional categories.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Supplementary Tables on Gene Annotations of 49 Bacillariophyta Genome Assemblies

<p>Supplementary Tables attaining to the manuscript entitled&nbsp;<strong>Annotation of protein-coding genes in 49 diatom genomes from the Bacillariophyta clade.&nbsp;</strong>These Supplementary Tables describe in part the data foundation and results of the actual annotation dataset that is available at <a href="https://doi.org/10.5281/zenodo.13767023" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13767023</a>.</p> <p><strong>Supplementary Table S1</strong>: Genome assemblies available at NCBI Datasets in June 2024.</p> <p><strong>Supplementary Table S2</strong>: Genome assemblies excluded from annotation.</p> <p><strong>Supplementary Table S3</strong>: Accession numbers of genome assemblies and RNAseq libraries used for annotating 49 diatom genomes. Table also lists repeat content of genome assemblies after masking with RepeatModeler2/RepeatMasker.</p> <p><strong>Supplementary Table S4</strong>: Software and container versions used for annotating 49 diatom genomes.</p> <p><strong>Supplementary Table S5</strong>: Summary of EnTAP functional annotation results. The actual functional annotations are included in the annotation dataset that is available at <a href="https://doi.org/10.5281/zenodo.13767023" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13767023</a>.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Metagenome-Assembled Genomes and Annotations for McGivern et al

<h3>Files:</h3> <ul> <li><code>reactorEMERGE_annotations.txt</code>: DRAM annotations for MAGs</li> <li><code>gene_lengths.txt</code>: gene length file used to calculate geTMM</li> <li><code>genes.gff.tar.gz</code>: gff file needed for metaT processing</li> <li><code>genes.faa.tar.gz</code>: amino acid sequences for MAG genes</li> <li><code>genes.fna.tar.gz</code>: nucleotide sequences for MAG genes, used as database for metaT mapping</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Metagenome assembled genome (MAG) annotations for Columbia River sediment bacteria and archaea

<p>Excel spreadsheet containing all annotations for metagenome assembled genomes (MAGs) that form part of a publication to be submitted titled:&nbsp;&quot;<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments&quot;.&nbsp;</strong></p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

MGBC-26640: high quality, non-redundant genome annotations (part 2)

<p>Genome annotations (GenBank flat files, .gbk) for the 26,640&nbsp;non-redundant, high-quality genomes of the MGBC (part 2:&nbsp;MGBC130000&nbsp;to&nbsp;MGBC167528).</p> <p>Please note that due to the 50Gb file size limit,&nbsp;part 1&nbsp;of the MGBC annotations can be found at DOI:&nbsp;10.5281/zenodo.5534741.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

MGBC-26640: high quality, non-redundant genome annotations (part 1)

<p>Genome annotations (GenBank flat files, .gbk) for the 276 genomes of the MCC and the 26,640&nbsp;non-redundant, high-quality genomes of the MGBC (part 1:&nbsp;MGBC000001 to&nbsp;MGBC129999).</p> <p>Please note that due to the 50Gb file size limit,&nbsp;part 2 of the MGBC annotations can be found at DOI:&nbsp;10.5281/zenodo.5532847.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Nontuberculous mycobacteria persistence in a cell model mimicking alveolar macrophages (genome assembly and annotation dataset)

<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following Nontuberculous mycobacteria (NTM) strains: <em>Mycobacterium smegmatis </em>mc<sup>2</sup>155 (reference strain), <em>Mycobacterium avium</em> ATCC25921 (reference strain),&nbsp;<em>M. avium </em>60/08 (clinical strain),&nbsp;<em>Mycobacterium fortuitum</em> ATCC6841&nbsp;(reference strain) and&nbsp;<em>M. fortuitum</em> 747/08 (clinical strain).</p> <p>All raw sequence reads used in this&nbsp;study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB30455).</p> <p>The associated article&nbsp;can be found here:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/31035520">https://www.ncbi.nlm.nih.gov/pubmed/31035520</a></p> <p>&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Annotated cable bacteria genomes

<p>Genomes of Candidatus Electrothrix and Candidatus Electronema annotated using Prokka version 1.14.5. Tables of Average Amino Acid Identity and Average Nucleotide Identity for these genomes included.</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Assembly and annotation of eleven Salix (shrub willow) genomes

<p>The shrub willows (<em>Salix</em> section <em>Vetrix</em>) are an emerging bioenergy crop in North America and Eurasia. However, genomics resources in this section are still quite limited, with only a few reference genomes available, despite many species in use in breeding programs. Here we present de novo assemblies and annotations of eleven shrub willow genomes from six species. Copy number variation of candidate sex determination genes within each genome was characterized and revealed remarkable differences in putative master regulator gene duplication and deletion. We also analyzed copy number and expression of candidate genes involved in floral secondary metabolism and identified substantial variation across genotypes, which can be used for parental selection in breeding programs. Lastly, we report on a genotype that produces only female descendants and identified gene presence/absence variation in the mitochondrial genome that may be responsible for this unusual inheritance.</p>

opencc-zeroDec 2022View details →
dryad36/100

Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data

<p class="MsoNormal">Whole genome sequencing enables us to ask fundamental questions about the genetic basis of adaptation, population structure, and epigenetic mechanisms, but usually requires a suitable reference genome for making sense of the sequence data. While the availability of reference genomes has significantly improvement in both taxonomic coverage and overall quality, this poses a challenge for researchers in determining which reference genome best suits their data. Here we compare the use of two different reference genomes for the three-spined stickleback (<em>Gasterosteus aculeatus</em>), one novel genome from a European individual and the published reference genome of a North American individual. Specifically, we investigate the impact of using a local reference versus one generated from a differentiated population on several commonly used metrics in population genomics. Through mapping genome resequencing data of 60 sticklebacks from across Europe and North America, we confirmed genome quality is an important factor in choosing a reference genome. A local reference genome did offer increased mapping efficiency and genotyping accuracy, likely stemming from the higher similarity in genome sequence and synteny. Despite comparable distributions of the metrics generated across the genome using SNP data (i.e., π, Tajima's D, and FST), window-based statistics using different references resulted in different outlier genes and enriched gene functions. In contrast, the marker-based analysis utilising DNA methylation distributions had a considerably higher overlap in outlier genes and functions when using different reference genomes. Overall, our results highlight how using a local reference genome can increase the resolution of genome scans when multiple similar-quality reference genomes are available. Such results have implications in the detection of signatures of selection.</p>

opencc-zeroJan 2023View details →
zenodo36/100

Sol Genome Annotations and Quantitative Traits Loci

<p>RDF data graphs</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

Costus pulverulentus genome annotations

<p><span>The spiral gingers (<em>Costus</em> L.) are a pantropical genus of herbaceous perennial monocots; the Neotropical clade of <em>Costus</em> radiated rapidly in the past few million years into over 60 species. The Neotropical spiral gingers have a rich history of evolutionary and ecological research that can motivate and inform modern genetic investigations. Here, we present the first two chromosome-level genome assemblies in the genus, for <em>C. pulverulentus</em> and <em>C. lasius</em>, and briefly compare their synteny. We assembled the <em>C. pulverulentus</em> genome from a combination of short-read data, Chicago and Dovetail Hi-C chromatin-proximity sequencing, and alignment with a linkage map. We assembled the <em>C. lasius</em> genome with Pacific Biosciences HiFi long reads and alignment to the <em>C</em>. <em>pulverulentus</em> genome. These two assemblies are the first published genomes for non-cultivated tropical plants. These genomes solidify the spiral gingers as a model system and will facilitate research on the poorly understood genetic basis of tropical plant diversification. </span></p>

opencc-zeroMar 2023View details →
zenodo36/100

Macroalgal deep genomics illuminate multiple paths to aquatic, photosynthetic multicellularity - Supplementary Data - ANNOTATIONS - Data S2

<p>Macroalgae are a polyphyletic group of multicellular aquatic organisms vital to global climate maintenance and have a wide variety of commercial applications. The lack of genomic datasets and poor physiological records preclude understanding their ecological roles and industrial potential. We <em>de novo</em> sequenced 121 macroalgal genomes from various climates spanning five major latitude parallels. The resultant genomic datasets illuminate the evolutionary mechanisms behind macroalgal diversification and specialization and reveal genetic bases for niche habitation facilitated by morphological complexity. Adhesome genes (e.g., cadherins, integrins, and lectins), extracellular matrix enzymes, and cytoskeletal organization regulating genes (e.g., spondins, Rho-type GTPases) predominantly distinguished macroalgal genomes from their microalgae correlates. Artificial neural networks could accurately classify an alga as micro- or macro- from set of significance-ranked genomic features (n = 251, entropy R<sup>2</sup> &gt; 0.99) as well as adhesome gene sets (n = 110, entropy R<sup>2</sup> &gt; 0.92). By deciphering the macroalgal adhesome, a clear picture of the genetic basis for the development and maintenance of complex algal tissues could be resolved in the three macroalgal phyla. Sequences from giant viruses were rampant in the macroalgal genomes and coded for zinc-finger transcription factors, ankyrins, Rieske proteins, and other exotic codomains. Lineage-specific retentions of transcription factors, cadherins, integrins, polysaccharide-acting enzymes, and receptor kinases, many with predicted viral origins, outline the divergent mechanisms facilitating multicellularity in these three macroalgal lineages. This work sheds new light on the evolution of multicellularity in three phyla (Rhodophyceae, Chlorophyceae, and Ochrophyceae v. Phaeophyceae) through the lens of large-scale genomics and paves the way for the genomic exploration of macroalgal biology.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Genome annotation file containing predicted genome features of Phytophthora agathidicida (Strain: 3770, Assembly:ASM2572299v1)

<p>This is the genome annotation file (gff3) containing predicted genome features of the&nbsp;<em>Phytophthora agathidicida </em>(Strain: 3770) genome published in Cox et al (2022). This annotation file is associated with the following entries at Genbank:</p> <p>Assembly:&nbsp; &nbsp; ASM2572299v1<br> Biosample:&nbsp; &nbsp;SAMN19597867<br> BioProject:&nbsp;&nbsp; PRJNA734652</p> <p>Included in the file are predicted functional annotations from Blastp search of all predicted proteins sequences against the Swiss-Prot sequence&nbsp;database&nbsp;(Release 23/02).</p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record