Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Annotated genome of Pseudomonas sp. MWU12-2345

<p>Annotated genome of Pseudomonas sp. MWU12-2345, isolated from peat and sandy bog soils in the Cape Cod National Seashore, Massachusetts.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Annotated genome of Pseudomonas sp. MWU12-2037

<p>Annotated genome of Pseudomonas sp. MWU12-2037, isolated from peat and sandy bog soils in the Cape Cod National Seashore, Massachusetts.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Genome and gene annotation for yeast strain SK1 used in "Deciphering the "m6A Code" via Antibody-Independent Quantitative Profiling"

<p>Genome and gene annotation used in &quot;Deciphering the &ldquo;m6A Code&rdquo; via Antibody-Independent Quantitative Profiling&quot; provided by Schraga Schwartz from his time at the Broad Institute.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Annotated Genome Pseudomonas Species MWU 12.2311

<p>Annotated genome Pseudomonas species MWU 12.2311 isolated from cultivated bogs in southeastern Massachusetts.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Enterobacter cloacae complex (E. bugandensis species, ST599) strain associated with a catheter-related bloodstream infection (CRBSI) (genome assembly and annotation dataset)

<p>This dataset includes the&nbsp;assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a&nbsp;<em><strong>Enterobacter&nbsp;cloacae </strong></em><strong>complex</strong><em><strong>&nbsp;(E. bugandensis </strong></em><strong>species</strong><em><strong>, </strong></em><strong>ST599</strong><em><strong>) </strong></em>strain associated with a catheter-related bloodstream infection&nbsp;(CRBSI) (genome anotation was performed using&nbsp;Prokka v1.14.5; https://github.com/tseemann/prokka)</p> <p>The raw sequence reads were&nbsp;deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB45360; Run&nbsp;Accession:&nbsp;ERR10044433).</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Genome and annotation files (v4_23) for Blumeria graminis f. sp. tritici isolate CHE_96224 (genome assembly: Bgt_genome_v3_16)

<p>Genome and annotation (v4_23) files for Blumeria graminis f. sp. tritici isolate CHE_96224 (genome assembly: Bgt_genome_v3_16)</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data

<p>This upload include data and scripts supporting&nbsp;the results described in the manuscript of&nbsp;<em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>.&nbsp;</em>Detailed description of the contents can be found in README.txt.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

RirC3: Rhizophagus irregularis reference genome and annotation

<p>Reference genome assembly of&nbsp;<em>Rhizophagus irregularis</em>&nbsp;isolate C3, made with PacBio Sequel II SMRT sequencing, and polished with Illumina reads. Annotation was produced with FunAnnotate.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Predicted genes annotated with Prokka for the 623 archaeal UBA genomes

<p>Gene prediction and annotation for the 623 archaeal UBA metagenome-assembled genomes using Prokka v1.12 with Pfam v31 and UniProt databases created on April 17, 2017 according to the Prokka instructions.</p> <p>These genomes are described in:</p> <p>Parks DH, et al. 2017. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol, doi:10.1038/s41564-017-0012-7/ .</p> <p>https://www.nature.com/articles/s41564-017-0012-7</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Revised transcript annotations for GRCh38 reference genome and Ensembl v87.

<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh38<br> Ensembl version: 87</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Revised transcript annotations for GRCh37 (hg19) reference genome and Ensembl v90.

<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh37<br> Ensembl version: 90</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Haematococcus lacustris genome assembly and functional annotation

<p>Nuclear, chloroplast and mitochondrial genome assembly and functional annotation of Haematococcus lacustris (formely Haematococcus pluvialis) K-0084</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Annotated black yeast genomes

<p>The funannotate &ldquo;annotate&rdquo; function was used to annotate genes using Eggnog, Pfam, Interproscan Gene Ontology (GO) terms, MEROPS, and Secreted proteins . FungiSMASH, the fungal-genome-specific version of antiSMASH6 was used to predict the biosynthetic gene clusters (BGCs).&nbsp;</p> <p>&nbsp;</p> <p>Genome annotade</p> <p><em>R. similis</em> LABIOMMI 1217<em>R. simili</em>s Poitiers<br><em>E. xenobiotica</em> CBS118157<br><em>E. spinifera </em>BMU00051<br><em>E. sideris</em> CBS121828<br><em>E. oligosperma</em> CBS72588</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
dryad36/100

Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024

<p>Expression quantitative trait loci (eQTLs) provide a key bridge between noncoding DNA sequence variants and organismal traits. The effects of eQTLs can differ among tissues, cell types, and cellular states, but these differences are obscured by gene expression measurements in bulk populations. We developed a one-pot approach to map eQTLs in <em>Saccharomyces cerevisiae</em> by single-cell RNA sequencing (scRNA-seq) and applied it to over 100,000 single cells from three crosses. We used scRNA-seq data to genotype each cell, measure gene expression, and classify the cells by cell-cycle stage. We mapped thousands of local and distant eQTLs and identified interactions between eQTL effects and cell-cycle stages. We took advantage of single-cell expression information to identify hundreds of genes with allele-specific effects on expression noise. We used cell-cycle stage classification to map 20 loci that influence cell-cycle progression. One of these loci influenced the expression of genes involved in the mating response. We showed that the effects of this locus arise from a common variant (W82R) in the gene <em>GPA1</em>, which encodes a signaling protein that negatively regulates the mating pathway. The 82R allele increases mating efficiency at the cost of slower cell-cycle progression and is associated with a higher rate of outcrossing in nature. Our results provide a more granular picture of the effects of genetic variants on gene expression and downstream traits.</p>

opencc-zeroMay 2024View details →
zenodo36/100

iPS2-sci-seq annotation genome

<p>Decompress SA.tar before using.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Draft genome and annotation of Raphidonema monicae

<p>Microalgae synthesize diverse lipids that perform a wide variety of structural, metabolic, and signalling roles in the cell. However, we know relatively little about lipid metabolism across different protist groups, and especially those adapted to extreme environments. As part of a project to investigate the lipid metabolism of <em>Raphidonema monicae</em> strain SAG 2030, a cold-adapted microalga isolated from Antarctica, a draft genome was assembled and annotated.</p> <p>The genome was assembled from a single DNA library sequenced with Illumina 150 bp paired-end reads and assembled by NovoGene (Hong Kong) with SOAPdenovo2. Novogene also perfomed curation and structural annotation of CDS and peptide sequences. The data were subsequently functionally annotated in our lab using BlastP, InterProScan and emapper.py, and the results were curated with OmicsBox software.</p> <p>The assembly sequence and annotation are provided as data files supporting the lipidomic investigation of the effects of abiotic stress and include the following:</p> <p>File 1. &lsquo;<strong>KAD1.seq</strong>&rsquo; is the assembled genome sequence contigs that have been curated by NovoGene, .fasta format.</p> <p>File 2. &lsquo;<strong>KAD1.gff</strong>&rsquo; is the annotation file .gff format for the gene CDS, .gff format.</p> <p>File 3. &lsquo;<strong>KAD1.cds</strong>&rsquo; contains the CDS nucleotide sequences, .fasta format.</p> <p>File 4. &lsquo;<strong>KAD1.pep</strong>&rsquo; The corresponding peptide sequences, .fasta format.</p> <p>File 5. &lsquo;<strong>kad.annotation.table.xls</strong>&rsquo; is the combined functional annotation of the peptide sequences using BlastP, InterProScan and Emapper.py in an excel-readable table format.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Genome and annotation files for Blumeria graminis f. sp. tritici isolate CHVD042201 (genome assembly: Bgt_CHVD042201 _genome_v1)

<p>Genome and annotation files for Blumeria graminis f. sp. tritici isolate CHVD042201 (genome assembly: Bgt_CHVD042201 _genome_v1)</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Training material for Genome Annotation

<p>Genome annotation is the process of attaching biological information to sequences. It consists of three main steps:</p> <ul> <li>identifying portions of the genome that do not code for proteins</li> <li>identifying elements on the genome, a process called gene prediction, and</li> <li>attaching biological information to these elements.</li> </ul>

opencc-by-4.0May 2018View details →
zenodo36/100

BRIDGES singletons annotated with local genomic features

<p>This dataset consists of a tarball archive containing 8,192 tab-delimited files (one per 7-mer sequence motif). Each file contains information about the status or value of 15 different genomic features at every possible site in hg19, centered at the 7-mer sequence motif indicated in the filename (or the reverse complement of that motif; e.g.&nbsp;<code>ACGATGC_annotated.txt</code>&nbsp;includes information for sites at 5&rsquo;-ACG<strong>A</strong>TGC-3&rsquo;&nbsp;<em>and</em>&nbsp;sites at 5&rsquo;-GCA<strong>T</strong>CGT-3&rsquo; motifs).</p> <p>Each file contains the following columns:</p> <ul> <li> <p><strong>AT_CG</strong>&nbsp;[indicator if site carries an A&gt;C or T&gt;G singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>AT_GC</strong>&nbsp;[indicator if site carries an A&gt;G or T&gt;C singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>AT_TA</strong>&nbsp;[indicator if site carries an A&gt;T or T&gt;A singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>GC_AT</strong>&nbsp;[indicator if site carries a G&gt;A or C&gt;T singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>GC_CG</strong>&nbsp;[indicator if site carries a G&gt;C or C&gt;G singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>GC_TA</strong>&nbsp;[indicator if site carries a G&gt;T or C&gt;A singleton (1) or not (0) in the BRIDGES data]</p> </li> <li> <p><strong>DP</strong>&nbsp;[average depth of coverage at site]</p> </li> <li> <p><strong>H3K4me1</strong>&nbsp;[indicator if site is within a H3K4me1 broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K4me3</strong>&nbsp;[indicator if site is within a H3K4me3 broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K9ac</strong>&nbsp;[indicator if site is within a H3K9ac broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K9me3</strong>&nbsp;[indicator if site is within a H3K9me3 broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K27ac</strong>&nbsp;[indicator if site is within a H3K27ac broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K27me3</strong>&nbsp;[indicator if site is within a H3K27me3 broad peak (1) or not (0)]</p> </li> <li> <p><strong>H3K36me3</strong>&nbsp;[indicator if site is within a H3K36me3 broad peak (1) or not (0)]</p> </li> <li> <p><strong>EXON</strong>&nbsp;[indicator if site is within an exon (1) or not (0)]</p> </li> <li> <p><strong>CpGI</strong>&nbsp;[indicator if site is within a CpG island (1) or not (0)]</p> </li> <li> <p><strong>RR</strong>&nbsp;[average recombination rate in the 10kb window centered at the site]</p> </li> <li> <p><strong>LAMIN</strong>&nbsp;[indicator if site is within an Lamin-Associated Domain (1) or not (0)]</p> </li> <li> <p><strong>DHS</strong>&nbsp;[indicator if site is within a DNase Hypersensitive region (1) or not (0)]</p> </li> <li> <p><strong>TIME</strong>&nbsp;[average recombination rate in the 10kb window centered at the site]</p> </li> <li> <p><strong>GC</strong>&nbsp;[average GC content in the 10kb window centered at the site]</p> </li> </ul> <p>Note that the chromosome and position of each site has been removed to protect sample privacy.</p> <p>Each file is then passed to an R script (available at&nbsp;<a href="https://github.com/carjed/smaug-genetics">https://github.com/carjed/smaug-genetics</a>) to estimate the effects each feature on the relative mutation rate using a logistic regression model (e.g.,&nbsp;<code>AT_GC ~ DP + ... + GC</code>). Each of the features used is available from data in the public domain; the provenance of these features is described in the associated paper, and additional scripts for processing the feature data can be found at at&nbsp;<a href="https://github.com/carjed/smaug-genetics">https://github.com/carjed/smaug-genetics</a>.</p> <p>&nbsp;</p> <p>The BRIDGES whole-genome sequencing study is described at&nbsp;<a href="https://doi.org/10.1101/108290">https://doi.org/10.1101/108290</a></p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

Annotations of sapSuaA1 genome assembly

<p>Annotations accompanying paper &quot;Genome Sequence of Flavor-Producing Yeast Saprochaete suaveolens NRRL Y-17571&quot;</p>

openother-openFeb 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record