Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Nicotiana tomentosiformis genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tomentosiformis</em>.</p> <p>The following files are available:</p> <ul> <li>ntom.fa.gz: reference genome sequence in fasta format</li> <li>ntom.gff3.gz: gene annotation in GFF3 format</li> <li>ntom.gtf.gz: gene annotation in GTF format</li> <li>ntom.tx.fa.gz: transcript sequences in fasta format</li> <li>ntom.cds.fa.gz: coding sequences in fasta format</li> <li>ntom.prot.fa.gz: protein sequences in fasta format</li> <li>ntom.tsv.gz: gene functional annotation in TSV format</li> <li>ntom.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntom.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntom.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntom.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium malaccense genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Syzygium malaccense</em>.</p> <p>The following files are available:</p> <ul> <li>smal.fa.gz: reference genome sequence in fasta format</li> <li>smal.gff3.gz: gene annotation in GFF3 format</li> <li>smal.gtf.gz: gene annotation in GTF format</li> <li>smal.tx.fa.gz: transcript sequences in fasta format</li> <li>smal.cds.fa.gz: coding sequences in fasta format</li> <li>smal.prot.fa.gz: protein sequences in fasta format</li> <li>smal.tsv.gz: gene functional annotation in TSV format</li> <li>smal.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium aqueum genome assembly and annotation

<p><em>De novo </em>genome assembly and annotation of <em>Syzygium aqueum</em>.</p> <p>The following files are available:</p> <ul> <li>saqu.fa.gz: reference genome sequence in fasta format</li> <li>saqu.gff3.gz: gene annotation in GFF3 format</li> <li>saqu.gtf.gz: gene annotation in GTF format</li> <li>saqu.tx.fa.gz: transcript sequences in fasta format</li> <li>saqu.cds.fa.gz: coding sequences in fasta format</li> <li>saqu.prot.fa.gz: protein sequences in fasta format</li> <li>saqu.tsv.gz: gene functional annotation in TSV format</li> <li>saqu.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Nicotiana tabacum genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tabacum</em>.</p> <p>The following files are available:</p> <ul> <li>ntab.fa.gz: reference genome sequence in fasta format</li> <li>ntab.gff3.gz: gene annotation in GFF3 format</li> <li>ntab.gtf.gz: gene annotation in GTF format</li> <li>ntab.tx.fa.gz: transcript sequences in fasta format</li> <li>ntab.cds.fa.gz: coding sequences in fasta format</li> <li>ntab.prot.fa.gz: protein sequences in fasta format</li> <li>ntab.tsv.gz: gene functional annotation in TSV format</li> <li>ntab.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntab.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntab.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntab.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium syzygioides genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Syzygium syzygioides</em>.</p> <p>The following files are available:</p> <ul> <li>ssyz.fa.gz: reference genome sequence in fasta format</li> <li>ssyz.gff3.gz: gene annotation in GFF3 format</li> <li>ssyz.gtf.gz: gene annotation in GTF format</li> <li>ssyz.tx.fa.gz: transcript sequences in fasta format</li> <li>ssyz.cds.fa.gz: coding sequences in fasta format</li> <li>ssyz.prot.fa.gz: protein sequences in fasta format</li> <li>ssyz.tsv.gz: gene functional annotation in TSV format</li> <li>ssyz.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Nicotiana sylvestris genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana sylvestris</em>.</p> <p>The following files are available:</p> <ul> <li>nsyl.fa.gz: reference genome sequence in fasta format</li> <li>nsyl.gff3.gz: gene annotation in GFF3 format</li> <li>nsyl.gtf.gz: gene annotation in GTF format</li> <li>nsyl.tx.fa.gz: transcript sequences in fasta format</li> <li>nsyl.cds.fa.gz: coding sequences in fasta format</li> <li>nsyl.prot.fa.gz: protein sequences in fasta format</li> <li>nsyl.tsv.gz: gene functional annotation in TSV format</li> <li>nsyl.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>nsyl.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>nsyl.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>nsyl.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Clove (Syzygium aromaticum) genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of clove (<em>Syzygium aromaticum</em>).</p> <p>The following files are available:</p> <ul> <li>saro.fa.gz: reference genome sequence in fasta format</li> <li>saro.gff3.gz: gene annotation in GFF3 format</li> <li>saro.gtf.gz: gene annotation in GTF format</li> <li>saro.tx.fa.gz: transcript sequences in fasta format</li> <li>saro.cds.fa.gz: coding sequences in fasta format</li> <li>saro.prot.fa.gz: protein sequences in fasta format</li> <li>saro.tsv.gz: gene functional annotation in TSV format</li> <li>saro.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Improved genome annotation of Rhynchosporium commune isolate UK7 using Illumina short reads of in vitro and in plantae conditions

<p>Improved genome annotation of the <em>Rhynchosporium commune</em> isolate UK7 using Illumina short reads of <em>in vitro</em> and <em>in plantae</em> conditions. The short reads used for the annotation are available at <a href="https://doi.org/10.5281/zenodo.5729968">https://doi.org/10.5281/zenodo.5729968</a> and <a href="https://doi.org/10.5281/zenodo.5729863">https://doi.org/10.5281/zenodo.5729863</a>.To create the gene models, we used tophat v. 2.0.14 to align short reads to the UK7 reference genome (Trapnell et al., 2009). The Intron splice site hints were generated using bam2hints, included in the AUGUSTUS v. 3.2.1 software (Stanke et al., 2006). Due to the very high RNA-sequencing depth available, intron splice hints were filtered for a minimum coverage of 20 reads to avoid an impact of spurious splice signals on gene prediction. To produce <em>ab initio </em>gene models, the BRAKER v. 1.0 pipeline (Hoff et al., 2016) combining GeneMark-ET <em>ab initio </em>gene model predictions and AUGUSTUS v. 3.2.1. GeneMark-ET was trained using the RNA-seq-based splice information as hints. AUGUSTUS was automatically trained using <em>ab initio </em>gene models that were fully supported by splice information. Finally, AUGUSTUS was used to predict gene models using both RNA-seq splice information and coding sequence hints based on exonerate protein alignments as extrinsic evidence.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Genome assembly and annotation of Pisum sativum cultivar ZW6 (PeaZW6)

<p>This reposity stores the genome assembly and gene annotation of Pisum sativm cultivar ZW6 (PeaZW6)</p> <p>Current Version : Release Candidate Version 2 (RC2)</p> <p>Associated NCBI BioProject :&nbsp;<strong>PRJNA730094</strong></p> <p>Correspondance&nbsp;: gaoshh@im.ac.cn</p> <p>&nbsp;</p> <p>pea.assembly.ZW6.RC2.fasta.gz&nbsp;- Full&nbsp;genome sequences</p> <p>pea.assembly.ZW6.RC2.chr.fasta.gz - Genome sequences with only chromosome molecules&nbsp;</p> <p>pea.assembly.ZW6.RC2.annotated.gff3 / gtf / bed - Gene annotation in GFF3 / GTF / BED formats</p> <p>pea.assembly.ZW6.RC2.annotated.cds.fasta - Gene coding sequences</p> <p>pea.assembly.ZW6.RC2.annotated.proteins.fasta - Gene protein sequences</p> <p>pea.assembly.ZW6.RC2.annotated.annotations.txt - Additional annotation information in tabular text format (TSV)</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

<p>The European starling, <em>Sturnus vulgaris</em>, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (<em>S. vulgaris</em> vAU), and a second North American genome (<em>S. vulgaris</em> vNA), as complementary reference genomes for population genetic and evolutionary characterisation. <em>S. vulgaris</em> vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (<em>Taeniopygia guttata</em>) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using <em>S. vulgaris</em> vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, <em>S. vulgaris</em> vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.</p>

opencc-zeroJul 2022View details →
zenodo40/100

Training data for 'Refining Manual Genome Annotations with Apollo (eukaryotes)' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial for manual curation of eukaryotic genome annotation using Apollo.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)

<p>Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Sei whole-genome sequence class annotations

<p>Sei sequence class whole-genome annotations are available in the following files:</p> <ul> <li> <p>sorted.hg38.tiling.bed.ipca_randomized_300.labels.merged.bed - The sorted, merged sequence class assignments from Louvain community clustering of the 30 million sequences, uniformly tiling the whole human genome. The fourth column is the sequence class number, with any sequence classes numbering 40-61 excluded from our analyses in the publication. Sequence classes 0-39 can be mapped to the following labels:&nbsp;<a href="https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names">https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names</a></p> </li> <li> <p>sorted.hg19.tiling.bed.ipca_randomized_300.labels.merged.bed - lifted over version of the hg38 BED file.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Genome annotations of Drosophila melanogaster and Drosophila simulans wild-type strains from long read sequencing assemblies

<p>Genome assemblies were performed for eight wild-type strains of Drosophila melanogaster and Drosophila simulans from Oxford Nanopore long read sequencing (please refer to Mohamed et al. Cells 2020 (doi:10.3390/cells9081776)). Assemblies were deposited in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB50024 (<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEBxxxx">https://www.ebi.ac.uk/ena/browser/view/</a>PRJEB50024).</p> <p>Transposable Element annotations: we used RepeatMasker 4.1.0 (<a href="http://repeatmasker.org/">http://repeatmasker.org/</a>) -species Drosophila, followed by OneCodeToFindThemAll (Bailly-Bechet et al. 2014) with default parameters.</p> <p>Gene annotations: We retrieved gtf files from FlyBase : <a>ftp.flybase.net/genomes/D</a><a>rosophila_melanogaster/dmel_r6,46_FB2022_03/gft/dmel-all-r6.46.gtf.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/gtf/dsim-all-</a><a>r2,02.gtf.gz</a>. The corresponding fasta files were also downloaded from FlyBase: <a>ftp.flybase.net/genomes/Drosophila_melanogaster/dmel_r6,46_FB2022_03/</a><a>fasta</a><a>/dmel-all-</a><a>chromosome-</a><a>r6.46.</a><a>fasta</a><a>.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/</a><a>fasta</a><a>/dsim-all-</a><a>chromosome-</a><a>r2,02.</a><a>fasta</a><a>.gz</a>. We used Liftoff (Shumate and Salzberg, 2020) to lift over gene annotations from the references to our genome assemblies. We used -flank 0.2 and only kept the &ldquo;gene&rdquo; and &ldquo;exon&rdquo; terms.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Genomic insights into local adaptation and future climate-induced vulnerability of a keystone forest tree in East Asia (The genome assembly and annotion fiile)

<p>The&nbsp;genome assembly and annotion fiile used in the manuscript:&nbsp;<strong>Genomic insights into local adaptation and future climate-induced vulnerability of a keystone forest tree in East Asia</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Mitochondrial genomes annotations for three previously submitted Rhizostomeae (OZ032132, OZ025205, OZ025288)

<p>This is a repository for the three Rhizostomeae mitochondrial genomes (OZ032132, OZ025205, OZ025288) that were annotated for analysis in the paper 'Complete linear mitochondrial genomes for Cephea cephea and Mastigias albipunctata (Scyphozoa: Rhizostomeae), with an analysis of phylogenetic relationships' by Tan KC, Collins AG and Ames CL.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Annotated genome files of Bacillus subtilis AV1

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
dryad40/100

Data from: Regulatory genome annotation for 33 insect species

<p>Annotation of newly-sequenced genomes frequently includes genes, but rarely covers important non-coding genomic features such as the cis -regulatory modules—e.g., enhancers and silencers—that regulate gene expression. Here, we begin to remedy this situation by developing a workflow for rapid initial annotation of insect regulatory sequences, and provide a searchable database resource with enhancer predictions for 33 genomes. Using our previously-developed SCRMshaw computational enhancer prediction method, we predict over 2.8 million regulatory sequences along with the tissues where they are expected to be active, in a set of insect species ranging over 360 million years of evolution. Extensive analysis and validation of the data provides several lines of evidence suggesting that we achieve a high true-positive rate for enhancer prediction. One, we show that our predictions target specific loci, rather than random genomic locations. Two, we predict enhancers in orthologous loci across a diverged set of species to a significantly higher degree than random expectation would allow. Three, we demonstrate that our predictions are highly enriched for regions of accessible chromatin. Four, we achieve a validation rate in excess of 70% using in vivo reporter gene assays. As we continue to annotate both new tissues and new species, our regulatory annotation resource will provide a rich source of data for the research community and will have utility for both small-scale (single gene, single species) and large-scale (many genes, many species) studies of gene regulation. In particular, the ability to search for functionally-related regulatory elements in orthologous loci should greatly facilitate studies of enhancer evolution even among distantly related species.</p>

opencc-zeroJul 2024View details →
zenodo40/100

Supplemental Results for Assembly, Annotation, and Analysis from HiFi reads of Gulf Toadfish Genome and Transcriptome fOpsBet2.1

<p>This repository contains gzipped tarballs of the results of the various assembly, annotation, and analysis steps performed during the assembly of the fOpsBet2.1 genome assembly for Opsanus beta at the University of Miami Rosenstiel School of Marine, Atmospheric, and Earth Science for the McDonald Toadfish Lab. These results are too numerous to include as supplemental data for a journal publication and so are available here for review. In this repository you will find results for:</p><p>Scripts:</p><p>-all bash and LSF scheduler job scripts used as part of the analysis, both exploratory and final.&nbsp;</p><p>QC:</p><p>-GenomeScope2 estimation of genome metrics from HiFi Reads</p><p>-QUAST genome statistics for each assembly step</p><p>-BUSCO completeness assessments for each assembly step&nbsp;</p><p>-inspector logs for polishing of initial assembly</p><p>-logs from Kraken2 contaminant screen</p><p>Assembly and Scaffolding:</p><p>-ntLINKS logs and intermediates for initial scaffolding</p><p>-ragtag logs and metrics for super-scaffolding to the ThaAma1.1 T. amazonica reference assembly</p><p>-mitoHIFI results for mitogenome assembly from HiFi reads, primary assembly, and purged alternate assembly</p><p>Annotation:</p><p>-PASA directory with full input and output for SQLite PASA assembly of transcriptome for gene predictors</p><p>-Results folder for Funannotate::annotate for gene models, annotations, and CDS/mRNA/protein fastas</p><p>-InterProSCan5 results for protein annotation used as input into Funannotate</p><p>-ghostKOALA KEGG assignment results for predicted proteins from funannotate results</p><p>Repetitive Elements:</p><p>-tidk telomere repeat analysis results</p><p>-TRAH satellite DNA analysis with subsequent analysis with HiCAT and StainedGlass</p><p>-RepeatModeler results for de novo TE prediction</p><p>-repclassifier results for TE curation</p><p>Comparative Analysis:</p><p>-OrthoFinder ortholog search for O. beta to several other vertebrates</p><p>-CAFE5 gene family expansion and contraction of Orthogroups from OrthoFinder results</p><p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Chardonnay genome assembly and annotations

<p>Annotations relating to the Chardonnay genome assembly (the study is published here: https://doi.org/10.1371/journal.pgen.1007807). The assembly is also available at NCBI: BioProject: PRJNA399599.</p> <p><strong>chardonnay_p-ctg.fasta / chardonnay_h-ctg.fasta:</strong></p> <p>Contig assembly for Chardonnay. chardonnay_p-ctg.fasta = primary contigs (haploid representation), chardonnay_h-ctg.fasta = haplotigs (alt contigs for phased regions in haploid assembly).</p> <p><strong>p-h.maker.gff / p-h.maker.proteins.faa / p-h.maker.transcripts.fna:</strong></p> <p>Maker-predicted gene annotations.</p> <p><strong>p-h.maker.draft-names.tsv / p-h.orthomcl.orthoGroups.tsv / p-h.KEGG.tsv:</strong></p> <p>Draft names (based on UniprotKB blastP hits), OrthoMCL annotations, and KEGG annotations for maker-predicted genes.</p> <p><strong>p-h.repeats.gff:</strong></p> <p>RepeatMasker-based repeat annotations.</p> <p><strong>chardonnay_primary_contigs_chromosome-order.fa / chardonnay_haplotigs_chromosome-order.fa:</strong></p> <p>Contigs placed in chromosome-order (using PN40024 as reference).</p> <p><strong>chardonnay_primary_contig_mappings.tsv / chardonnay_haplotig_mappings.tsv:</strong></p> <p>Mapping coordinates for chromosome-ordered contigs.</p> <p><strong>kmer-based_parentage.primary_contigs.bed / kmer-based_parentage.haplotigs.bed:</strong></p> <p>Parentage assignments using the kmer-based method described in the study.</p> <p><strong>SNP-based_parentage.primary_contigs.bed / SNP-based_parentage.haplotigs.bed:</strong></p> <p>Parentage assignments using SNP-based method (view haplotig assignments against primary contigs) described in the study.</p> <p><strong>p-ctg.gene-expansion-candidates.tsv / h-ctg.gene-expansion-candidates.tsv / p-ctg.gene-expansion-candidates.bed / h-ctg.gene-expansion-candidates.bed:</strong></p> <p>Gene expansion candidates (TSV = 1 row per predicted orthogroup, BED = annotations for viewing).</p> <p><strong>PN_and_CH.FAR2.msa.png: </strong></p> <p>Multi Sequence Alignment for expansion of FAR2-like genes in described in Chardonnay genome assembly publication described in the study.</p>

opencc-by-4.0Nov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record