Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo36/100

179 high quality metagenome-assembled genomes sequences and annotations

<p>We analyzed seven sediment samples collected adjacent to ferromanganese nodules from the Clarion&ndash;Clipperton&nbsp;Fracture Zone&nbsp;(CCFZ) in the eastern Pacific Ocean.&nbsp;Through deep metagenomic sequencing, assembly, and binning, we reconstructed 179 high quality metagenome-assembled genomes (MAGs).&nbsp;This archive contains these&nbsp;genomes sequences and annotations.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Genome annotation of Hibiscus trionum

<p><strong>Genome annotation of <em>Hibiscus trionum</em></strong></p> <p>Author information<br> Shizuka Koshimizu<sup>1,2</sup>, Sachiko Masuda<sup>3</sup>, Arisa Shibata<sup>3</sup>, Takayoshi Ishii<sup>4</sup>, Ken Shirasu<sup>3</sup>, Atsushi Hoshino<sup>5,6</sup>, Masanori Arita<sup>1,2</sup></p> <p>1 Bioinformation and DDBJ Center, National Institute of Genetics, Mishima 411-8540, Japan<br> 2 Graduate Institute for Advanced Studies, SOKENDAI, Mishima 411-8540, Japan<br> 3 Center for Sustainable Resource Science, RIKEN, Yokohama 230-0045, Japan<br> 4 Arid Land Research Center, Tottori University, Tottori, 680-001, Japan<br> 5 National Institute for Basic Biology, Okazaki 444-8585, Japan<br> 6 Graduate Institute for Advanced Studies, SOKENDAI, Okazaki 444-8585, Japan</p> <p><br> Contents<br> <strong>htrionum_v1.0_genome.fasta</strong>: Genome sequence<br> <strong>htrionum_v1.0_transcripts.gff3</strong>: Gene annotation<br> <strong>htrionum_v1.0_transcripts.gtf</strong>: Gene annotation<br> <strong>htrionum_v1.0_transcripts.fna</strong>: Transcript sequences<br> <strong>htrionum_v1.0_pep.faa</strong>: Protein sequences<br> <strong>Description.txt</strong>: Functional descriptions<br> <strong>InterProScan.txt</strong>: The results of InterProScan for all proteins</p> <p>&nbsp;</p> <p>Reference</p> <p>Genome and transcriptome analyses reveal genes involved in the formation of fine ridges on petal epidermal cells in&nbsp;<em>Hibiscus trionum</em></p> <p><em>DNA Research</em>, 2023;, dsad019,&nbsp;<a href="https://doi.org/10.1093/dnares/dsad019">https://doi.org/10.1093/dnares/dsad019</a></p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Serratia fonticola EBS19 Whole genome sequence data fasta file annotated

<p>Whole genome sequence data of <em>Serratia fonticola</em>&nbsp;<strong>EBS19 </strong>strain.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Danaus plexippus genome annotation and orthofinder results

<p>Funannotate v1.5.3 docker image was used to train Augustus v3.2.3, predict gene models, and perform functional annotation. As input for optimizing the performance of Augustus v3.2.3, funannotate used 2,404 PASA v2.3.3 gene models. To obtain this training gene model set, transcripts were de novo assembled with Trinity v2018-2.8.3 under settings --SS_lib_type RF, using all poly(A) RNA-seq paired reads after adapter removal with Trimmomatic v0.32. These transcripts were aligned to the genome under PASA using BLAT v36, obtaining a first set of gene models. The 500 longest non-redundant ORFs associated with the PASA gene models were used to train TransDecoder v5.2.0. Then the gene models were selected according to their abundance as estimated by Kallisto v0.44.0 under settings --rf-stranded using the Trinity normalized reads. Ultimately, BRAKER v2.0.3b trained Augustus with the retained gene models.<br>For gene prediction, funannotate aligned mRNAs and proteins from the previous annotation (official gene set 2, OGS2) with minimap v2.14-r883 under settings -ax splice --cs -u b -G 3000, and Diamond blastx v0.8.22, respectively. Protein alignments were further refined by funannotate, including 3 kb upstream and downstream of the region of alignment, and subsequently executing Exonerate v2.4.0. Additionally, funannotate parsed the introns supported by alignments of poly(A) RNA-seq reads generated with HISAT v2.1 under settings --rna-strandness RF --max-intronlen 10,000. This combination of hints (protein alignments, transcript alignments, and intron locations) was used by Augustus to predict a second set of 16,756 gene models. Of them, 9,695 were dubbed as highly supported, i.e. had more than 90% of their model supported either by intron hints, transcript alignments, or protein alignments. GeneMark-ET v4.35, under settings --max_intron 3,000 --soft_mask 2,000, was also run independently to predict a third set of gene models but only relying on intron hints.<br>The PASA, Augustus highly supported, Augustus not highly supported, and GeneMark prediction sets were combined by EVidenceModeler, assigning them 10, 5, 1, and 1 relative weights, respectively. The predictions were further filtered by removing genes shorter than 50 aa in length, or that had high sequence similarity (diamond blastp --sensitive --evalue 1e-10) to the repeat database included in funannotate, or that had more than 90% of the model intersecting regions masked by RepeatMasker. The filtered set of gene models was updated in order to include UTR information by two executions of the PASA annotation comparison using the Trinity transcripts and filtering gene models according to transcripts per million as calculated by Kallisto. Alternative transcripts were only kept if they were at least 10% as highly expressed as the most highly expressed transcript per gene.<br>Non-coding genes were annotated with the following tools: tRNA genes, tRNAscan-SE v.2.0; rRNA genes, RNAmmer v.1.2; and for a variety of other RNA genes, Infernal v1.1.1. Specifically, for miRNA-encoding genes, we used BLASTn to locate the most recent annotation of these genes. Lastly, FEELnc classified lncRNAs from the transcripts assembled by StringTie v1.3.2d, and considering protein-coding predictions described above.</p><p>Homology relationships of the predicted protein coding genes were inferred with OrthoFinder (v2.2.6) relative to <i>Bombyx mori, Plutella xylostella, Papilio xuthus, Amyelois transitella, Heliconius melpomene, Plodia interpunctella </i>and<i> Drosophila melanogaster</i>. OrthoFinder was run with the settings "-S diamond -M msa".</p>

opencc-by-4.0May 2021View details →
dryad36/100

A high-quality genome assembly and annotation of the gray mangrove, Avicennia marina

Open the record for dataset details and reuse information.

publicFeb 2022View details →
dryad36/100

Genome assembly and annotation of the Dark-branded Bushbrown butterfly Mycalesis mineus (Nymphalidae: Satyrinae)

Open the record for dataset details and reuse information.

publicFeb 2024View details →
dryad36/100

Assembly and annotation of eleven Salix (shrub willow) genomes

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

Re-annotated genomes for plastomes and mitogenomes

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad36/100

A high-quality genome assembly and annotation of the dark-eyed junco Junco hyemalis, a recently diversified songbird

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Costus pulverulentus genome annotations

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Annotation of genome of brown-marbled grouper (Epinephelus fuscoguttatus)

Open the record for dataset details and reuse information.

publicMay 2021View details →
dryad36/100

A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad36/100

Data from: An annotated draft genome of the mountain hare (Lepus timidus)

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad36/100

Reference genome and annotation for Teleopsis dalmanni

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad36/100

Chromosome-level assembly of two pearl millet (Cenchrus americanus) genomes, functional annotation and transcriptomes

Open the record for dataset details and reuse information.

publicFeb 2024View details →
dryad36/100

Annotated genome assemblies for Geoscapheus dilatatus, Panesthia cribrata and Neogeoscapheus hanni

Open the record for dataset details and reuse information.

publicFeb 2024View details →
dryad36/100

Deep-learning-based annotation of 230 superasterid genomes reveals a harmonized dataset of 91,366 NLRs

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Genome sequence assembly and annotation of MATA and MATB strains of <em>Yarrowia lipolytica</em>

Open the record for dataset details and reuse information.

publicOct 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record