Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

133

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

133 results for “gene annotation”

Learn how ShareScore rates datasets ↗
dryad40/100

Data for: Heterosigma akashiwo transcriptome gene annotations

<p>Heterosigma akashiwo is a eukaryotic, cosmopolitan, and unicellular alga (class: Raphidophyceae), and produces fish-killing blooms. There is a substantial scientific and practical interest in its ecophysiological characteristics that determine bloom dynamics and its adaptation to broad climate zones. A well-annotated genomic/genetic sequence information enables researchers to characterize organisms using modern molecular technology. In the present study, we conducted H. akashiwo RNA sequencing, a de novo transcriptome assembly of 84,693,530 high-quality deduplicated short-read sequences. The obtained RNA reads were assembled by Trinity assembler and 144,777 contigs were identified with N50 values of 1085. The raw data were deposited in the NCBI SRA database (BioProject PRJDB6241 and PRJDB15108), and the assemblies are available in NCBI TSA database (ICRV01).  Total 60,877 open reading frames with the length of 150 bp or greater were predicted. Here, the top Gene Ontology terms, the pfam hits, and the BLAST hits were annotated for all the predicted genes, and shared as text files.</p>

opencc-zeroMar 2023View details →
dryad40/100

Annotation of genes encoding enzymes across marine phytoplankton genomes

<p>Phytoplankton cells span a large size range, from picoplankton (&lt;2µm), nanoplankton (2 to 20µm), microplankton (20 to 200µm) to macroplankton (200 to &lt;2000µm). Cell size interacts with multiple selective pressures, including cellular metabolic rate, light absorption, nutrient uptake, cell nutrient quotas, trophic interactions and diffusional exchanges with the environment. Beyond simple size, cells of different shapes differ in surface area to volume ratio. For example, more elongated cells, such as pennate diatoms, have a larger surface area to volume ratio compared to more rounded cells, such as centric diatoms, of equivalent biovolume, which can in turn influence diffusional exchanges between cells and their environment. We assembled metadata on diverse marine phytoplankters, in parallel with genomic or transcriptomic data annotations to identify genes encoding enzymes, to facilitate analyses of genomic patterns of encoded enzymes across diverse taxa, sizes, growth forms and origins of strains.</p>

opencc-zeroApr 2023View details →
zenodo40/100

Genome sequences and gene annotations for two Ophryocystis lineages

<p>Assembly, annotation, and gene sequences for the&nbsp;<em>Ophryocystis&nbsp;</em>lineages sequenced in &quot;Genome sequence of <em>Ophryocystis elektroscirrha</em>, an apicomplexan parasite of monarch butterflies: cryptic diversity and response to host-sequestered plant chemicals.&quot; Each of the two lineages has three associated files:&nbsp;a genome sequence file (.fa), an annotation in .gff3 format, and gene sequences in .fna format. Sequences generated for&nbsp;<em>Ophryocystis elektroscirrha&nbsp;</em>come from direct DNA extraction and sequencing effort and are hosted elsewhere on NCBI as well. The other lineage, prefixed&nbsp;Ophryocystis-elektroscirrha_like, was bioinformatically extracted from the genome of an infected host. As such, we are less confident in its completeness and it is not archived elsewhere.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Genomes of cactophilic Drosophila species and their respective gene and TE annotations + QC data

<p>The genomes deposited in this repository refers to the data used in the study entitled "Transposable elements contribute to the evolution of host shift-related genes in cactophilic<em> Drosophila</em> species", from Oliveira D. S., Larue A., Nunes W. V. B., Sabot F., Bodel&oacute;n A., Garc&iacute;a Guerreiro M. P., Vieira C., Carareto C. M. A.</p> <p>Each genome has its following assembly (fasta), gene annotation (gff), and TE annotation (gtf). The quality control for the nanopore genomes can be accessed on QC_nanopore_genomes.zip, and the quality control for the RNA-seq data on QC_RNAseq.zip.</p> <p>Additional files, as code and input files to reproduce the specific analysis of the manuscript, are also provided in the github repository: https://github.com/OliveiraDS-hub/Pipelines-Cactophilic-Drosophila-Species</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
dryad40/100

Annotation of genes encoding enzymes across marine phytoplankton genomes

Open the record for dataset details and reuse information.

publicApr 2023View details →
dryad40/100

Data for: Heterosigma akashiwo transcriptome gene annotations

Open the record for dataset details and reuse information.

publicMar 2023View details →
zenodo36/100

Annotated genes harboring major effect markers (R2 ≥ 15%). Highlighted in green are genes annotated from Rhodes et al. 2014,2017, in orange genes annotated as similar to Peroxidase, in yellow new annotations from sorghum genome in Atlas. In the first three columns start and stop position on the sorghum genome and transcript name, followed by the nearest marker name and the distance of the gene from the nearest marker, then a column where are shown the GWAS methods and target traits for which the linked SNP was significant, the last column shows the category of the genes.

<p><strong>We conducted a comprehensive genomics study to map genomic loci determining the production of antioxidants in sorghum grains. Encouraging results were obtained and published in peer-reviewed article with impact factor (https://doi.org/10.1371/journal.pone.0225979). Annotated genes harboring major effect markers (R<sup>2</sup> &ge; 15%) were identified and will be of worldwide interest. </strong></p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Whole genome assembly and gene annotation of a diploid genotype of Brachiaria ruziziensis (syn. Urochloa ruziziensis)

<p>In this work, we have presented a comprehensive analysis of the molecular mechanism linked to aluminium tolerance in <em>Brachiaria</em> species. By assembling and annotating a diploid genotype of <em>B. ruziziensis</em> we have developed the capability for genomic-based studies of desirable phenotypic traits. Using this resource, we have identified three QTLs associated to root architecture and vigour during Al<sup>3+</sup> stress in a hybrid population from a high and low tolerant accession. We have also identified a number of genes and molecular responses that impact on different aspects of signalling, cell-wall composition and active transports as a response to aluminium stress. <em>Brachiaria </em>tolerance appears to build in the same genes than in rice. However, we found that external mechanisms such as sequestration of Al<sup>3+</sup> common in other grasses might be not that important in <em>Brachiaria. </em>Also, contrasting regulation in the same genotype after 8 or 72 hours of Al<sup>3+</sup> stress of numerous genes involved in RNA translation can explain the different levels of tolerance among different Brachiaria species. The newly annotated draft genome represents an important base upon which study other aspects of <em>Brachiaria</em> biology.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus

<p>The Pro(tein)/Gene corpus was developed at the JULIE Lab Jena under supervision of Prof. Udo Hahn.</p> <p>The goals of the annotation project were</p> <ul> <li>to construct a consistent and (as far as possible) subdomain-independent/-comprehensive protein-annotated corpus</li> <li>to differentiate between protein families and groups, protein complexes, protein molecules, protein variants (e.g. alleles) and elliptic enumerations of proteins.</li> </ul> <p>The corpus has the following annotation levels / entity types:</p> <ul> <li>protein</li> <li>protein_familiy_or_group</li> <li>protein_complex</li> <li>protein_variant</li> <li>protein_enum</li> </ul> <p>For definitions of the annotation levels, please refer to the Proteins-guidelines-final.doc file that is found in the download package.</p> <p>To achieve a large coverage of biological subdomains, document from multiple other protein / gene corpora were reannotated. For further coverage, new document sets were created. All documents are abstracts from PubMed/MEDLINE. The corpus is made up of the union of all the documents in the different subcorpora.<br> All document are delivered as MMAX2 (http://mmax2.net/) annotation projects.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Neurogenomic divergence during speciation by reinforcement of mating behaviors in chorus frogs (Pseudacris) – De novo reference transcriptome raw data, contigs and gene annotations

<p>RNA-Seq raw data used in the assembly and annotation of a reference transcriptome for the Upland Chorus Frog, <em>Pseudacris feriarum</em>. Raw data were&nbsp;obtained by sequencing of four tissue types: Brain, eyes, testis, and somatic. Assembled contigs (Trinity) and gene annotations (Trinotate) are also provided.</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Amino acid sequences of annotated genes in Pelargonium zonale

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo36/100

Gene Ontology annotation for Candida auris genome

<p>GO annotation for Candida auris genome B8441</p>

opencc-by-4.0Nov 2023View details →
dryad36/100

Data from: Gene modelling and annotation for the Hawaiian bobtail squid, Euprymna scolopes

<p>Coleoid cephalopods possess numerous complex, species-specific morphological and behavioural adaptations, e.g., a uniquely structured nervous system that is the largest among the invertebrates. The Hawaiian bobtail squid Euprymna scolopes is one of the most established cephalopod species. With its recent publication of the chromosomal-scale genome assembly and regulatory genomic data, it also emerges as a key model for cephalopod gene regulation and evolution. However, the latest genome assembly has been lacking a native gene model set. Our manuscript describes the generation of new long-read transcriptomic data and, combined with a plethora of available transcriptomic datasets, a new reference annotation for <em>E. scolopes</em>. </p>

opencc-zeroDec 2023View details →
zenodo36/100

Vibrant_metabolic_gene_annotations_vOTUs

<p>Vibrant v. 1.2.0 (default parameters; Kieft, Zhou, and Anantharaman 2020) was used to provide gene annotations for vOTUs that are &lt;10kb or not detected by VirSoter. <span>Although annotation is available for these vOTUs, we chose not to include the data in our results and discussion because of the uncertainty surrounding their validity.</span></p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Dataset associated to "Annotation matters: the effect of structural gene annotation on orthology"

<p>Dataset including the input and output files in "Annotation matters: the effect of structural gene annotation on orthology".</p> <ul> <li>Input proteomes for OMA and their corresponding splice files are in the OMAproteomes zipped folder. The OMA results for each annotation method are in the zipped folders with the method name (e.g. Augustus.zip).</li> <li>The fasta files (proteomes) are the same for OrthoFinder input in the cases of UniProt and Augustus (as they only have one isoform per gene). In these cases, the OrthoFinder folders (e.g. OFUniProt.zip), include the OrthoFinder output for that proteomes set. For Ensembl and NCBI, given the different approach each orthology method follows, the specific orthofinder proteomes are also included in the OrthoFinder (OF) zipped folder (e.g. OFtopNCBI.zip), in their corresponding primary_transcripts subfolder.&nbsp;</li> <li>The folder GSTDBenchmarOutput.zip contains the results from the Generalized Species Tree Discordance Benchmark.</li> <li>topNCBI/topEnsembl correspond to the original proteomes downloaded from the databases.</li> <li>priNCBI/primEnsembl correspond to the proteomes sets which include only the genes found on the primary assembly (reference sequences).</li> <li>For the species code to species name correspondance, please check the Code-Species.csv file.</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Genome and gene annotation for yeast strain SK1 used in "Deciphering the "m6A Code" via Antibody-Independent Quantitative Profiling"

<p>Genome and gene annotation used in &quot;Deciphering the &ldquo;m6A Code&rdquo; via Antibody-Independent Quantitative Profiling&quot; provided by Schraga Schwartz from his time at the Broad Institute.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Predicted genes annotated with Prokka for the 623 archaeal UBA genomes

<p>Gene prediction and annotation for the 623 archaeal UBA metagenome-assembled genomes using Prokka v1.12 with Pfam v31 and UniProt databases created on April 17, 2017 according to the Prokka instructions.</p> <p>These genomes are described in:</p> <p>Parks DH, et al. 2017. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol, doi:10.1038/s41564-017-0012-7/ .</p> <p>https://www.nature.com/articles/s41564-017-0012-7</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Gene and repeat annotation for common eider (Somateria mollissima)

<p>Here we provide the gene and repeat annotation for common eider (Somateria mollissima). It is unfortunately currently not possible to upload repeat annotation tracks to an international nucleotide sequence database such as ENA. While uploading the gene annotation is possible, some of the cross references to different databases in the functional annotation is removed. Further, the names of the entries in the publicly available genome assemblies on ENA have different names that what is found in the annotation tracks here, so we also provide the FASTA files for the assemblies. Ideally, all this should have been available via ENA.</p> <p>We annotated the genome assemblies using a pre-release version of the EBP-Nor genome annotation pipeline (<a href="https://github.com/ebp-nor/GenomeAnnotation">https://github.com/ebp-nor/GenomeAnnotation</a>). First, AGAT (https://zenodo.org/record/7255559) agat_sp_keep_longest_isoform.pl and agat_sp_extract_sequences.pl were used on the GRCg7b (GCA_016699485.1) chicken genome assembly and annotation to generate one protein (the longest isoform) per gene. Miniprot (Li, 2023) was used to align the proteins to the curated assemblies. UniProtKB/Swiss-Prot (Consortium et al., 2022) release 2022_03 in addition to the vertebrata part of OrthoDB v11 (Kuznetsov et al., 2022) were also aligned separately to the assemblies. Red (Girgis, 2015) was run via redmask (<a href="https://github.com/nextgenusfs/redmask">https://github.com/nextgenusfs/redmask</a>) on the assemblies to mask repetitive areas. In addition, we ran Earl Grey (Baril et al., 2023) to annotate transposable elements. GALBA (Brůna et al., 2023; Buchfink et al., 2015; Hoff and Stanke, 2018; Li, 2023; Stanke et al., 2006) was run with the chicken proteins using the miniprot mode on the masked assemblies. The funannotate-runEVM.py script from Funannotate was used to run EvidenceModeler (Haas et al., 2008) on the alignments of chicken proteins, UniProtKB/Swiss-Prot proteins, vertebrata proteins and the predicted genes from GALBA. The resulting predicted proteins were compared to the protein repeats that Funannotate distributes using DIAMOND blastp&nbsp; and the predicted genes were filtered based on this comparison using AGAT. The filtered proteins were compared to the UniProtKB/Swiss-Prot release 2022_03 using DIAMOND (Buchfink et al., 2015) blastp to find gene names and InterProScan&nbsp; was used to discover functional domains. AGATs agat_sp_manage_functional_annotation.pl was used to attach the gene names and functional annotations to the predicted genes. EMBLmyGFF3 (Norling et al., 2018) was used to combine the fasta files and GFF3 files into a EMBL format for submission to ENA.</p> <p>The assemblies provided here can also be found at ENA under accessions PRJEB61097 (pseudo-haplotype one with sex chromosomes; https://www.ebi.ac.uk/ena/browser/view/PRJEB61097) and PRJEB62037 (pseudo-haplotype two; https://www.ebi.ac.uk/ena/browser/view/PRJEB62037).</p> <p><strong></strong>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Shaw Lab Protein Coding Gene Annotation Table

<p>1. Original Table was downloaded from https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz</p> <p>2. Table was then merged with https://ftp.ncbi.nih.gov/gene/DATA/gene2ensembl.gz on April 8, 2024</p> <p>3. The table was appended with CellMark 2.0 https://pubmed.ncbi.nlm.nih.gov/36300619/</p> <p>4. Appended genes encoding surfaceome proteins GESPs from the supplementary table S1-S40 https://pubmed.ncbi.nlm.nih.gov/35121907/</p> <p>5. Appended mouse homology genes from https://www.informatics.jax.org/downloads/reports/HOM_MouseHumanSequence.rpt source code used in the processing is https://github.com/shawlab-moffitt/Human2MouseGeneSymbolConversion</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Automated cell annotation in scRNA-seq data using unique marker gene sets

<p>Single-cell RNA sequencing has revolutionized the study of cellular heterogeneity, yet accurate cell type annotation remains a significant challenge. Inconsistent labels, technological variability, and limitations in transferring annotations from reference datasets hinder precise annotation. This study presents a novel approach for accurate cell type annotation in scRNA-seq data using unique marker gene sets. By manually curating cell type names and markers from 280 publications, we verified marker expression profiles across these datasets and unified nomenclatures to consistently identify 166 cell types and subtypes. Our customized algorithm, which builds on the AUCell method, achieves accurate cell labeling at single-cell resolution and surpasses the performance of reference-based tools like Azimuth, especially in distinguishing closely related subtypes. To enhance accessibility and practical utility for researchers, we have also developed a user-friendly application that automates the cell typing process, enabling efficient verification and supporting comprehensive downstream analyses. The desktop application can be accessed at <a href="https://omnibusx.com/">https://omnibusx.com</a>.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record