Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

345

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

345 results for “genome annotation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Alliance of Genome Resources Disease Annotations

<p>Tab separated formatted spreadsheets of disease annotations from the Alliance of Genome Resources. Annotations are to terms in the Disease Ontology (DO), evidence codes from the Evidence Code Ontology (ECO).&nbsp; PubMed ids (PMID) are provided for the source of the annotation assertions.</p> <p>File includes annotations for the following organisms:</p> <ul> <li>Homo sapiens (human; NCBI:txid 9606)</li> <li>Caenorhabditis elegans (nematode; NCBI:txid 6239)</li> <li>Danio rerio (zebrafish;NCBI:txid 7955)</li> <li>Drosophila melanogaster (fruit fly; NCBI:txid 7227)</li> <li>Mus musculus (mouse; NCBI:txid10090)</li> <li>Rattus norvegicus (rat; NCBI:txid 10116)</li> <li>Saccharomyces cerevisiae (yeast; NCBI:txid 559292)</li> <li>Xenopus laevis (African clawed frog; NCBI:txid 8355)</li> <li>Xenopus tropicalis (Western clawed frog; NCBI:txid 8364)</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Supplemental material of "An annotated whole-genome multilocus sequence typing schema for scalable high resolution typing of Streptococcus pyogenes"

<p>This supplemental material includes the genome assemblies, associated metadata and analysis results for five datasets used to define a publicly available annotated wgMLST schema for <em>S. pyogenes</em> and to evaluate its suitability for high resolution typing. A brief description for each file in the dataset is available in the included README file. Raw sequencing data and sample metadata for the 265 isolates included in Dataset1 have been deposited in the European Nucleotide Archive (ENA) under Project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB49967?show=reads">PRJEB49967</a>.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Whole genome sequence and annotation dataset of rare actinobacteria, Barrientosiimonas humi gen. nov., sp. nov. 39T from Antarctica

<p>The present data files are the source files of the annotation output from the whole genome sequencing of rare actinobacteria, <em>Barrientosiimonas humi gen. nov., sp. nov.</em> 39<sup>T</sup> from Antarctica.</p> <p>The dataset of the whole-genome sequence of <em>B. humi</em> had been deposited in European Nucleotide Archive (ENA) repository under the accession number PRJEB44986 / ERP129097, direct URL to data:<strong> </strong><a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB44986">https://www.ebi.ac.uk/ena/browser/view/PRJEB44986</a></p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Whole genome sequence and annotation of Penstemon davidsonii

<p><em><span>Penstemon</span></em><span> is the most speciose flowering plant genus endemic to North America. <em>Penstemon</em> species' diverse morphology and adaptation to various environments have made them a valuable model system for studying evolution, but the absence of publicly available reference genomes limits possible research directions. Here we report the first reference genome assembly and annotation for <em>Penstemon</em> <em>davidsonii</em>. Using PacBio long-read sequencing and Hi-C scaffolding technology, we constructed a de novo reference genome of 437,568,744 bases, with a contig N50 of 40 Mb and L50 of 5. The annotation includes 18,199 gene models, and both the genome and transcriptome assembly contain over 95% complete eudicot BUSCOs. This genome assembly will serve as a valuable reference for studying the evolutionary history and genetic diversity of the <em>Penstemon</em> genus.</span></p>

opencc-zeroOct 2023View details →
zenodo40/100

Genomes of cactophilic Drosophila species and their respective gene and TE annotations + QC data

<p>The genomes deposited in this repository refers to the data used in the study entitled "Transposable elements contribute to the evolution of host shift-related genes in cactophilic<em> Drosophila</em> species", from Oliveira D. S., Larue A., Nunes W. V. B., Sabot F., Bodel&oacute;n A., Garc&iacute;a Guerreiro M. P., Vieira C., Carareto C. M. A.</p> <p>Each genome has its following assembly (fasta), gene annotation (gff), and TE annotation (gtf). The quality control for the nanopore genomes can be accessed on QC_nanopore_genomes.zip, and the quality control for the RNA-seq data on QC_RNAseq.zip.</p> <p>Additional files, as code and input files to reproduce the specific analysis of the manuscript, are also provided in the github repository: https://github.com/OliveiraDS-hub/Pipelines-Cactophilic-Drosophila-Species</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad40/100

Chromosome-level genome assembly and annotation of the emblematic silver-lipped pearl oyster, Pinctada maxima Jameson, 1901

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad40/100

Data from: Regulatory genome annotation for 33 insect species

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad40/100

Whole genome sequence and annotation of Penstemon davidsonii

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad40/100

Annotation of genes encoding enzymes across marine phytoplankton genomes

Open the record for dataset details and reuse information.

publicApr 2023View details →
zenodo36/100

Genome annotation of Macquarie perch

<p>Intermediate and final files generated from the repeat masking and protein coding gene prediction of the Macquarie perch genome (NCBI Bioproject:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA516983">PRJNA516983</a>)</p> <p>Repeat_Annotation.tar.gz: RepeatMasker and RepeatModeler output - contains de novo repeat library , repeat-masked (soft-masked) genome and repeat annotation in gff3 format.</p> <p>BUSCO_Genome.tar.gz: BUSCO completeness (actinopterygii_odb9) calculation based on whole-genome sequence (-m genome).&nbsp;</p> <p>BUSCO_Protein.tar.gz: BUSCO completeness (actinopterygii_odb9) calculation based on predicted proteins (-m prot)</p> <p>MP.gff3: BRAKER2 annotation in gff3 format</p> <p>MP.gtf: BRAKER2 annotation in gtf format</p> <p>MP.codingseq.fna: BRAKER2 coding sequence output (DNA sequences)</p> <p>MP.faa: BRAKER2 translated coding sequence output (protein sequences)</p> <p>MP_nointernalstop.faa: Filtered BRAKER2 translated coding sequence (no&nbsp;sequences with&nbsp;internal stop codon)</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Annotated genes harboring major effect markers (R2 ≥ 15%). Highlighted in green are genes annotated from Rhodes et al. 2014,2017, in orange genes annotated as similar to Peroxidase, in yellow new annotations from sorghum genome in Atlas. In the first three columns start and stop position on the sorghum genome and transcript name, followed by the nearest marker name and the distance of the gene from the nearest marker, then a column where are shown the GWAS methods and target traits for which the linked SNP was significant, the last column shows the category of the genes.

<p><strong>We conducted a comprehensive genomics study to map genomic loci determining the production of antioxidants in sorghum grains. Encouraging results were obtained and published in peer-reviewed article with impact factor (https://doi.org/10.1371/journal.pone.0225979). Annotated genes harboring major effect markers (R<sup>2</sup> &ge; 15%) were identified and will be of worldwide interest. </strong></p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Whole genome assembly and gene annotation of a diploid genotype of Brachiaria ruziziensis (syn. Urochloa ruziziensis)

<p>In this work, we have presented a comprehensive analysis of the molecular mechanism linked to aluminium tolerance in <em>Brachiaria</em> species. By assembling and annotating a diploid genotype of <em>B. ruziziensis</em> we have developed the capability for genomic-based studies of desirable phenotypic traits. Using this resource, we have identified three QTLs associated to root architecture and vigour during Al<sup>3+</sup> stress in a hybrid population from a high and low tolerant accession. We have also identified a number of genes and molecular responses that impact on different aspects of signalling, cell-wall composition and active transports as a response to aluminium stress. <em>Brachiaria </em>tolerance appears to build in the same genes than in rice. However, we found that external mechanisms such as sequestration of Al<sup>3+</sup> common in other grasses might be not that important in <em>Brachiaria. </em>Also, contrasting regulation in the same genotype after 8 or 72 hours of Al<sup>3+</sup> stress of numerous genes involved in RNA translation can explain the different levels of tolerance among different Brachiaria species. The newly annotated draft genome represents an important base upon which study other aspects of <em>Brachiaria</em> biology.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Megalopta genalis genome annotation files

<p>This repository contains genome annotation files for the genome assembly reported in Kapheim et al. 2020.</p>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Updated functional annotation of the Mycobacterium bovis AF2122/97 reference genome - datasets

<p>This is an archive of the repository held under https://github.com/dmnfarrell/gordon-group/tree/master/mbovis_annotation</p> <p>It contains a notebook and required input files for updating the MTB/Mbovis AF2122/97 genomes with new protein product annotations from literature.</p> <p>D Farrell UCD February 2020</p> <p>Updates to M.bovis genome (Aug 2019):<br> added 611 product annotations<br> added 689 gene names<br> added 5 manually edited entries from pdb hits<br> removed locus_tags for repeat_region features<br> added prefix to tRNA and mobile_element features for consistency</p> <p><br> References<br> Updated functional annotation of the Mycobacterium bovis AF2122/97 reference genome (https://www.biorxiv.org/content/10.1101/757823v1)</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Assemblies and annotations from the paper "Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria"

<p>These files are the assemblies and annotations produced and analyzed in the revised version of the manuscript entitled &quot;Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria&quot;.</p>

opencc-by-4.0Dec 2019View details →
dryad36/100

A high-quality genome assembly and annotation of the gray mangrove, Avicennia marina

<p class="CxSpFirst">The gray mangrove [<i>Avicennia marina</i> (Forsk.) Vierh.] is the most widely distributed mangrove species, ranging throughout the Indo-West Pacific. It presents remarkable levels of geographic variation both in phenotypic traits and habitat, often occupying extreme environments at the edges of its distribution. However, subspecific evolutionary relationships and adaptive mechanisms remain understudied, especially across populations of the West Indian Ocean. High-quality genomic resources accounting for such variability are also sparse. Here we report the first chromosome-level assembly of the genome of <i>A. marina</i>. We used a previously release draft assembly and proximity ligation libraries Chicago and Dovetail HiC for scaffolding, producing a 456,526,188 bp long genome. The largest 32 scaffolds (22.4 Mb to 10.5 Mb) accounted for 98 % of the genome assembly, with the remaining 2% distributed among much shorter 3,759 scaffolds (62.4 Kb to 1 Kb). We annotated 45,032 protein-coding genes using tissue-specific RNA-seq data in combination with <i>de novo</i> gene prediction, from which 34,442 were associated to GO terms. Genome assembly and annotated set of genes yield a 96.7% and 95.1% completeness score, respectively, when compared with the eudicots BUSCO dataset. Furthermore, an F<sub>ST</sub> survey based on resequencing data successfully identified a set of candidate genes potentially involved in local adaptation, and revealed patterns of adaptive variability correlating with a temperature gradient in Arabian mangrove populations. Our <i>A. marina </i>genomic<i> </i>assembly provides a highly valuable resource for genome evolution analysis, as well as for identifying functional genes involved in adaptive processes and speciation.</p>

opencc-zeroMay 2020View details →
dryad36/100

Data from: An annotated draft genome of the mountain hare (Lepus timidus)

<p>Hares (genus Lepus) provide clear examples of repeated and often massive introgressive hybridization and striking local adaptations. Genomic studies on this group have so far relied on comparisons to the European rabbit (Oryctolagus cuniculus) reference genome. Here, we report the first de novo draft reference genome for a hare species, the mountain hare (Lepus timidus), and evaluate the efficacy of whole-genome re-sequencing analyses using the new reference versus using the rabbit reference genome. The genome was assembled using the ALLPATHS-LG protocol with a combination of overlapping pair and mate-pair Illumina sequencing (77x coverage). The assembly contained 32,294 scaffolds with a total length of 2.7 Gb and a scaffold N50 of 3.4 Mb. Re-scaffolding based on the rabbit reference reduced the total number of scaffolds to 4,205 with a scaffold N50 of 194 Mb. A correspondence was found between 22 of these hare scaffolds and the rabbit chromosomes, based on gene content and direct alignment. We annotated 24,578 protein coding genes by combining ab-initio predictions, homology search, and transcriptome data, of which 683 were solely derived from hare-specific transcriptome data. The hare reference genome is therefore a new resource to discover and investigate hare-specific variation. Similar estimates of heterozygosity and inferred demographic history profiles were obtained when mapping hare whole-genome re-sequencing data to the new hare draft genome or to alternative references based on the rabbit genome. Our results validate previous reference-based strategies and suggest that the chromosome-scale hare draft genome should enable chromosome-wide analyses and genome scans on hares.</p>

opencc-zeroOct 2020View details →
zenodo36/100

Itag2.3 Tomato Genome Annotation, RDF graph

<p>Annotation of the tomato genome, ITAG2.3 (ftp://ftp.sgn.cornell.edu/genomes/Solanum_lycopersicum/annotation/ITAG2.3_release/). GFF file&#39;s were converted into a RDF graph.</p>

opencc-zeroNov 2015View details →
zenodo36/100

COMBAT TB Tuberculosis genome annotation database

<p>This is a Neo4j format *M. tuberculosis* reference annotation database. To view it you need&nbsp;a Neo4j (v 2.3)&nbsp;instance. There is a script (<strong>run_db.sh</strong>) that will start a Neo4j instance for&nbsp;you based on this data, using docker. Run that (<strong>bash run_db.sh</strong>) and connect to http://localhost:7474.</p> <p>The database was created by the COMBAT TB project&nbsp;(http://christoffels.sanbi.ac.za/index.php/projects/combat-tb) at the South African National Bioinformatics Institute (SANBI).<br /> <br /> Authors: Thoba Lose, Peter van Heusden, Ziphozakhe Mashologu, Alan Christoffels .</p> <p>The COMBAT TB project is funded by the South African Medical Research Council (MRC) and was supported by the South African<br /> Research Chairs Initiative of the Department of Science and Technology and National Research Foundation of South Africa.</p>

opencc-by-sa-4.0Jun 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record