Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations.
<p>Genes predicted by OrthoFiller not found in public genome annotations for the species analysed in the paper titled "OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations."</p>
DeepAnnotation: A novel interpretable deep learning-based genomic selection model that integrates comprehensive functional annotations
<p>1. Update package, example dataset, and demo code of DeepAnnotation</p> <p>2. Update the transformed genotype data, the phenotype data, the comprehensive functional annotation data for Duroc prepared by RNAfold, DeepSEA, easyMF models, and the four types of input data for training DeepAnnotation model</p> <p>3. Add the conserved functional annotation</p> <p> </p>
Genome and annotation of microalgae Trebouxia lynnae
<p>This dataset includes:</p> <ul> <li>Genome sequence in fasta format of the symbiont lichen algae <em>Trebouxia lynnae</em></li> <li>Structural genome annotation in gff3 format</li> <li>Funcional genome annotation in tabular format</li> </ul> <p>It has been published in the paper "From spores to gametes: A sexual life cycle in a symbiotic Trebouxia microalga" in Algal Research. DOI: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.algal.2024.103744" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.algal.2024.103744</span></span></a></p>
Gene Ontology annotation for Candida auris genome
<p>GO annotation for Candida auris genome B8441</p>
A root-specific NLR network confers resistance to plant parasitic nematodes - genomic sequences and annotations
<p>Sequence and annotation data associated with "A root-specific NLR network confers resistance to plant parasitic nematodes"</p>
Annotated genome assemblies for Geoscapheus dilatatus, Panesthia cribrata and Neogeoscapheus hanni
<p>Genetic changes that enabled the evolution of eusociality have long captivated biologists. More recently, attention has focussed on the consequences of eusociality on genome evolution. Studies have reported higher molecular evolutionary rates in eusocial hymenopteran insects compared with their solitary relatives. To investigate the genomic consequences of eusociality in termites, we sequenced genomes from three of their non-eusocial cockroach relatives. Using a phylogenomic approach, we found that termite genomes experienced lower rates of synonymous mutations than those of cockroaches, possibly as a result of longer generation times. We identified higher rates of nonsynonymous mutations in termite genomes than in cockroach genomes, and identified pervasive relaxed selection in the former (24–31% of the genes analysed) compared with the latter (2–4%). We infer that this is due to a reduction in effective population size, rather than gene-specific effects (e.g., indirect selection of caste-biased genes). We found no obvious signature of increased genetic load in termites, and postulate efficient purging of deleterious alleles at the colony level. Additionally, we identified genomic adaptations that may underpin caste formation, such as genes involved in post-translational modifications. Our results provide insights into the evolution of termites and the genomic consequences of eusociality more broadly.</p>
Data from: Genome-scale annotation of protein binding sites via language model and geometric deep learning
<p>The dataset contains the training and test sets of protein binding sites with DNA, RNA, peptide, protein, ATP, HEM, Zn2+, Ca2+, Mg2+ and Mn2+. Each protein is associated with 3 lines indicating the protein name (PDB accession code and chain), sequence and residue labels (0 for non-binding and 1 for binding), respectively. The ESMFold-predicted structures are also provided.</p>
Annotation files for B. asplanchnoidis genome (Stelzer et al. 2021, BMC Biology)
<p>This dataset contains the annotation files for the genome of Brachionus asplanchnoidis (Rotifera).</p> <p>The genome has been published here:</p> <p>Stelzer, C.P., Blommaert, J., Waldvogel, A.M. <em>et al.</em> Comparative analysis reveals within-population genome size variation in a rotifer is driven by large genomic elements with highly abundant satellite DNA repeat elements. <em>BMC Biol</em> <strong>19</strong>, 206 (2021).</p> <p>https://doi.org/10.1186/s12915-021-01134-w</p> <p> </p>
Genome assembly and annotation for temperate coral Astrangia poculata
<p>Initial release of the genome assembly, gene prediction, and functional annotation for the temperate coral <em>Astrangia poculata. </em></p> <p>The genomic resources made available here are described in <a href="https://www.biorxiv.org/content/10.1101/2023.09.22.558704v1.full">Stankiewicz et al., 2023</a>. The files are as follows:</p> <p> </p> <p>Genome assembly:</p> <ul> <li>apoculata.genome.fasta.gz --unmasked version of the genome assembly</li> <li>apoculata.genome.masked.fasta.gz --masked version of the genome assembly</li> <li> <p><span>apoculata_circularized_mitogenome.fasta --circularized mitochondrial genome assembly</span></p> </li> </ul> <p>Annotation:</p> <ul> <li>apoculata_GeneAnnotation_combined.txt.gz: function annotation</li> <li>apoculata.gff3.gz: gene predictions in GFF3 </li> <li>apoculata.gtf.gz: gene predictions in GTF </li> <li>apoculata_cds.fasta.gz: coding sequences </li> <li>apoculata_longest_trans_cds.fasta.gz: coding sequences filtered to just the longest translatable per gene </li> <li>apoculata_mrna.fasta.gz: mRNA sequences </li> <li>apoculata_proteins.fasta.gz: protein sequences</li> </ul>
Chromosome-scale genome assembly and de novo annotation of Alopecurus aequalis.
<p><em>Alopecurus aequalis</em> is a winter annual or short-lived perennial bunchgrass which has in recent years emerged as the dominant agricultural weed of barley and wheat in certain regions of China and Japan, causing significant yield losses. Its robust tillering capacity and high fecundity, combined with the development of both target and non-target-site resistance to herbicides means it is a formidable challenge to food security. Here we report on a chromosome-scale assembly of <em>A. aequalis</em> with a genome size of 2.83 Gb. The genome contained 33,758 high-confidence protein-coding genes with functional annotation. Comparative genomics revealed that the genome structure of <em>A. aequalis</em> is more similar to <em>Hordeum vulgare </em>rather than the more closely related <em>Alopecurus myosuroides</em>. The datasets provided here are the assembly FASTA file (lpAloAequ1.1.prim.cur.20230912.fasta.gz), the high-confidence protein-coding genes (Alaeq_EIv0.2.release_HC_genes.gff3.gz) and the full annotation which includes both low and high confidence features of all biotypes (Alaeq_EIv0.2.release.gff3.gz) </p>
Diaphorina citri new genome annotations and GO terms (2021)
<p>Additional genome annotations for the <em>Diaphorina citri</em> genome (Diaci_v3) including: TEs, intergenic regions and putative promotors, also additional gene ontology terms annotated using EGGnog. </p>
Genomic, transcriptomic and proteomic comparison of MRSA CC398 isolates collected from human and wild animal samples (Genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following methicillin-resistant <em>Staphylococcus aureus</em> (MRSA) strains: MRSA CC398 isolates recovered from humans, namely C5621 and C9017, and from a wild boar, namely OR418.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB35102).</p>
Annotated genome Pseudomonas species MWU 12.3088
<p>Annotated genome Pseudomonas species MWU 12.3088 isolated from cultivated cranberry bogs in southeastern Massachusetts. </p>
Annotated genomes of Pseudomonas spp. isolated from wild and cultivated cranberry bogs
<p><em>Pseudomonas</em> spp. were isolated from rhizospheres of wild and cultivated cranberry bogs in southeastern Massachusetts. Genomes were translated and annotated.</p>
Annotated genome Pseudomonas species MWU 12.2319
<p>Annotated genome Pseudomonas species MWU 12.2319 isolated from cultivated cranberry bogs in southeastern Massachusetts.</p>
Annotation of the Priestia megaterium MWU16_30321 genome
<p>Annotation of the <em>Priestia megaterium</em> MWU16_30321 genome </p>
Annotated genome Pseudomonas species MWU 12.3091
<p>Annotated genome Pseudomonas species MWU 12.3091 isolated from cultivated cranberry bogs in southeastern Massachusetts.</p>
Annotated Genome Pseudomonas sp MWU 13-2517
<p>Annotated genome Pseudomonas sp MWU 13-2517 isolated from wild cranberry bogs in southeastern Massachusetts.</p>
Annotated Genome Pseudomonas sp MWU 13-2862
<p>Annotated genome Pseudomonas sp MWU 13-2862 isolated from wild cranberry bogs in southeastern Massachusetts. </p>
A high-quality genome assembly and annotation of the dark-eyed junco Junco hyemalis, a recently diversified songbird
<p>The dark-eyed junco (<i>Junco hyemalis</i>) is one of the most common passerines of North America, and has served as a model organism in studies related to ecophysiology, behavior and evolutionary biology for over a century. It is composed by at least six distinct, geographically structured forms of recent evolutionary origin presenting remarkable variation in phenotypic traits, migratory behavior and habitat. Here we report a high-quality genome assembly and annotation of the dark-eyed junco generated using a combination of shotgun libraries and proximity ligation Chicago<sup>TM</sup> and Dovetail HiC<sup>TM</sup> libraries. The final assembly is 1,031,523,571 bp long, with 98.3% of the sequence located in 30 full or nearly full chromosome scaffolds, and with a N50/L50 of 71,3 Mb/5 scaffolds. We identified 19,026 functional genes combining gene prediction and similarity approaches, of which 15,967 were associated to GO terms. Genome assembly and annotated set of genes yielded 95.4% and 96.2% completeness scores, respectively, when compared with the BUSCO avian dataset. This new assembly for <i>J. hyemalis </i>provides a valuable resource for genome evolution analysis, as well as for identifying functional genes involved in adaptive processes and speciation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.