Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
Data from: Chromosome scale genome assemblies and annotations for Poales species Carex cristatella, Carex scoparia, Juncus effusus and Juncus inflexus
<p>The majority of sequenced genomes in the Monocots are from species belonging to the Poaceae, which includes many commercially important crops. Here, we expand the number of sequenced genomes from the Monocots to include the genomes of four related Cyperids: <em>Carex cristatella</em> and <em>Carex scoparia</em> from Cyperaceae and <em>Juncus effusus</em> and <em>Juncus inflexus</em> from Juncaceae. The high-quality, chromosome-scale genome sequences from these four Cyperids were assembled by combining whole-genome shotgun sequencing of Nanopore long reads, Illumina short reads, and Hi-C sequencing data. Some members of the Cyperaceae and Juncaceae are known to possess holocentric chromosomes. We examined the repeat landscapes in our sequenced genomes to search for potential repeats associated with centromeres. Several large satellite repeat families, comprising 3.2% to 9.5% of our sequenced genomes, showed dispersed distribution of large repeat clusters across all <em>Carex</em> chromosomes, with few instances of these repeats clustering in the same chromosomal regions. In contrast, most large <em>Juncus</em> satellite repeats were clustered in a single location on each chromosome, with sporadic instances of large satellite repeats throughout the Juncus genomes. Recognizable transposable elements account for about 20% of the assemblies, with the <em>Carex</em> genomes containing more DNA transposons than retrotransposons while the converse is true for the <em>Juncus</em> genomes. These genome sequences and annotations will facilitate better comparative analysis within monocots.</p>
Gene annotation of complete genome sequence of Achromobacter sp. strain E1
Open the record for dataset details and reuse information.
Genome annotations for: Multi-omics approaches define novel aphid effector candidates associated with virulence and avirulence phenotypes
<div> <p><span><span>Peter Thorpe</span></span><span><span>1</span></span><span><span>, Simone Altmann</span></span><span><span>1</span></span><span><span>, Rosa Lopez-Cobollo</span></span><span><span>2</span></span><span><span>, Nadine Douglas</span></span><span><span>3</span></span><span><span>, Javaid Iqbal</span></span><span><span>2</span></span><span><span>, Sadia Kanvil</span></span><span><span>2</span></span><span><span>, </span><span>Jean-Christophe Simon</span></span><span><span>4</span></span><span><span>, </span><span>James </span><span>C. </span><span>Carolan</span></span><span><span>3</span></span><span><span>, Jorunn Bos</span></span><span><span>1</span><span>*</span></span><span><span>, Colin Turnbull</span></span><span><span>2</span><span>*</span></span><span> </span></p> </div> <div> <p><span><span>1</span></span><span><span>School of Life Sciences, University of Dundee, UK;</span> </span><span><span>2</span></span><span><span>Department of Life Sciences, Imperial College London, UK; </span></span><span><span>3</span></span><span><span>Department of Biology, Maynooth University, </span><span>Republic of Ireland</span><span>; </span></span><span><span>4</span></span> <span><span>INRAE , France</span><span>. *Authors for correspondenc</span><span>e: </span></span><a href="mailto:j.bos@dundee.ac.uk" target="_blank" rel="noreferrer noopener"><span><span>j.bos@dundee.ac.uk</span></span></a><span><span>, </span></span><a href="mailto:c.turnbull@imperial.ac.uk" target="_blank" rel="noreferrer noopener"><span><span>c.turnbull@imperial.ac.uk</span></span></a><span><span>. </span></span><span> </span></p> <p> </p> <p><span>This is a repository for the version3 gene predictions and annotation for the pea aphid used for the publication:</span></p> <p> </p> <p><strong><span><span><span>Multi-omics approaches define novel aphid effector candidates associated with </span><span>virulence and </span><span>avirulence</span> <span>phenotypes</span></span><span> </span></span></strong></p> <p> </p> <div> <p><span><span>ABSTRACT</span></span></p> </div> <div> <p><span><span>Background</span></span><span><span>. Compatibility between aphids and plant hosts is genetically </span><span>determined</span><span> by both interacting organisms. For example, plants may carry resistance (R) genes or deploy chemical defences. Aphid saliva </span><span>contains</span><span> many proteins that are secreted into host tissues. </span><span>S</span><span>ubset</span><span>s</span><span> of these proteins are predicted to act as effectors, either subverting or triggering host immunity. However, associating </span><span>particular effectors</span><span> with virulence or </span><span>avirulence</span><span> outcomes presents challenges due to the combinatorial complexity. Here we use defined aphid and host genetics to test for co-segregation of expressed aphid transcripts and proteins with virulent or avirulent phenotypes.</span></span><span> </span></p> </div> <div> <p><span><span>Results</span><span>. </span></span><span><span>We compared virulent and avirulent pea aphid parental genotypes, and their bulk segregant F</span></span><span><span>1</span></span><span><span> progeny on </span></span><span><span>Medicago </span><span>truncatula</span> </span><span><span>genotypes</span></span> <span><span>carrying or lacking the </span></span><span><span>RAP1 </span></span><span><span>resistance </span><span>quantitative trait locus</span><span>. </span><span>D</span><span>ifferential expression </span><span>analysis based on </span><span>RNA sequencing </span><span>of whole bod</span><span>y and head samples, </span><span>in combination with proteomics of saliva and salivary glands</span><span>,</span> <span>enabled us </span><span>to pinpoint proteins </span><span>associated</span><span> with virulence/</span><span>avirulence</span><span> phenotypes. </span></span><span><span><span>There was relatively </span><span>little impact</span><span> of </span></span></span><span><span><span>host genotype, </span></span></span><span><span><span>whereas</span><span> l</span></span></span><span><span>arge numbers of transcripts and proteins were differentially expressed between parental aphids, </span><span>likely a</span><span> reflection of their classification as divergent biotypes within the pea aphid species complex. Many fewer </span><span>transcripts</span> <span>intersected with the equivalent differential expression patterns in the bulked F</span></span><span><span>1</span></span><span><span> progeny, providing an effective filter for removing </span><span>genomic </span><span>background effects</span></span><span><span>. </span><span>Overall, t</span><span>here were more upregulated genes detected in the </span><span>F</span></span><span><span>1</span></span><span> <span>avirulent </span><span>dataset </span><span>compared with the virulent one. </span><span>Some of the</span><span> differentially expressed transcripts </span><span>were also found in the differentially expressed proteomes</span><span>, with a</span><span>minopeptidase N prot</span><span>eins </span><span>being </span><span>t</span><span>he most frequent</span><span> differentially expressed</span><span> family</span></span><span><span>. </span><span>In addition</span><span>, a </span><span>substantial</span> <span>proportion </span><span>(26%) </span><span>of salivary proteins lack annotations, suggesting that </span><span>many </span><span>novel functions </span><span>remain</span><span> to be discovered. </span></span><span> </span></p> </div> <div> <p><span><span>Conclusions.</span></span><span><span> Especially when combined with tightly controlled genetics of both insect and host, multi-</span><span>omic</span><span> approaches are powerful tools for revealing and filtering candidate lists down to plausible genes for further functional analysis as putative </span><span>aphid </span><span>effectors.</span></span><span> </span></p> </div> </div>
Assemblies, associated annotation files, and analysis source data of Platanus x acerifloia genome
<p><em>Platanus</em> <span>× </span><em>acerifolia </em>(London plane; Platanaceae) is a major ornamental tree used worldwide. Platanaceae is one of the last early-diverging eudicot families without a complete nuclear genome assembly. Here, we assembled a high-quality, chromosome-level reference genome for <em>P.</em> <span>× </span><em>acerifolia.</em></p>
The gene structure annotation, gene function annotation and TE annatition files of the Glyphodes pyloalis's genome
Open the record for dataset details and reuse information.
Genome assemblies and annotations of wild grape V. davidii Föex (0940) and a cultivated grape V. vinifera L. Manicure Finger
<p>We newly sequenced genome of a wild grape (V. davidii (0940) (namely Vd)and a somatic of Manicure Finger (namely MF), and constructed phase-resolved genomes using Hi-C and HiFi reads.</p> <p>MF_hap1.fa.gz</p> <p>#Sequence of haplotype 1 from the MF genome</p> <p>MF_hap1.gff3.gz</p> <p>#Annotation of haplotype 1 from the MF genome</p> <p>MF_hap2.fa.gz</p> <p>#Sequence of haplotype 2 from the MF genome</p> <p>MF_hap2.gff3.gz</p> <p>#Annotation of haplotype 2 from the MF genome</p> <p>Vd_hap1.fa.gz</p> <p>#Sequence of haplotype 1 from the Vd genome</p> <p>Vd_hap1.gff3.gz</p> <p>#Annotation of haplotype 1 from the Vd genome</p> <p>Vd_hap2.fa.gz</p> <p>#Sequence of haplotype 2 from the Vd genome</p> <p>Vd_hap2.gff3.gz</p> <p>#Annotation of haplotype 2 from the Vd genome</p> <p> </p> <p> </p>
Population genetics of Paramecium mitochondrial genomes; genome assemblies and annotation files
<p>Because of issues arising during submission to GenBank of Paramecium mitochondrial genomes due to the highly unconventional nature of the genetic code used in these genomes, we are initially making the genomes publicly available here (while we are still working on a submission to official databases).</p>
ST131_4071_genome_assembly_annotation_files_theExtra2
<p>SRR3290149_prokka.gff.gz and ERR1971796_prokka.gff.gz</p>
Annotation for human genomic variation during the BMP4-induced conversion from embryonic stem cells to trophoblast by bone
<p>Whole genomic data of three cell lines were sequenced. These three cell lines are two hESC lines (H1 & H9) by invasion assays and iPSC cell line MRucR. These three cell lines are two hESC lines (H1 & H9) by invasion assays and iPSC cell line MRucR. Paired-end DNA libraries were prepared according to manufacturer’s instructions (Illumina Truseq Library Construction). And then, valid sequencing data is mapped to the reference genome (UCSC hg19) by Burrows-Wheeler Aligner (BWA) software. Reads that aligned to genomic regions were collected for mutation identification and subsequent analysis. Samtools mpileup and bcftools are used to do variant calling and identify SNP, indels. Control-free (Boeva V et al.2012) is utilized to do CNV detection. And BreakDancer (Chen K et al.2009) is applied to detect SV information. And, Single nucleotide variant (SNV) and small somatic insertions and deletions (indels) were identified using Strelka2.</p>
Genome, Annotation, and auxiliary files for the Puccinia striiformis f. sp. tritici DK0911 genome
<p>These are the genome, annotation, and auxiliary files for the Puccinia striiformis f. sp. tritici DK0911 genome. This dataset is associated with the publication "Distinct life histories impact dikaryotic genome evolution in the rust fungus Puccinia striiformis causing stripe rust in wheat" and Bioproject accession number PRJNA588102. </p>
Round goby Neogobius melanostomus genome annotation
<p>Annotation file for the round goby genome (sequence deposited as “RGoby_Basel_V2”, BioProject accession PRJNA549924, BioSample SAMN12099445, GenBank genome accession VHKM00000000, release date July 22 2019). </p> <p>Supplementary Material S1 for Adrian-Kalchhauser et al, BMC Biology, The round goby genome provides insights into mechanisms that may facilitate biological invasions.</p> <p>The round goby genome assembly was annotated using Maker v2.31.8. Two iterations were run with assembled transcripts from round goby embryonic tissue and data from eleven other actinopterygian species available in the ENSEMBL database as well as the SwissProt protein set from the uniprot database as evidence (downloaded March 2, 2016). In addition, an initial set of reference sequences obtained from a closely related species, the sand goby (<em>Pomatoschistus minutus</em>), sequenced by the IMAGO marine genomes project of the CeMEB consortium at University of Gothenburg, Sweden was included. The second maker iteration was run after first training the gene modeler SNAP version 2006-07-28 based on the results from the first run. Augustus v3.2.2 was run with initial parameter settings from zebrafish (<em>Danio rerio</em>). Repeat regions in the genome were masked using RepeatMasker known elements and repeat libraries from Repbase as well as <em>de novo</em> identified repeats from the round goby genome assembly obtained from a RepeatModeler analysis.</p>
The gene structure annotation, gene function annotation and TE annatition files for the Cibotium barometz isolate CiBa-2024 genome
<p>This dataset comprises comprehensive annotation files for the genome of Cibotium barometz (Golden Chicken Fern), isolate CiBa-2024. It includes gene structure predictions, functional annotations, and transposable element (TE) identifications, complementing the chromosome-level genome assembly. The gene structure annotation provides detailed information on predicted gene models, including exon-intron boundaries and coding sequences. Functional annotations offer insights into the potential roles of identified genes, including Gene Ontology (GO) terms, protein domains, and pathway associations. The TE annotation file details the classification and distribution of transposable elements within the genome. These annotations were generated using state-of-the-art bioinformatics tools and databases, offering a valuable resource for researchers studying fern genomics, plant evolution, and the genetic basis of C. barometz's unique biological features, including its medicinal properties. This dataset aims to facilitate further research in comparative genomics, functional studies, and the exploration of fern biology and evolution.</p>
Genome assembly and gene annotations for a Morus alba var. zhenzhubai
Open the record for dataset details and reuse information.
COG and Pfam annotation results of the genome sequences of four novel Endozoicomonas strains associated with the octocoral Litophyton in a long-term aquarium facility
<p>COG and Pfam annotation files from DOE-JGI Microbial Genome Annotation Pipeline (MGAP) version 4(1), for four <em>Endozoicomonas</em> strains associated with the tropical octocoral Litophyton in a long-term aquarium facility. Data correspond to the assemblies of NE35, NE40, NE41, and NE43, available under the BioProject accession numbers <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075803">PRJNA1075803</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075804">PRJNA1075804</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075805">PRJNA1075805</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/1075806">PRJNA1075806</a>, respectively. Results were submitted to the Integrated Microbial Genomes and Microbiomes system v7 (IMG/M) (2) for comparative analysis. The genome annotations can be interactively accessed on IMG/M (https://img.jgi.doe.gov/cgi-bin/m/main.cgi) using the following identifiers: 8036267134 (strain NE35), 8036277142 (strain NE40), 8045494135 (strain NE41); 8036272146 (strain NE43). </p> <p>This dataset is part of the following study:</p> <p>Marques M, da Silva DMG, Santos E, Baylina N, Peixoto R, Kyrpides NC, Woyke T, Whitman WB, Keller-Costa T, Costa R. 2024. Genome sequences of four novel <em>Endozoicomonas </em>strains associated with a tropical octocoral in a long-term aquarium facility. Microbiology Resource Announcements</p> <p> </p> <p>Other reference sources:</p> <p>(1) Huntemann M, Ivanova NN, Mavromatis K, James Tripp H, Paez-Espino D, Palaniappan K, Szeto E, Pillay M, Chen IMA, Pati A, Nielsen T, Markowitz VM, Kyrpides NC. 2015. The standard operating procedure of the DOE-JGI Microbial Genome Annotation Pipeline (MGAP v.4). Stand Genomic Sci 10:1–6.</p> <p>(2) Chen IMA, Chu K, Palaniappan K, Ratner A, Huang J, Huntemann M, Hajek P, Ritter SJ, Webb C, Wu D, Varghese NJ, Reddy TBK, Mukherjee S, Ovchinnikova G, Nolan M, Seshadri R, Roux S, Visel A, Woyke T, Eloe-Fadrosh EA, Kyrpides NC, Ivanova NN. 2023. The IMG/M data management and analysis system v.7: content updates and new features. Nucleic Acids Res 51:D723–D732.</p>
Annotation of eukaryotic genomes with Maker datasets
<p>Input and output datasets.</p>
Annotation of eukaryotic genomes with Maker datasets
<p>Input and output datasets.</p>
Annotation of eukaryotic genomes with Maker datasets
<p>Input and output datasets.</p>
Annotation of eukaryotic genomes with Maker datasets
<p>Input and output datasets.</p>
Genome annotation associated with the publication "Chromosome-level genome assembly of the Cape cliff lizard (Hemicordylus capensis)"
<p>Genome annotation associated with the publication "Chromosome-level genome assembly of the Cape cliff lizard (<em>Hemicordylus capensis</em>)"</p> <p>rHemCap1.1.gff3 - Genome annotation in GFF3 format<br> rHemCap1.1.proteins.fa - Multi-fasta file of protein coding genes<br> rHemCap1.1.cds-transcripts.fa - Multi-fasta file of transcripts (CDS)</p> <p> </p>
Genome assembly and annotation files of Colletotrichum graminicola strain TZ-3
<p>The maize anthracnose stalk rot and leaf blight diseases caused by the fungal pathogen Colletotrichum graminicola is emerging as an important threat to corn production worldwide. In this work, we provide an improved genome assembly of C. graminicola strain TZ-3 by using the PacBio Sequel II and Illumina high-throughput sequencing technologies. We hope that this genome information will facilitate the pathogenesis research of this plant pathogen.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.