Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
FC309 genome assembly and annotation files
Open the record for dataset details and reuse information.
Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024
Open the record for dataset details and reuse information.
Whole genome assembly and annotation of the King Angelfish (Holacanthus passer) gives insight into the evolution of marine fishes of the Tropical Eastern Pacific
Open the record for dataset details and reuse information.
Annotated genome of M. phaseolina MP2_2471002
<p>Annotated genome of <em>Macrophomina phaseolina </em>MP2_2471002</p>
Updated genome annotation for Physarum polycephalum
<p>This dataset contains a genome assembly (previously released by another group) for the slime mold <em>Physarum polycephalum</em>, along with our updated annotation file (see details below). The updated annotation reflects a significant enrichment in U12-type introns in <em>P. polycephalum </em>as described in this preprint: https://doi.org/10.1101/2020.10.12.336362</p> <p>We downloaded the <em>P. polycephalum</em> genome assembly and annotation from <a href="http://www.physarum-blast.ovgu.de/">http://www.physarum-blast.ovgu.de/</a>, and RNA-seq for <em>P. polycephalum</em> from NCBI’s SRA database (accession numbers DRR047256, ERR089824-ERR089827, and ERR557103-ERR557120). To reannotate the genome, we combined <em>de novo</em> and reference-based approaches. First, we generated a <em>de novo</em> transcriptome from the aggregate RNA-seq data using Trinity (Grabherr et al. 2011). We also separately mapped the reads to the genome using HISAT2 (Kim et al. 2019), allowing for non-canonical splice sites (--pen-noncansplice 0), followed by StringTie (M. Pertea et al. 2016) to incorporate the mapped reads with the existing annotations and generate additional putative transcript structures. Coding-sequence annotations for the assembled transcripts, informed by additional homology information from the SwissProt (UniProt Consortium 2008) protein database, were generated using TransDecoder (Brian J. Haas et al. 2013), and further refined with the de novo transcriptome via PASA (B. J. Haas 2003). In addition, an AUGUSTUS (Stanke et al. 2008) annotation was generated from the mapped reads using BRAKER1 (Hoff et al. 2015) explicitly allowing for AT-AC splice boundaries (--allow_hinted_splicesites=atac). Lastly, the AUGUSTUS- and StringTie-based gene predictions were merged using gffcompare (G. Pertea and Pertea 2020), and updated again using PASA. To gauge the quality of our annotations versus those previously available, we performed a BUSCO (Simão et al. 2015) analysis against conserved eukaryotic genes; the previous annotations contained matches to 60.1% of eukaryotic BUSCO groups (54.5% single-copy; 27.1% fragmented; 12.8% missing); our annotation increased this percentage to 73.3% (64.4% single-copy; 18.5% fragmented; 8.2% missing).</p>
Genome, annotations and SNPs for the green peafowl and associated scripts
<p><span><span>Both anthropogenic impacts and historical climate change could contribute to population decline and species extinction, but their relative importance has yet to be determined. Emerging approaches based on genomic, climatic and anthropogenic data provide a promising analytical framework to address this question. This study applied such an integrative approach to examine potential drivers for endangerment of the green peafowl (<i>Pavo muticus</i>). Several demographic reconstructions based on population genomes congruently retrieved a drastic population declination since the mid-Holocene. Furthermore, comparison between historical and modern genomes suggested genetic diversity decrease during the last 50 years. However, climate-based ecological niche models predicted general range stationarity during these periods and imply little impact of climate change. Further analyses suggested that human activity intensities were negatively correlated with the green peafowl's effective population sizes and significantly associated with its survival statuses (extirpation or persistence). Archaeological and historical records corroborate the critical role of humans, leaving the footprint of low genomic diversity and high inbreeding in the surviving populations. This study sheds light on the potential deep-time effects of human disturbance on species endangerment and on the whole, offers immediately a multi-evidential approach in examining underlying forces for population declines.</span></span></p>
Data from: Genome assembly and annotation of a Drosophila simulans strain from Madagascar
Drosophila simulans is a closely relative of the genetic model D. melanogaster. In an effort to improve the genomic resources for D. simulans, we assembled and annotated the genome of a strain originating from Madagascar (M252), the ancestral range of D. simulans. The comparison of the M252 genome to other available D. simulans assemblies confirmed its high quality, but also highlighted genomic regions that are difficult to assemble with NGS data. The annotation of M252 provides a clear improvement with alternative splicing for 52% of the multiple-exon genes, UTRs for 70% of the genes, 225 novel genes and 781 pseudogenes being reported. We anticipate that the M252 genome will be a valuable resource for many research questions.
Data from: Genome assembly and annotation of Arabidopsis halleri, a model for heavy metal hyperaccumulation and evolutionary ecology
The self-incompatible species Arabidopsis halleri is a close relative of the self-compatible model plant Arabidopsis thaliana. The broad European and Asian distribution and heavy metal hyperaccumulation ability make A. halleri a useful model for ecological genomics studies. We used long-insert mate-pair libraries to improve the genome assembly of the A. halleri ssp. gemmifera Tada mine genotype (W302) collected from a site with high contamination by heavy metals in Japan. After five rounds of forced selfing, heterozygosity was reduced to 0.04%, which facilitated subsequent genome assembly. Our assembly now covers 196 Mb or 78% of the estimated genome size and achieved scaffold N50 length of 712 kb. To validate assembly and annotation, we used synteny of A. halleri Tada mine with a previously published high-quality reference assembly of a closely related species, Arabidopsis lyrata. Further validation of the assembly quality comes from synteny and phylogenetic analysis of the HEAVY METAL ATPASE4 (HMA4) and METAL TOLERANCE PROTEIN1 (MTP1) regions using published sequences from European A. halleri for comparison. Three tandemly duplicated copies of HMA4, key gene involved in cadmium and zinc hyperaccumulation, were assembled on a single scaffold. The assembly will enhance the genomewide studies of A. halleri as well as the allopolyploid Arabidopsis kamchatica derived from A. lyrata and A. halleri.
Red clover (Trifolium pratense L.) genome assembly and annotation
<p>De Vega, J. J. <em>et al.</em> Red clover (<em>Trifolium pratense</em> L.) draft genome provides a platform for trait improvement. <em>Sci. Rep.</em> <strong>5</strong>, 17394; doi: 10.1038/srep17394 (2015).</p>
microbetag : building a thorough database of genome-scale KO annotations
<p>In this repository we keep internal data for the <em><a href="https://hariszaf.github.io/microbetag/">microbetag</a> </em>microbial co-occurrence network annotator.</p> <p><em>microbetag</em> makes use of 2-column files for each genome, indicating the KO term found and a KEGG module in which this terms takes part into. <br>As a single KO term might participates in more than one KEGG modules, the same KO might be more than once in an annotation file. </p> <table> <tbody> <tr> <td> <div>chem_xref.tar.gz</div> </td> <td> <p>The MNXref namespace</p> <ol> <li>The identifier of a chemical compound in an external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#XREF">XREF</a>]</li> <li>The corresponding identifier in the MNXref namespace [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#MNX_ID">MNX_ID</a>]</li> <li>The description given by the external resource [<a href="https://www.metanetx.org/mnxdoc/mnxref.html#STRING">STRING</a>]</li> </ol> <p>MNXref 4.0 release notes: - The third column (evidence tag for the mapping) was suppressed - The descriptions were completed - Deprecated identifiers were moved into they own table below</p> </td> </tr> <tr> <td> <div>gtdb_modelseed_gems.zip</div> </td> <td> <p>for all the GTDB genomes their corresponding <a href="https://patricbrc.org/">PATRIC</a> annotations were gathered. Then, using <a href="https://github.com/ModelSEED/ModelSEEDpy">modelseedpy</a> we constructed their genome scale metabolic reconstructions</p> </td> </tr> <tr> <td> <div>gtdb_kofam_scan_per_module.tar.gz</div> </td> <td> <p>all representative genomes of <a href="https://gtdb.ecogenomic.org/">GTDB</a> (v.202) were parsed and their corresponding `.faa` files were retrieved from the <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/">NCBI FTP</a>. Then the <a href="https://github.com/takaram/kofam_scan">kofam_scan</a> tool was used to annotate them and finally a <a href="https://github.com/hariszaf/microbetag/blob/clean/mappings/gtdb_mappings/gtdb_annotations_per_module.py">manual script </a>was used to keep KOs of each genome per module. </p> </td> </tr> <tr> <td>SeedSet.pkl.gz</td> <td> <p>A pickle file with the seeds of each GEM included in the <em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the <em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. Example:</p> <p>PATRIC SeedSet<br>373.172 [cpd00891, cpd00136, cpd00199, cpd01772, cpd00...<br>397278.5 [cpd00891, cpd00136, cpd01772, cpd02698, cpd08...</p> </td> </tr> <tr> <td>NonSeedSet.pkl.gz</td> <td> <p>A pickle file with the non seeds of each GEM included in the <em>gtdb_modelseed_gems.zip </em>file and related to the KEGG MODULES based on the <em>seedId_keggId_module.tsv </em>file you can find on microbetag's GitHub page. Example:</p> <p>PATRIC NonSeedSet<br>64187.548 [cpd00508, cpd00869, cpd00774, cpd03830, cpd00...<br>74426.1719 [cpd00204, cpd00447, cpd20171, cpd03470, cpd00...</p> </td> </tr> <tr> <td>seeds_per_genome.pkl.gz</td> <td> <p>A pickle file with a binary representation of the seeds per genome . Example:</p> <p> cpd00493 cpd00296 cpd11431 cpd00063 cpd15717 ... <br>2162051.4 1 0 0 1 0 0 0 0 0 0 ... </p> </td> </tr> <tr> <td> <div>nonseeds_per_genome.pkl.gz</div> </td> <td>Like above for non-seeds.</td> </tr> <tr> <td> <div>phen_classes.zip</div> </td> <td>A list of pickle files with the re-trained classes of phenDB for the prediction of functional traits on a genome.</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p>
Engelmann spruce Se404v1 - genome annotations (Maker)
<p>BTL lab at the GSC, Vancouver Canada</p><p>Curated and released by K.K. Gagalova</p><p> </p><p>Please cite "Spruce giga-genomes: structurally similar yet distinctive with differentially expanding gene families and rapidly evolving genes<a href="https://doi.org/10.1111/tpj.15889"><strong>" Gagalova et al. - https://doi.org/10.1111/tpj.15889</strong></a></p>
Sitka spruce Q903v1 - genome annotations (Maker)
<p>BTL lab at the GSC, Vancouver Canada</p><p>Curated and released by K.K. Gagalova</p><p> </p><p>Please cite "Spruce giga-genomes: structurally similar yet distinctive with differentially expanding gene families and rapidly evolving genes<a href="https://doi.org/10.1111/tpj.15889"><strong>" Gagalova et al. - https://doi.org/10.1111/tpj.15889</strong></a></p>
Data for "Visual Analytics for Enhancing Quality and trust in Genome-Wide Expression Clustering and Annotation"
<p><a href="https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/bb993a532f92298639b70bf7f2cd758d/raw/e2def8df7c20551e16f296c7c303308e5418c5e9/index.html">RasteredHPA23v06BloodReport (https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/bb993a532f92298639b70bf7f2cd758d/raw/e2def8df7c20551e16f296c7c303308e5418c5e9/index.html)</a></p> <p><a href="https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/087802569e028f96648dd5387bf4aba7/raw/4c9d65f96127ff9620509cbcf2b8d01435455705/index.html">RasteredHPA23v06brainReport (https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/087802569e028f96648dd5387bf4aba7/raw/4c9d65f96127ff9620509cbcf2b8d01435455705/index.html)</a></p> <p><a href="https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/bf5d7b90428a514bbdbe7e138863d04c/raw/6cb942e57d7a703e6cf78027813aa65e340b7c78/index.html">RasteredHPA23v06CellineReport (https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/bf5d7b90428a514bbdbe7e138863d04c/raw/6cb942e57d7a703e6cf78027813aa65e340b7c78/index.html)</a></p> <p><a href="https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/44315c0313299e9a8fe5d3ba3dac49fa/raw/200be056379260764b5c94064f38e7b8f33fbcc5/index.html">RasteredHPA23v06SinglecellReport (https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/44315c0313299e9a8fe5d3ba3dac49fa/raw/200be056379260764b5c94064f38e7b8f33fbcc5/index.html)</a></p> <p><a href="https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/d6ce47837daf019acc5e22baef379b79/raw/88600436126299a6a56984b3671f1e98630e6c19/index.html">RasteredHPA23v06TissueReport (https://rictjo.github.io/?https://gist.githubusercontent.com/rictjo/d6ce47837daf019acc5e22baef379b79/raw/88600436126299a6a56984b3671f1e98630e6c19/index.html)</a></p> <p>goto gist then check revision for report datestamps</p>
Three genomes of Pseudomonas syringae pv. actinidiae and their annotation files were used for comparative genomic studies
Open the record for dataset details and reuse information.
Dataset - Genome annotation with Maker
<p>These are the input and output files for the genome annotation workflow with Maker.</p>
Annotated Protein files of 120 Phytophthora genomes
<p>These are the annotated protein files of 120 Phytophthora genomes</p>
Single cell Iso-Sequencing enables rapid genome annotation for scRNAseq analysis
<p>Single <span>cell RNA sequencing (scRNAseq) is a powerful technique that continues to expand across various biological applications. However, incomplete 3' UTR annotations can impede single cell analysis resulting in genes that are partially or completely uncounted. Performing scRNAseq with incomplete 3' UTR annotations can hinder the identification of cell identities and gene expression patterns and lead to erroneous biological inferences. We demonstrate that performing single cell isoform sequencing (ScISOr-Seq) in tandem with scRNAseq can rapidly improve 3' UTR annotations. Using threespine stickleback fish (</span><em>Gasterosteus aculeatus</em><span>), we show that gene models resulting from a minimal embryonic ScISOr-Seq dataset retained 26.1% greater scRNAseq reads than gene models from Ensembl alone. Furthermore, pooling our ScISOr-Seq isoforms with a previously published adult bulk Iso-Seq dataset from stickleback, and merging the annotation with the Ensembl gene models, resulted in a marginal improvement (+0.8%) over the ScISOr-Seq only dataset. In addition, isoforms identified by ScISOr-Seq included thousands of new splicing variants. The improved gene models obtained using ScISOr-Seq lead to successful identification of cell types and increased the reads identified of many genes in our scRNAseq stickleback dataset. Our work illuminates ScISOr-Seq as a cost-effective and efficient mechanism to rapidly annotate genomes for scRNAseq.</span></p>
Aralia elata genome and annotation
<p>We report a chromosome-level genome of Aralia elata with a total length of 1.08 Gb and a contig N50 of 1.2 Mb using SMRT, HiC, and next-generation sequencing. The Aralia elata genome is the first chromosome-level genome of genus Aralia. This dataset also includes the annotation of Aralia elata genome and the gene expression profiles.</p>
Xcc BrA1 re-sequenced genome contigs, annotation file and Trinitiy De-nove assembled DEG's
<p>Assembled genome contigs of Xcc strain BrA1 (Behlau et al 2017) which was re-sequenced by our lab and used in a copper stress Transcriptome study. A hybrid RNA-seq analysis pipeline was used and the DEG Trinity De-novo assembled transcripts are also included.</p>
De novo assembly and annotation of parasitic trematode genomes
<p>Contained in this release are 19 genome assemblies and annotations of parasitic trematodes, encompassing 13 species. This included representatives of the <em>Schistosoma</em> (<em>n</em> = 13 assemblies), <em>Trichobilharzia</em> (<em>n</em> = 2 assemblies), <em>Heterobilharzia americana</em> (<em>n</em> = 2 assemblies) and <em>Dicrocoelium dendriticum </em>(<em>n </em>= 1 assembly). The <em>Schistosoma curassoni</em> assembly has been released previously (10.5281/zenodo.6594833) but a new annotation is included with the original assembly here. </p> <p>These genomes were assembled from a variety of sources including stored parasites from museum collections, established laboratory strains and wild-caught isolates sampled from natural hosts in endemic regions. Using a combination of DNA sequencing approaches, all genomes were assembled into chromosomal-scale scaffolds. This was followed by genome annotation based on short-read RNA sequencing (RNA-seq) and long-read isoform sequencing (Iso-seq) transcriptomic data.</p> <p>Included here are the primary assemblies (representing a non-redundant haploid genome) for each species (*.primary.fa), alternate loci (alternate representations of loci found in a largely haploid assembly; *.haplotypes.fa) and annotations (*.gff3). Metadata for each assembly can be found in the included spreadsheets (metadata.xlsx). </p> <p>This data is part of a pre-publication release. For information on the proper use of pre-publication data shared by the Wellcome Trust Sanger Institute (including details of any publication moratoria), please see https://www.sanger.ac.uk/about/research-policies/open-access-science/.</p> <p>This repository will be updated with a complete list of collaborators/authors prior to publication. Please contact Duncan Berger (db22@sanger.ac.uk) with questions regarding pre-publication use of this dataset. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.