Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
104
datasets available to search
ShareScore release 0.9.0
Dataset results
104 results for “functional annotation”
FURNA: a database of functional annotations of RNA structures (part 1)
<p>A copy of the FURNA database curated on June 9, 2024 (part 1). Download both part 1 (10.5281/zenodo.11664059) and part 2 (10.5281/zenodo.11672037) to combine them by:</p> <p>$ cat xaa xab > furna.tar.bz2</p> <p>$ rm xaa xab</p> <p>$ tar -xvf furna.tar.bz2</p>
Taxonomic and functional annotations of transcripts and proteins derived from a Metatranscriptomic study of microbial eukaryotes from Lake Pavin
<p>These data were obtain as part of a metatranscriptomic study (Monjot <em>et al.,</em> 2023, 2024). All scripts to obtain these annotations are available at https://github.com/amonjot/SSN_Monjot_2024. The sequencing data (i.e. metatranscriptomic) used to obtain this taxonomic and functional information are archived at ENA under accession number PRJEB61515.</p> <p>This repository also contains various protein sequence similarity networks (Lagoon_output.zip). As these files are very time-consuming to produce, we have provided them to complete all the steps in Monjot <em>et al,</em> 2024. All procedures to produce them are present on the following github repository : https://github.com/amonjot/SSN_Monjot_2024.</p> <p> </p>
AnnoTree: Functionally Annotated Tree of Life
<p><strong>Current version:</strong><br> 2019-02-01: This file is a MySQL (v5.7) dump for AnnoTree, a functionally annotated tree of life. It includes the databases 'gtdb_bacteria_RS86' and 'gtdb_archaea_RS86'. Both databases include <a href="https://pfam.xfam.org/">Pfam</a> (v27), <a href="https://www.genome.jp/kegg/">KEGG</a> Ontology (from the <a href="https://www.uniprot.org/help/uniref">UniRef100</a> database downloaded March 6, 2018), and <a href="http://tigrfams.jcvi.org/cgi-bin/index.cgi">TIGRFAM</a> (v15) annotations for the representative genomes in <a href="http://gtdb.ecogenomic.org/">GTDB</a> release 03-RS86.</p> <p>Please refer to <a href="http://annotree.uwaterloo.ca">http://annotree.uwaterloo.ca</a> for the full site.</p> <p> </p> <p><strong>Older versions:</strong><br> 2018-10-22: MySQL dump file (v5.7) containing the 'gtdb_bacteria' database for GTDB release 02-RS83 for use with AnnoTree. Functional annotations include <a href="https://pfam.xfam.org/">Pfam</a> (v27) and <a href="https://www.genome.jp/kegg/">KEGG</a> Ontology (from the <a href="https://www.uniprot.org/help/uniref">UniRef100</a> database downloaded March 6, 2018). Corresponds to AnnoTree version 1.0.0 before the integration of Archaea.</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain JAMR
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis acuminata strain DAVA01
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain MBL-DK2009
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Teleaulax amphioxeia
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis ovum strain DoSS3195
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
<p class="MsoNormal">Whole genome sequencing enables us to ask fundamental questions about the genetic basis of adaptation, population structure, and epigenetic mechanisms, but usually requires a suitable reference genome for making sense of the sequence data. While the availability of reference genomes has significantly improvement in both taxonomic coverage and overall quality, this poses a challenge for researchers in determining which reference genome best suits their data. Here we compare the use of two different reference genomes for the three-spined stickleback (<em>Gasterosteus aculeatus</em>), one novel genome from a European individual and the published reference genome of a North American individual. Specifically, we investigate the impact of using a local reference versus one generated from a differentiated population on several commonly used metrics in population genomics. Through mapping genome resequencing data of 60 sticklebacks from across Europe and North America, we confirmed genome quality is an important factor in choosing a reference genome. A local reference genome did offer increased mapping efficiency and genotyping accuracy, likely stemming from the higher similarity in genome sequence and synteny. Despite comparable distributions of the metrics generated across the genome using SNP data (i.e., π, Tajima's D, and FST), window-based statistics using different references resulted in different outlier genes and enriched gene functions. In contrast, the marker-based analysis utilising DNA methylation distributions had a considerably higher overlap in outlier genes and functions when using different reference genomes. Overall, our results highlight how using a local reference genome can increase the resolution of genome scans when multiple similar-quality reference genomes are available. Such results have implications in the detection of signatures of selection.</p>
Metabolomics dataset relating to the pubblication: "Combining CRISPRi and metabolomics for functional annotation of compound libraries"
<p>Metabolomics dataset relating to the pubblication: "Combining CRISPRi and metabolomics for functional annotation of compound libraries"</p>
Dataset - Functional annotation for protein sequences
<p>These are the input and output files for the functional annotation of protein sequences workflow tests.</p>
Supplementary datasets for eRNA community and Functional annotations
Open the record for dataset details and reuse information.
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
Open the record for dataset details and reuse information.
Chromosome-level assembly of two pearl millet (Cenchrus americanus) genomes, functional annotation and transcriptomes
Open the record for dataset details and reuse information.
Database for mi-faser: Functional sequencing read annotation for high precision microbiome analysis
<p><strong>[Database for mi-faser]</strong></p> <p><strong>mi-faser: </strong><em>microbiome - functional annotation of sequencing reads</em></p> <p>A super-fast ( < 20min/10GB of reads ) and accurate ( > 90% precision ) method for annotation of molecular functionality encoded in sequencing read data without the need for assembly or gene finding.</p> <p>Web Service: http://services.bromberglab.org/mifaser/|<br> Repository: https://bitbucket.org/bromberglab/mifaser_base/</p>
Taxonomic and functional annotations of the Integrated non-redundant Gene Catalog 9.9
<p>The Integrated non-redundant Gene Catalog (<strong>IGC</strong>) 9.9 is a database of 9.9 million genes from 1267 individual fecal samples together with the Homo sapiens database (MetaHIT project, grant agreement 201052). This repository contains the taxonomic and functional annotation of the IGC database.</p> <p><strong>full_taxonomy_MetaHIT99.tsv </strong>: Taxonomic assignment of proteins from IGC database with the sequence aligner DIAMOND against the non-redundant NCBI database, with an e-value threshold of 10<sup>-4</sup> </p> <p><strong>KEGG89_IGC_hs99.table</strong> : Functional annotation of proteins from IGC database with KEGG resource with an e-value threshold of 10<sup>-5</sup>, a bit-score threshold of 60 and using the sensitive mode of DIAMOND</p>
Functional annotation
Open the record for dataset details and reuse information.
DRAM raw annotations for "Cover Crop Root Exudates Impact Soil Microbiome Functional Trajectories in Agricultural Soils" Seitz et al 2024
<p>Additional File 5: <span>Raw DRAM MAG annotations. </span></p>
The gene structure annotation, gene function annotation and TE annatition files of the Glyphodes pyloalis's genome
Open the record for dataset details and reuse information.
The gene structure annotation, gene function annotation and TE annatition files for the Cibotium barometz isolate CiBa-2024 genome
<p>This dataset comprises comprehensive annotation files for the genome of Cibotium barometz (Golden Chicken Fern), isolate CiBa-2024. It includes gene structure predictions, functional annotations, and transposable element (TE) identifications, complementing the chromosome-level genome assembly. The gene structure annotation provides detailed information on predicted gene models, including exon-intron boundaries and coding sequences. Functional annotations offer insights into the potential roles of identified genes, including Gene Ontology (GO) terms, protein domains, and pathway associations. The TE annotation file details the classification and distribution of transposable elements within the genome. These annotations were generated using state-of-the-art bioinformatics tools and databases, offering a valuable resource for researchers studying fern genomics, plant evolution, and the genetic basis of C. barometz's unique biological features, including its medicinal properties. This dataset aims to facilitate further research in comparative genomics, functional studies, and the exploration of fern biology and evolution.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.