Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.7.1
Dataset results
345 results for “genome annotation”
179 high quality metagenome-assembled genomes sequences and annotations
<p>We analyzed seven sediment samples collected adjacent to ferromanganese nodules from the Clarion–Clipperton Fracture Zone (CCFZ) in the eastern Pacific Ocean. Through deep metagenomic sequencing, assembly, and binning, we reconstructed 179 high quality metagenome-assembled genomes (MAGs). This archive contains these genomes sequences and annotations. </p>
Genome annotation of Hibiscus trionum
<p><strong>Genome annotation of <em>Hibiscus trionum</em></strong></p> <p>Author information<br> Shizuka Koshimizu<sup>1,2</sup>, Sachiko Masuda<sup>3</sup>, Arisa Shibata<sup>3</sup>, Takayoshi Ishii<sup>4</sup>, Ken Shirasu<sup>3</sup>, Atsushi Hoshino<sup>5,6</sup>, Masanori Arita<sup>1,2</sup></p> <p>1 Bioinformation and DDBJ Center, National Institute of Genetics, Mishima 411-8540, Japan<br> 2 Graduate Institute for Advanced Studies, SOKENDAI, Mishima 411-8540, Japan<br> 3 Center for Sustainable Resource Science, RIKEN, Yokohama 230-0045, Japan<br> 4 Arid Land Research Center, Tottori University, Tottori, 680-001, Japan<br> 5 National Institute for Basic Biology, Okazaki 444-8585, Japan<br> 6 Graduate Institute for Advanced Studies, SOKENDAI, Okazaki 444-8585, Japan</p> <p><br> Contents<br> <strong>htrionum_v1.0_genome.fasta</strong>: Genome sequence<br> <strong>htrionum_v1.0_transcripts.gff3</strong>: Gene annotation<br> <strong>htrionum_v1.0_transcripts.gtf</strong>: Gene annotation<br> <strong>htrionum_v1.0_transcripts.fna</strong>: Transcript sequences<br> <strong>htrionum_v1.0_pep.faa</strong>: Protein sequences<br> <strong>Description.txt</strong>: Functional descriptions<br> <strong>InterProScan.txt</strong>: The results of InterProScan for all proteins</p> <p> </p> <p>Reference</p> <p>Genome and transcriptome analyses reveal genes involved in the formation of fine ridges on petal epidermal cells in <em>Hibiscus trionum</em></p> <p><em>DNA Research</em>, 2023;, dsad019, <a href="https://doi.org/10.1093/dnares/dsad019">https://doi.org/10.1093/dnares/dsad019</a></p>
Serratia fonticola EBS19 Whole genome sequence data fasta file annotated
<p>Whole genome sequence data of <em>Serratia fonticola</em> <strong>EBS19 </strong>strain.</p>
Danaus plexippus genome annotation and orthofinder results
<p>Funannotate v1.5.3 docker image was used to train Augustus v3.2.3, predict gene models, and perform functional annotation. As input for optimizing the performance of Augustus v3.2.3, funannotate used 2,404 PASA v2.3.3 gene models. To obtain this training gene model set, transcripts were de novo assembled with Trinity v2018-2.8.3 under settings --SS_lib_type RF, using all poly(A) RNA-seq paired reads after adapter removal with Trimmomatic v0.32. These transcripts were aligned to the genome under PASA using BLAT v36, obtaining a first set of gene models. The 500 longest non-redundant ORFs associated with the PASA gene models were used to train TransDecoder v5.2.0. Then the gene models were selected according to their abundance as estimated by Kallisto v0.44.0 under settings --rf-stranded using the Trinity normalized reads. Ultimately, BRAKER v2.0.3b trained Augustus with the retained gene models.<br>For gene prediction, funannotate aligned mRNAs and proteins from the previous annotation (official gene set 2, OGS2) with minimap v2.14-r883 under settings -ax splice --cs -u b -G 3000, and Diamond blastx v0.8.22, respectively. Protein alignments were further refined by funannotate, including 3 kb upstream and downstream of the region of alignment, and subsequently executing Exonerate v2.4.0. Additionally, funannotate parsed the introns supported by alignments of poly(A) RNA-seq reads generated with HISAT v2.1 under settings --rna-strandness RF --max-intronlen 10,000. This combination of hints (protein alignments, transcript alignments, and intron locations) was used by Augustus to predict a second set of 16,756 gene models. Of them, 9,695 were dubbed as highly supported, i.e. had more than 90% of their model supported either by intron hints, transcript alignments, or protein alignments. GeneMark-ET v4.35, under settings --max_intron 3,000 --soft_mask 2,000, was also run independently to predict a third set of gene models but only relying on intron hints.<br>The PASA, Augustus highly supported, Augustus not highly supported, and GeneMark prediction sets were combined by EVidenceModeler, assigning them 10, 5, 1, and 1 relative weights, respectively. The predictions were further filtered by removing genes shorter than 50 aa in length, or that had high sequence similarity (diamond blastp --sensitive --evalue 1e-10) to the repeat database included in funannotate, or that had more than 90% of the model intersecting regions masked by RepeatMasker. The filtered set of gene models was updated in order to include UTR information by two executions of the PASA annotation comparison using the Trinity transcripts and filtering gene models according to transcripts per million as calculated by Kallisto. Alternative transcripts were only kept if they were at least 10% as highly expressed as the most highly expressed transcript per gene.<br>Non-coding genes were annotated with the following tools: tRNA genes, tRNAscan-SE v.2.0; rRNA genes, RNAmmer v.1.2; and for a variety of other RNA genes, Infernal v1.1.1. Specifically, for miRNA-encoding genes, we used BLASTn to locate the most recent annotation of these genes. Lastly, FEELnc classified lncRNAs from the transcripts assembled by StringTie v1.3.2d, and considering protein-coding predictions described above.</p><p>Homology relationships of the predicted protein coding genes were inferred with OrthoFinder (v2.2.6) relative to <i>Bombyx mori, Plutella xylostella, Papilio xuthus, Amyelois transitella, Heliconius melpomene, Plodia interpunctella </i>and<i> Drosophila melanogaster</i>. OrthoFinder was run with the settings "-S diamond -M msa".</p>
A high-quality genome assembly and annotation of the gray mangrove, Avicennia marina
Open the record for dataset details and reuse information.
Genome assembly and annotation of the Dark-branded Bushbrown butterfly Mycalesis mineus (Nymphalidae: Satyrinae)
Open the record for dataset details and reuse information.
Assembly and annotation of eleven Salix (shrub willow) genomes
Open the record for dataset details and reuse information.
Re-annotated genomes for plastomes and mitogenomes
Open the record for dataset details and reuse information.
Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis
Open the record for dataset details and reuse information.
A high-quality genome assembly and annotation of the dark-eyed junco Junco hyemalis, a recently diversified songbird
Open the record for dataset details and reuse information.
Costus pulverulentus genome annotations
Open the record for dataset details and reuse information.
Annotation of genome of brown-marbled grouper (Epinephelus fuscoguttatus)
Open the record for dataset details and reuse information.
A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections
Open the record for dataset details and reuse information.
Data from: An annotated draft genome of the mountain hare (Lepus timidus)
Open the record for dataset details and reuse information.
Reference genome and annotation for Teleopsis dalmanni
Open the record for dataset details and reuse information.
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
Open the record for dataset details and reuse information.
Chromosome-level assembly of two pearl millet (Cenchrus americanus) genomes, functional annotation and transcriptomes
Open the record for dataset details and reuse information.
Annotated genome assemblies for Geoscapheus dilatatus, Panesthia cribrata and Neogeoscapheus hanni
Open the record for dataset details and reuse information.
Deep-learning-based annotation of 230 superasterid genomes reveals a harmonized dataset of 91,366 NLRs
Open the record for dataset details and reuse information.
Genome sequence assembly and annotation of MATA and MATB strains of <em>Yarrowia lipolytica</em>
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.