Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
505
datasets available to search
ShareScore release 0.9.0
Dataset results
505 results for “Complete genomes”
Figure 1 in The complete mitochondrial genome of Parnassius actius (Lepidoptera: Papilionidae: Parnassinae) with the related phylogenetic analysis
Figure 1. Circular map of the mitochondrial genome of P. actius. The abbreviations for the genes are as follows: COI, COII and COIII refer to the cytochrome oxidase subunits; Cytb refers to cytochrome B; ATP6 and ATP8 refer to subunits 6 and 8 of ATPase; ND1–6 grefers to components of NADH dehydrogenase. The tRNAs are indicated by the IUPAC-IUB single letter amion acid codes, while L1, L2, S1, S2 denote tRNALeu(CUN), tRNALeu(UUR), tRNASer(AGN) and tRNASer(UCN), respectively. Gene names that are not underlined indicate the direction of transcription from left to right, and with underline indicates right to left. The P. actius mitogenome was sequenced by using 7 short fragments (SF1–SF7) and 7 long fragments (LF1–LF7) as templates, shown as single lines within a circle.
Figure 2 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 2. Relative Synonyous Codon Usage (RSCU) in L. melli, H. cunea, and A. formosae mitogenomes. Codons are provided on the x-axis.
Figure 6 in The complete mitochondrial genome of Lemyra melli (Daniel) (Lepidoptera: Erebidae) and a comparative analysis within the Noctuoidea
Figure 6. Phylogenetic analysis inferred from the concatenated nucleotide sequences of 13 PCGs in the mitogenome. D. melanogaster and E. regina were used as outgroups. The numbers above the branches specify bootstrap percentages (1 000 replicates) from software RAXMAL and the ML method. A, B, C, D, and E indicated the superfamily of Noctuoidea, Bombycoidea, Geometroidea, Pyraloidea and Tortricidea, respectively.
Figure 7 in The complete mitochondrial genome of Parnassius actius (Lepidoptera: Papilionidae: Parnassinae) with the related phylogenetic analysis
Figure 7. Bayesian Inference and maximum likelihood phylograms of the 45 Parnassius species in this study (Numbers on each node correspond to the posterior probability values of the BI analysis (left) and the ML bootstrap percentage values for 1000 replicates (right), posterior probability values below 0.5 or bootstrap percentage values below 50% was not shown on the diagram).
Figure 5 in General methods to obtain and analyze the complete mitochondrial genome of aphid species: Eriosoma lanigerum (Hemiptera: Aphididae) as an example
Figure 5. Bayesian inference (BI) and Maximum likelihood (ML) phylogenetic tree inferred from 11 aphids mt genome sequences. The support values on the nodes are the bootstrap (BS) values and the Bayesian posterior probabilities (BPP).
Figure 4 in General methods to obtain and analyze the complete mitochondrial genome of aphid species: Eriosoma lanigerum (Hemiptera: Aphididae) as an example
Figure 4. The secondary structure of the 22 transfer RNAs (tRNAs) in the Eriosoma lanigerum mt genome.
Figure 3 in General methods to obtain and analyze the complete mitochondrial genome of aphid species: Eriosoma lanigerum (Hemiptera: Aphididae) as an example
Figure 3. Circular map of the Eriosoma lanigerummt genome. Gene names not underlined indicate the direction of transcription in
BLASTDB NCBI GeneBank viruses complete genomes 2023/01/01
<p>Database: NCBI GeneBank Viruses Complete genomes 2023_01_01<br> 86,641 sequences; 2,391,589,518 total bases<br> Date: Jan 1, 2023 4:32 PM<br> BLASTDB Version: 5</p> <p> </p>
Data from: Refinement of the Antarctic fur seal (Arctocephalus gazella) reference genome increases continuity and completeness
Open the record for dataset details and reuse information.
Raw Nanopore data for "Nanopore Long-Read Guided Complete Genome Assembly of Hydrogenophaga intermedia, and Genomic Insights into 4-Aminobenzenesulfonate, p-Aminobenzoic Acid and Hydrogen Metabolism in the Genus Hydrogenophaga"
<p>This is the raw Nanopore dataset (fast5) for Hydrogenophaga intermedia PBC. The gDNA was prepared using the now obsolete SQK-NSK007 kit and sequenced on a MINION R9 Flowcell. </p>
The complete genome sequence and comparative genomic analyses of four phages (NJ-P3, NB-P21, NC-P34 and NN-P42)
<p>We downloaded and reanalyzed the raw data of four phages genomes (NJ-P3, NB-P21, NC-P34, NN-P42). This is the reassembled whole genomes and comparative genomic analyses of four phages.</p>
The complete mitochondrial genome of dwarf form of Sthenoteuthis oualaniensis (Cephalopoda: Ommastrephidae) from the South China Sea
<p>The purpleback flying squid (<em>Sthenoteuthis oualaniensis</em>) is a pelagic squid with tremendous potential for commercial exploitation. <em>Sthenoteuthis oualaniensis</em> comprises two forms in the South China Sea, dwarf form and medium-sized form. In this study, we described the complete mitochondrial genome of dwarf form of <em>Sthenoteuthis oualaniensis.</em> The genome is 20320 bp in length, encoding the standard set of 13 protein-coding genes, 20 tRNA genes and 2 rRNA genes, with circular organization. The overall base composition of the whole mitochondrial genome was A (37.23%), T (32.78%), G (10.53%) and C (19.46%) with an AT bias of 70.01%. The longest protein-coding genes of these species was <em>ND5</em>, whereas the shortest <em>ATP8</em>.</p>
OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations.
<p>Genes predicted by OrthoFiller not found in public genome annotations for the species analysed in the paper titled "OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations."</p>
Data from: high repeat content in the genomes of sparrows: the importance of genome assembly completeness for transposable element discovery
<p>Transposable elements (TE) play critical roles in shaping genome evolution. However, the highly repetitive sequence content of TEs is a major source of assembly gaps. This makes it difficult to decipher the impact of these elements on the dynamics of genome evolution. The increased capacity of long-read sequencing technologies to span highly repetitive regions of the genome should provide novel insights into patterns of TE diversity. Here we report the generation of highly contiguous reference genomes using PacBio long read and Omni-C technologies for three species of sparrows in the family Passerellidae. To assess the influence of sequencing technology on TE annotation, we compared these assemblies to three chromosome-level sparrow assemblies recently generated by the Vertebrate Genomes Project and nine other sparrow species generated using a variety of short- and long-read technologies. All long-read based assemblies were longer in length (range: 1.12-1.41 Gb) than short-read assemblies (0.91-1.08 Gb). Assembly length was strongly correlated with the amount of repeat content, with longer genomes showing much higher levels of repeat content than typically reported for the avian order Passeriformes. Repeat content for the Bell's sparrow (31.2% of genome) was the highest level reported to date for a songbird genome assembly and was more in line with woodpecker (order Piciformes) genomes. CR1 LINE elements retained from an expansion that occurred 25-30 million years ago were the most abundant TEs in the song sparrow genome. Although the other five sparrow species also exhibit evidence for a spike in CR1 LINE activity at 25-30 million years ago, LTR elements stemming from more recent expansions were the most abundant elements in these species. LTRs were uniquely abundant in the Bell's sparrow genome deriving from two recent peaks of activity. Higher levels of repeat content (79.2-93.7%) were found on the W chromosome relative to the Z (20.7-26.5) or autosomes (16.1-30.9%). These patterns support a dynamic model of transposable element expansion and contraction underpinning the seemingly constrained and small sized genomes of birds. Our work highlights how the resolution of difficult-to-assemble regions of the genome with new sequencing technologies promises to transform our understanding of avian genome evolution.</p>
The complete chloroplast genome of Mimusops elengi (Sapotaceae)
<p><span>The first complete chloroplast genome sequences of <i><span>Mimusops elengi</span></i> (Sapotaceae) were reported in this study. The cpDNA of <i><span>M</span></i><i><span>.</span></i><i><span> elengi</span></i> is 159,719 bp in length, contains a large single-copy region (LSC) of 88,935 bp and a small single-copy region (SSC) of 18,606 bp, which were separated by a pair of inverted repeat (IR) regions of 26,089 bp. The genome contains 132 genes, including 87 protein-coding genes, 8 ribosomal RNA genes, and 37 transfer RNA genes. The overall GC content of the whole genome is 36.8%. Phylogenetic analysis of 12 chloroplast genomes within the family Sapotaceae suggests that the sister relationship of <i><span>Autranella</span></i> and <i><span>Tieghemella</span></i><span> is strongly supported. </span><i><span>Minusops</span></i> genus is close to <i><span>Autranella</span></i> and <i><span>Tieghemella</span></i>, although the support value is still low.</span></p>
Complete mitochondrial genome sequence of the Atlantic Mudskipper (Periophthalmus barbarus) (Linnaeus, 1766) (Perciformes: Gobiidae)
<p>The complete mitochondrial genome of the Atlantic mudskipper (<em>Periophthalmus barbarus</em>) was determined in this study. The specimen was collected from Abonnema, Nigeria (4.73075, 6.77565). The complete mitogenome sequence of <em>P. barbarus</em> would be useful for further studies on molecular phylogenetic relationship and population genetics of the subfamily Oxudercinae.</p>
HiFi Metagenomic Sequencing Enables Assembly of Accurate and Complete Genomes from Human Gut Microbiota.
<p>We reported 102 complete metagenome assembled genomes (cMAGs) from five human fecal HiFi sequencing samples.</p> <p>102_cMAGs_fna.tar.gz: Fasta sequence files of 102 cMAGs.</p> <p>gc_skew_figures.tar.gz: GC-skew pattern figures of 102 cMAGs. (SVG format)</p> <p>coverage_plots.tar.gz: Genome coverage plot of 102 cMAGs.</p>
Supplementary Materials from the article Characterization and molecular evolution analysis of Periploca forrestii inferred from its complete chloroplast genome sequence
<p>Table S1. Base composition of chloroplast genome in <em>P. forrestii</em>, Table S2. The lengths of introns and exons for the splitting genes, Table S3. The GC content of the codons from <em>P. forrestii </em>chloroplast genome, Table S4. Preferred codons in chloroplast genome of <em>P. forrestii</em>, Table S5. Long repeat sequences in the <em>P. forrestii </em>chloroplast genome, Figure S1. Codon bias analysis of P. forrestii chloroplast genome. (A) Neutrality plot analysis; (B) Analysis of PR2 bias plot; (C) Analysis on ENC and GC3 relationship.</p>
Complete telomere-to-telomere genomes uncover virulence evolution conferred by chromosome fusion in oomycete plant pathogens
<p><span>Variations in chromosome number are occasionally observed among oomycetes, a group that includes many plant pathogens, but the emergence of such variations and their effects on genome and virulence evolution remain ambiguous. We generated complete telomere-to-telomere genome assemblies for <em>Phytophthora sojae</em>, <em>Globisporangium ultimum</em>, <em>Pythium oligandrum</em>, and <em>G. spinosum</em>. Reconstructing the karyotype of the most recent common ancestor in Peronosporales revealed that frequent chromosome fusion and fission drove changes in chromosome number. Centromeres enriched with <em>Copia</em>-like transposons may contribute to chromosome fusion and fission events. Chromosome fusion facilitated the emergence of pathogenicity genes and their adaptive evolution. Effectors tended to duplicate in the sub-telomere regions of fused chromosomes, which exhibited evolutionary features distinct to the non-fused chromosomes. By integrating ancestral genomic dynamics and structural predictions, we have identified secreted Ankyrin repeat-containing proteins (ANKs) as a novel class of effectors in <em>P. sojae</em>. Phylogenetic analysis and experiments further revealed that ANK is a specifically expanded effector family in oomycetes. These results revealed chromosome dynamics in oomycete plant pathogens, and provided novel insights into karyotype and effector evolution.</span></p>
Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan
<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.