Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
669
datasets available to search
ShareScore release 0.9.0
Dataset results
669 results for “comparative genomics”
Data from: Genomic differentiation during speciation-with-gene-flow: comparing geographic and host-related variation in divergent life history adaptation in Rhagoletis pomonella
Open the record for dataset details and reuse information.
Data for: Endophyte genomes support greater metabolic gene cluster diversity compared with non-endophytes in Trichoderma
Open the record for dataset details and reuse information.
Data from: Remarkably conserved plastid genomes of Quercus Group Cerris in China: comparative and phylogenetic analyses
Open the record for dataset details and reuse information.
Data from: Comparative genomic analysis of nine Sphingobium strains: insights into their evolution and Hexachlorocyclohexane (HCH) degradation pathways
Open the record for dataset details and reuse information.
Data from: Comparative genomics of color morphs In the coral Montastraea cavernosa
Open the record for dataset details and reuse information.
Data from: Comparative phylogeographic inference with genome-wide data from aggregated population-pairs
Open the record for dataset details and reuse information.
Data from: An initial comparative genomic autopsy of wasting disease in sea stars
Open the record for dataset details and reuse information.
FIGURE 4 in Comparative mitochondrial genomics of Shoveliteratura triangula (Orthoptera Tettigoniidae, Meconematinae) and the first description of a female specimen
FIGURE 4. Phylogenetic reconstruction of the Meconematinae using mitochondrial PCGs and rRNA of the concatenated dataset. Applicable posterior probability values are shown. Numbers in the ML tree represent SH-aLRT support/ultrafast bootstrap support values.
FIGURE 2 in Comparative mitochondrial genomics of Shoveliteratura triangula (Orthoptera Tettigoniidae, Meconematinae) and the first description of a female specimen
FIGURE 2. Relative synonymous codon usage of Shoveliteratura triangula mitochondrial protein-coding genes. Codon families are shown on the x-axis.
Data from: The first complete mitochondrial genome of the Indian Tent Turtle, Pangshura tentoria (Testudines: Geoemydidae): characterization and comparative analysis
Characterization of complete mitogenome is a widely used genomics study for species delineation and evolutionary research. However, the sequences and structural motifs contained within the mitogenome have been rarely examined to understand the phylogeny and evolutionary history among Testudines. Hence, the mitogenomic features of several Testudines taxa are still anonymous to the scientific communities. The present study decodes the first complete mitochondrial genome of the Indian Tent Turtle, Pangshura tentoria (16,657 bp) by using next-generation sequencing. This denovo assembly encodes 37 genes: 13 protein coding genes (PCGs), 22 transfer RNA (tRNAs), two ribosomal RNA (rRNAs), and one control region (CR). The mitogenome contained 19 intergenic spacer and six overlapping regions. Most of the genes were encoded on majority strand, except for one PCG (NADH dehydrogenase subunit 6) and eight tRNAs. Most of the PCGs were started with an ATG initiation codon, except for cytochrome oxidase subunit 1 with 'GTG' and NADH dehydrogenase subunit 5 with 'ATA'. The termination codons, 'TAA' and 'AGA' were observed in NADH dehydrogenase subunit 4l and NADH dehydrogenase subunit 6 respectively. The Relative Synonymous Codon Usage analysis revealed the maximum abundance of Alanine, Isoleucine, Leucine, and Threonine. The non-synonymous/synonymous ratios were <1 in all PCGs, which indicates strong negative selection among all Geoemydid species. The study also found the typical cloverleaf secondary structure in most of the tRNA genes, except for Serine (trnS1) with lack of the conventional DHU arm. The Wobble base pairing was observed in the different stems (DHU, acceptor, and anticodon) of 11 tRNAs. The comparative study of Geoemydid mitogenomes revealed the occurrence of tandem repeats was frequent in the 3´ end of CR. Further, two copies of a unique tandem repeat 'TTCTCTTT' were identified in P. tentoria. The Bayesian and Maximum Likelihood phylogenetic trees using concatenation of 13 PCGs revealed the close relationships of P. tentoria with Batagur trivittata in the studied dataset. All the Geoemydid species showed distinct clustering with high bootstrap support congruent with previous evolutionary hypotheses. We suggest that the generations of more mitogenomes of Geoemydid species, especially for Batagurinae subfamily, are required to improve our understanding their in-depth phylogenetic and evolutionary relationships.
FIGURE 8 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 8. Tree showing the phylogenetic relationships among Lepidoptera species, constructed using Maximum Likelihood method. Bootstrap values (1000 repetitions) of the branches are indicated. Drosophila incompta (NC_025936) and Anopheles gambiae (NC_002084) were used as outgroups.
FIGURE 7 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 7. Features present in the A+T-rich region of Smerinthus planus. The ATATG motif is yellow shaded. The poly-T stretch is dotted underlined while the poly-A stretch is underlined. The single microsatellite T/A repeat sequence is indicated by gray shaded.
FIGURE 5 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 5. Putative secondary structures of the 22 tRNA genes of the Smerinthus planus mitochondrial genome.
FIGURE 6 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 6. Sequence alignment of PCG and tRNA regions from 27 different Lepidoptera species. 6A, The color sequence part is motif AAGATAGAAACCAACCTGGCTYACACCGGTTTGAACTCAGATCATGTAAG; 6B, The color sequence part is motif GAAGAATGAACTAAAGCAGAAACWGGAGTWGGAGCWGCTATAGCWGCWGG; 6C, The color sequence part is motif AAYCCWGAAACTAATTCTCTTTCHCCTTCAGCAAAATCAAAAGGAGTWCG.
FIGURE 4. A in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 4. A. The Relative Synonymous Codon Usage (RSCU) of the mitochondrial genome of 31 species in the Lepidoptera. Codon families are plotted on the X axis. B. The Relative Synonymous Codon Usage (RSCU) of the mitochondrial genome of S. planus in the Lepidoptera. Codon families are plotted on the X axis.
FIGURE 2 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 2. Comparison of the codon usage patterns within the mitochondrial genome of different Lepidoptera species.
FIGURE 3 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 3. Codon distribution patterns in various Lepidoptera species. The y-coordinate is the proportion of codons per 100 codons.
FIGURE 1 in The mitochondrial genome of Smerinthus planus (Lepidoptera: Sphingidae) and its comparative analysis with other Lepidoptera species
FIGURE 1. Map of the mitogenome of S. planus. The tRNA genes are labeled according to the IUPAC-IUB single-letter amino acids: cox1, cox2 and cox3 refer to the cytochrome c oxidase subunits; cob refers to cytochrome b; nad1-nad6 refer to NADH dehydrogenase components; rrnL and rrnS refer to ribosomal RNAs. Gene named above the bar are located on major strand, while the others are located on minor strand.
Data from: Comparative genomics of pathogenic and nonpathogenic beetle-vectored fungi in the genus Geosmithia
Geosmithia morbida is an emerging fungal pathogen which serves as a model for examining the evolutionary processes behind pathogenicity because it is one of two known pathogens within a genus of mostly saprophytic, beetle-associated, fungi. This pathogen causes thousand cankers disease in black walnut trees and is vectored into the host via the walnut twig beetle. G. morbida was first detected in western US and currently threatens the timber industry concentrated in eastern US. We sequenced the genomes of G. morbida in a previous study and two non-pathogenic Geosmithia species in this work and compared these species to other fungal pathogens and nonpathogens to identify genes under positive selection in Geosmithia morbida that may be associated with pathogenicity. Geosmithia morbida possesses one of the smallest genomes among the fungal species observed in this study, and one of the smallest fungal pathogen genomes to date. The enzymatic profile in this pathogen is very similar to its non-pathogenic relatives. Our findings indicate that genome reduction or retention of a smaller genome may be an important adaptative force during the evolution of a specialized lifestyle in fungal species that occupy a specific niche, such as beetle vectored tree pathogens. We also present potential genes under selection in G. morbida that could be important for adaptation to a pathogenic lifestyle.
Data from: Multi-DICE: R package for comparative population genomic inference under hierarchical co-demographic models of independent single-population size changes
Population genetic data from multiple taxa can address comparative phylogeographic questions about community-scale response to environmental shifts, and a useful strategy to this end is to employ hierarchical co-demographic models that directly test multi-taxa hypotheses within a single, unified analysis while benefiting in statistical power from aggregating datasets. This approach has been applied to classical phylogeographic datasets such as mitochondrial barcodes as well as reduced-genome polymorphism datasets that can yield 10,000s of SNPs, produced by emergent technologies such as RAD-seq and GBS. A strategy for the latter had been accomplished by adapting the site frequency spectrum to a novel summarization of population genomic data across multiple taxa called the aggregate site frequency spectrum (aSFS), which potentially can be deployed under various inferential frameworks including approximate Bayesian computation, random forest, and composite likelihood optimization. Here, we introduce the R package Multi-DICE, a wrapper program that exploits existing simulation software for straight-forward and flexible execution of hierarchical model-based inference using the aSFS, which is derived from genomic-scale data, as well as mitochondrial data. We validate several novel software features such as applying alternative inferential frameworks, enforcing a minimal threshold of time surrounding event pulses, and specifying flexible hyperprior distributions. In sum, Multi-DICE provides comparative analysis within the familiar R environment while allowing a high degree of user customization, and will thus serve as a valuable tool for comparative phylogeography and population genomics.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.