Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
669
datasets available to search
ShareScore release 0.9.0
Dataset results
669 results for “comparative genomics”
Data from: Pinpointing genes underlying annual/perennial transitions with comparative genomics
Background: Transitions between perennial and an annual life history occur often in plant lineages, but the genes that control whether a plant is an annual or perennial are largely unknown. To identify genes that confer differences between annuals and perennials we compared the gene content of four pairs of sister lineages (Arabidopsis thaliana/Arabidopsis lyrata, Arabis montbretiana/Arabis alpina, Arabis verna/Aubrieta parviflora and Draba nemorosa/Draba hispanica) in the Brassicaceae in which each pair contains one annual and one perennial, plus one extra annual species (Capsella rubella). Results: After sorting all genes in all nine species into gene families, we identified five families in which well-annotated genes are present in the perennials A. lyrata and A. alpina, but are not present in any of the annual species. For the eleven genes in perennials in these families, an orthologous pseudogene or otherwise highly diverged gene was found in the syntenic region of the annual species in six cases. The five candidate families identified encode: a kinase, an oxidoreductase, a lactoylglutathione lyase, a F-box protein and a zinc finger protein. By comparing the active gene in the perennial to the pseudogene or heavily altered gene in the annual, dN and dS were calculated. The low dN/dS values in one kinase suggest that it became pseudogenized more recently, while the other kinase, F-box, oxidoreductase and zinc-finger became pseudogenized closer to the divergence between the annual-perennial pair. Conclusions: We identified five gene families that may be involved in the life history switch from perennial to annual. Considering the dN and dS data and whether syntenic pseudogenes were found and the potential functions of the genes, the F-box family is considered the most promising candidate for future functional studies to determine if it affects life history.
Comparative genomics reveals divergent thermal selection in warm- and cold-tolerant marine mussels
<p class="p1">Investigating the history of natural selection among closely related species can elucidate how genomes diverge in response to disparate environmental pressures. Molecular evolutionary approaches can be integrated with knowledge of gene functions to examine how evolutionary divergence may affect ecologically-relevant traits such as temperature tolerance and species distribution limits. Here, we integrate transcriptome-wide analyses of molecular evolution with knowledge from physiological studies to develop hypotheses regarding the functional classes of genes under positive selection in one of the world's most widespread invasive species, the warm-tolerant marine mussel <i>Mytilus </i><i>galloprovincialis</i>. Based on existing physiological information, we test the hypothesis that genomic functions previously linked to divergent temperature adaptation at the whole-organism level show accelerated molecular divergence between warm-adapted<i> M. </i><i>galloprovincialis</i> and cold-adapted congeners. Combined results from codon model tests and analyses of polymorphism and divergence reveal that divergent selection has affected genomic functions previously associated with species-specific expression responses to heat stress, namely oxidative stress defense and cytoskeletal stabilisation. Examining specific loci implicated in thermal tolerance among <i>Mytilus</i> species (based on interspecific biochemical or expression patterns), we find close functional similarities between known thermotolerance candidate genes under positive selection and positively selected loci under predicted genomic functions (those associated with divergent expression responses). Taken together, our findings suggest a contribution of temperature-dependent selection in the molecular divergence between warm- and cold-adapted <i>Mytilus </i>species that is largely consistent with results from physiological studies. More broadly, this study provides an example of how independent experimental evidence from ecophysiological investigations can inform evolutionary hypotheses about molecular adaptation in closely related non-model species.</p>
Data from: Phylogenetic and comparative genomics of the family Leptotrichiaceae and introduction of a novel fingerprinting MLVA for Streptobacillus moniliformis
Background: The Leptotrichiaceae are a family of fairly unnoticed bacteria containing both microbiota on mucous membranes as well as significant pathogens such as Streptobacillus moniliformis, the causative organism of streptobacillary rat bite fever. Comprehensive genomic studies in members of this family have so far not been carried out. We aimed to analyze 47 genomes from 20 different member species to illuminate phylogenetic aspects, as well as genomic and discriminatory properties. Results: Our data provide a novel and reliable basis of support for previously established phylogeny from this group and give a deeper insight into characteristics of genome structure and gene functions. Full genome analyses revealed that most S. moniliformis strains under study form a heterogeneous population without any significant clustering. Analysis of infra-species variability for this highly pathogenic rat bite fever organism led to the detection of three specific variable number tandem analysis loci with high discriminatory power. Conclusions: This highly useful and economical tool can be directly employed in clinical samples without laborious prior cultivation. Our and prospective case-specific data can now easily be compared by using a newly established MLVA database in order to gain a better insight into the epidemiology of this presumably under-reported zoonosis.
Data from: Comparative genomics of chemosensory protein genes reveals rapid evolution and positive selection in ant-specific duplicates
Gene duplications can have a major role in adaptation, and gene families underlying chemosensation are particularly interesting due to their essential role in chemical recognition of mates, predators and food resources. Social insects add yet another dimension to the study of chemosensory genomics, as the key components of their social life rely on chemical communication. Still, chemosensory gene families are little studied in social insects. Here we annotated chemosensory protein (CSP) genes from seven ant genomes and studied their evolution. The number of functional CSP genes ranges from 11 to 21 depending on species, and the estimated rates of gene birth and death indicate high turnover of genes. Ant CSP genes include seven conservative orthologous groups present in all the ants, and a group of genes that has expanded independently in different ant lineages. Interestingly, the expanded group of genes has a differing mode of evolution from the orthologous groups. The expanded group shows rapid evolution as indicated by a high dN/dS (nonsynonymous to synonymous changes) ratio, several sites under positive selection and many pseudogenes, whereas the genes in the seven orthologous groups evolve slowly under purifying selection and include only one pseudogene. These results show that adaptive changes have played a role in ant CSP evolution. The expanded group of ant-specific genes is phylogenetically close to a conservative orthologous group CSP7, which includes genes known to be involved in ant nestmate recognition, raising an interesting possibility that the expanded CSPs function in ant chemical communication.
Data from: Assembly and comparative analysis of transposable elements from low coverage genomic sequence data in Asparagales
The research field of comparative genomics is moving from a focus on genes to a more holistic view including the repetitive complement. This study aimed to characterize relative proportions of the repetitive fraction of large, complex genomes in a non-model system. The monocotyledonous plant order Asparagales (onion, asparagus, agave) comprises some of the largest angiosperm genomes and represents variation in both genome size and structure (karyotype). Anonymous, low coverage, single-end Illumina data from eleven exemplar Asparagales taxa were assembled using a de novo method. Resulting contigs were annotated using a reference library of available monocot repetitive sequences. Mapping reads to contigs provided rough estimates of relative proportions of each type of transposon in the nuclear genome. The results were parsed into general repeat types and synthesized with genome size estimates and a phylogenetic context to describe the pattern of transposable element evolution among these lineages. The major finding is that while some lineages in Asparagales exhibit conservation in repeat proportions, there is generally wide variation in types and frequency of repeats. This approach is an appropriate first step in characterizing repeats in evolutionary lineages with a paucity of genomic resources.
Data from: Comparing genomic signatures of domestication in two Atlantic salmon (Salmo salar L.) populations with different geographical origins
Selective breeding and genetic improvement have left detectable signatures on the genomes of domestic species. The elucidation of such signatures is fundamental for detecting genomic regions of biological relevance to domestication and improving management practices. In aquaculture, domestication was carried out independently in different locations worldwide, which provides opportunities to study the parallel effects of domestication on the genome of individuals that have been selected for similar traits. In the present study, we aimed to detect potential genomic signatures of domestication in two independent pairs of wild/domesticated Atlantic salmon populations of Canadian and Scottish origins respectively. Putative genomic regions under divergent selection were investigated using a 200K SNP array by combining three different statistical methods based either on allele frequencies (LFMM, Bayescan) or haplotype differentiation (Rsb). We identified 337 and 270 SNPs potentially under divergent selection in wild and hatchery populations of Canadian and Scottish origins respectively. We observed little overlap between results obtained from different statistical methods, highlighting the need to test complementary approaches for detecting a broad range of genomic footprints of selection. The vast majority of the outliers detected were population-specific but we found four candidate genes that were shared between the populations. We propose that these candidate genes may play a role in the parallel process of domestication. Overall, our results suggest that genetic drift may have override the effect of artificial selection and/or point towards a different genetic basis underlying the expression of similar traits in different domesticated strains. Finally, it is likely that domestication may predominantly target polygenic traits (e.g., growth) such that its genomic impact might be more difficult to detect with methods assuming selective sweeps.
Comparative genomic analysis reveals cellulase plays an important role in the pathogenicity of Setosphaeria turcica f. sp. Zeae
<p><i><span>Setosphaeria turcica</span></i><span> f. sp. <i>sorghi</i> and <i>S. turcica</i> f. sp. <i>zeae</i>, the two formae speciales of <i>S. turcica</i>, cause northern leaf blight disease of sorghum and corn, respectively, and often cause serious economic losses. They show obvious host specialization and have a close evolutionary relationship. Genomic sequencing can provide more information for understanding the virulence mechanisms of pathogens. However, the complete genomic sequence of <i>S. turcica</i> f. sp. <i>sorghi</i> has not yet been reported, and no comparative genomic information is available for the two formae speciales. In this study, based on the analysis of genomic structure, there were more protein-coding genes in <i>S. turcica</i> f. sp. <i>sorghi</i> than <i>S. turcica</i> f. sp. <i>zeae</i></span><span>, </span><span>showing positive selection in the evolution of <i>S. turcica</i>. The results of genomic functional analysis showed that the two formae speciales had a large number of identical protein-coding genes, while there were also specific protein families, including metabolic pathway proteins, transport proteins, </span><span>CAZy</span><span>s, pathogen and host interaction proteins. We also investigated the expression of specific effector-coding genes in <i>S. turcica</i> f. sp. <i>zeae</i>, and found that the endo-1, 4-β-D-glucanase coding gene </span><i><span>CEL</span><span>2</span></i><span>, an important component of cellulase, was significantly up-regulated during the interaction process. Finally, gluconolactone inhibited cellulase activity and decreased infection rate and pathogenicity, which indicates that cellulase is essential for maintaining virulence. These findings demonstrate that cellulase plays an important role in the pathogenicity of <i>S. turcica</i> f. sp. <i>zeae</i>. </span></p>
Data from: Comparative genomics reveals insight into virulence strategies of plant pathogenic oomycetes
The kingdom Stramenopile includes diatoms, brown algae, and oomycetes. Plant pathogenic oomycetes, including Phytophthora, Pythium and downy mildew species, cause devastating diseases on a wide range of host species and have a significant impact on agriculture. Here, we report comparative analyses on the genomes of thirteen straminipilous species, including eleven plant pathogenic oomycetes, to explore common features linked to their pathogenic lifestyle. We report the sequencing, assembly, and annotation of six Pythium genomes and comparison with other stramenopiles including photosynthetic diatoms, and other plant pathogenic oomycetes such as Phytophthora species, Hyaloperonospora arabidopsidis, and Pythium ultimum var. ultimum. Novel features of the oomycete genomes include an expansion of genes encoding secreted effectors and plant cell wall degrading enzymes in Phytophthora species and an over-representation of genes involved in proteolytic degradation and signal transduction in Pythium species. A complete lack of classical RxLR effectors was observed in the seven surveyed Pythium genomes along with an overall reduction of pathogenesis-related gene families in H. arabidopsidis. Comparative analyses revealed fewer genes encoding enzymes involved in carbohydrate metabolism in Pythium species and H. arabidopsidis as compared to Phytophthora species, suggesting variation in virulence mechanisms within plant pathogenic oomycete species. Shared features between the oomycetes and diatoms revealed common mechanisms of intracellular signaling and transportation. Our analyses demonstrate the value of comparative genome analyses for exploring the evolution of pathogenesis and survival mechanisms in the oomycetes. The comparative analyses of seven Pythium species with the closely related oomycetes, Phytophthora species and H. arabidopsidis, and distantly related diatoms provide insight into genes that underlie virulence.
Data from: Comparative population genomics reveals key barriers to dispersal in Southern Ocean penguins
The mechanisms that determine patterns of species dispersal are important factors in the production and maintenance of biodiversity. Understanding these mechanisms helps to forecast the responses of species to environmental change. Here we used a comparative framework and genome-wide data obtained through RAD-seq to compare the patterns of connectivity among breeding colonies for five penguin species with shared ancestry, overlapping distributions, and differing ecological niches, allowing an examination of the intrinsic and extrinsic barriers governing dispersal patterns. Our findings show that at-sea range and oceanography underlie patterns of dispersal in these penguins. The pelagic niche of emperor (Aptenodytes forsteri), king (A. patagonicus), Adélie (Pygoscelis adeliae) and chinstrap (P. antarctica) penguins facilitates gene flow over thousands of kilometres. In contrast, the coastal niche of gentoo penguins (P. papua) limits dispersal, resulting in population divergences. Oceanographic fronts also act as dispersal barriers to some extent. We recommend that forecasts of extinction risk incorporate dispersal and that management units are defined by at-sea range and oceanography in species lacking genetic data.
Data from: Genome sequencing and comparative analysis of three Chlamydia pecorum strains associated with different pathogenic outcomes
Background: Chlamydia pecorum is the causative agent of a number of acute diseases, but most often causes persistent, subclinical infection in ruminants, swine and birds. In this study, the genome sequences of three C. pecorum strains isolated from the faeces of a sheep with inapparent enteric infection (strain W73), from the synovial fluid of a sheep with polyarthritis (strain P787) and from a cervical swab taken from a cow with metritis (strain PV3056/3) were determined using Illumina/Solexa and Roche 454 genome sequencing. Results: Gene order and synteny was almost identical between C. pecorum strains and C. psittaci. Differences between C. pecorum and other chlamydiae occurred at a number of loci, including the plasticity zone, which contained a MAC/perforin domain protein, two copies of a >3400 amino acid putative cytotoxin gene and four (PV3056/3) or five (P787 and W73) genes encoding phospholipase D. Chlamydia pecorum contains an almost intact tryptophan biosynthesis operon encoding trpABCDFR and has the ability to sequester kynurenine from its host, however it lacks the genes folA, folKP and folB required for folate metabolism found in other chlamydiae. A total of 15 polymorphic membrane proteins were identified, belonging to six pmp families. Strains possess an intact type III secretion system composed of 18 structural genes and accessory proteins, however a number of putative inc effector proteins widely distributed in chlamydiae are absent from C. pecorum. Two genes encoding the hypothetical protein ORF663 and IncA contain variable numbers of repeat sequences that could be associated with persistence of infection. Conclusions: Genome sequencing of three C. pecorum strains, originating from animals with different disease manifestations, has identified differences in ORF663 and pseudogene content between strains and has identified genes and metabolic traits that may influence intracellular survival, pathogenicity and evasion of the host immune system.
Data from: Population genomics of pearl millet (Pennisetum glaucum (L.) R. Br.): comparative analysis of global accessions and Senegalese landraces
Background: Pearl millet is a staple food for people in arid and semi-arid regions of Africa and South Asia due to its high drought tolerance and nutritional qualities. A better understanding of the genomic diversity and population structure of pearl millet germplasm is needed to support germplasm conservation and genetic improvement of this crop. Here we characterized two pearl millet diversity panels, (i) a set of global accessions from Africa, Asia, and the America, and (ii) a collection of landraces from multiple agro-ecological zones in Senegal. Results: We identified 83,875 single nucleotide polymorphisms (SNPs) in 500 pearl millet accessions, comprised of 252 global accessions and 248 Senegalese landraces, using genotyping by sequencing (GBS) of PstI-MspI reduced representation libraries. We used these SNPs to characterize genomic diversity and population structure among the accessions. The Senegalese landraces had the highest levels of genetic diversity (π), while accessions from southern Africa and Asia showed lower diversity levels. Principal component analyses and ancestry estimation indicated clear population structure between the Senegalese landraces and the global accessions, and among countries in the global accessions. In contrast, little population structure was observed across in the Senegalese landraces collections. We ordered SNPs on the pearl millet genetic map and observed much faster linkage disequilibrium (LD) decay in Senegalese landraces compared to global accessions. A comparison of pearl millet GBS linkage map with the foxtail millet (Setaria italica) and sorghum (Sorghum bicolor) genomes indicated extensive regions of synteny, as well as some large-scale rearrangements in the pearl millet lineage. Conclusions: We identified 83,875 SNPs as a genomic resource for pearl millet improvement. The high genetic diversity in Senegal relative to other regions of Africa and Asia supports a West African origin of this crop, followed by wide diffusion. The rapid LD decay and lack of confounding population structure along agro-ecological zones in Senegalese pearl millet will facilitate future association mapping studies. Comparative population genomics will provide insights into panicoid crop evolution and support improvement of these climate-resilient crops.
FIGURE 1 in Complete mitochondrial genomes of three crickets (Orthoptera: Gryllidae) and comparative analyses within Ensifera mitogenomes
FIGURE 1. Comparison of AT skews of Grylloidea, Gryllotalpoidea and Tettigonioidea.
FIGURE 3 in Mitochondrial genome of Abraxas suspecta (Lepidoptera: Geometridae) and comparative analysis with other Lepidopterans
FIGURE 3. Codon distribution in members of the Lepidoptera. CDspT = codons per thousand codons.
FIGURE 5 in Mitochondrial genome of Abraxas suspecta (Lepidoptera: Geometridae) and comparative analysis with other Lepidopterans
FIGURE 5. Putative secondary structures of the 23 tRNA genes of the A. suspecta mitogenome.
Comparative genomic analysis of the mutant Rhodotorula mucilaginosa JH-R23 provides insight into the high-yield carotenoid mechanism
Open the record for dataset details and reuse information.
Figure 6 from: Zhang Q-H, Huang P, Chen B, Li T-J (2018) The complete mitochondrial genome of Orancistrocerus aterrimus aterrimus and comparative analysis in the family Vespidae (Hymenoptera, Vespidae, Eumeninae). ZooKeys 790: 127-144. https://doi.org/10.3897/zookeys.790.25356
Figure 6 The phylogenetic relationships were established by the 13 PCGs using ML (A) and BI (B) methods. Numbers abutting branches were bootstrap percentages with 1000 replicates (A) and Bayesian posterior probabilities (B). Red pentagram refers to the mitogenome sequences of O.a.aterrimus.
Figure 5 from: Zhang Q-H, Huang P, Chen B, Li T-J (2018) The complete mitochondrial genome of Orancistrocerus aterrimus aterrimus and comparative analysis in the family Vespidae (Hymenoptera, Vespidae, Eumeninae). ZooKeys 790: 127-144. https://doi.org/10.3897/zookeys.790.25356
Figure 5 Secondary structures of 23 tRNAs of O.a.aterrimus mitochondrial genome. Watson-Crick bonds are showed by dashes, GU pairs by filled dots, and AG and UU by open dots.
Figure 2 from: Zhang Q-H, Huang P, Chen B, Li T-J (2018) The complete mitochondrial genome of Orancistrocerus aterrimus aterrimus and comparative analysis in the family Vespidae (Hymenoptera, Vespidae, Eumeninae). ZooKeys 790: 127-144. https://doi.org/10.3897/zookeys.790.25356
Figure 2 Mitochondrial gene arrangement of 12 species of Vespidae. The red fonts indicate the rearrangement of the genes.
Figure 4 from: Zhang Q-H, Huang P, Chen B, Li T-J (2018) The complete mitochondrial genome of Orancistrocerus aterrimus aterrimus and comparative analysis in the family Vespidae (Hymenoptera, Vespidae, Eumeninae). ZooKeys 790: 127-144. https://doi.org/10.3897/zookeys.790.25356
Figure 4 Relative synonymous codon usage (RSCU) in Vespidae. Codon families are displayed along the x-axis.
Figure 1 from: Zhang Q-H, Huang P, Chen B, Li T-J (2018) The complete mitochondrial genome of Orancistrocerus aterrimus aterrimus and comparative analysis in the family Vespidae (Hymenoptera, Vespidae, Eumeninae). ZooKeys 790: 127-144. https://doi.org/10.3897/zookeys.790.25356
Figure 1 The mitochondrial genome of O.a.aterrimus. Arrows indicate the direction of genes. Abbreviations of the gene name are as follows: nad1-4 and nad4L act as nicotinamide adenine dinucleotide hydrogen dehydrogenase subunits 1-6 and 4L; cox1, cox2, and cox3 act as the cytochrome C oxidase subunits; cytb act as cytochrome b; atp8 and atp6 act as adenosine triphosphate synthase subunits 6 and 8; rrnL and rrnS act as large and small rRNA subunits; In addition, CR indicates control region and NCR indicates non-coding region.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.