Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Data from: Phylogenomics from whole genome sequences using aTRAM
Novel sequencing technologies are rapidly expanding the size of data sets that can be applied to phylogenetic studies. Currently the most commonly used phylogenomic approaches involve some form of genome reduction. While these approaches make assembling phylogenomic data sets more economical for organisms with large genomes, they reduce the genomic coverage and thereby the long-term utility of the data. Currently, for organisms with moderate to small genomes (<1000 Mbp) it is feasible to sequence the entire genome at modest coverage (10−30×). Computational challenges for handling these large data sets can be alleviated by assembling targeted reads, rather than assembling the entire genome, to produce a phylogenomic data matrix. Here we demonstrate the use of automated Target Restricted Assembly Method (aTRAM) to assemble 1107 single-copy ortholog genes from whole genome sequencing of sucking lice (Anoplura) and out-groups. We developed a pipeline to extract exon sequences from the aTRAM assemblies by annotating them with respect to the original target protein. We aligned these protein sequences with the inferred amino acids and then performed phylogenetic analyses on both the concatenated matrix of genes and on each gene separately in a coalescent analysis. Finally, we tested the limits of successful assembly in aTRAM by assembling 100 genes from close- to distantly related taxa at high to low levels of coverage.
Data from: Comparative species divergence across eight triplets of spiny lizards (Sceloporus) using genomic sequence data
Species divergence is typically thought to occur in the absence of gene flow, but many empirical studies are discovering that gene flow may be more pervasive during species formation. Although many examples of divergence with gene flow have been identified, only few clades have been investigated in a comparative manner, and fewer have been studied using genome-wide sequence data. We contrast species divergence genetic histories across eight triplets of North American Sceloporus lizards using a maximum likelihood implementation of the isolation–migration (IM) model. Gene flow at the time of species divergence is modeled indirectly as variation in species divergence time across the genome or explicitly using a migration rate parameter. Likelihood ratio tests (LRTs) are used to test the null model of no gene flow at speciation against these two alternative gene flow models. We also use the Akaike information criterion to rank the models. Hundreds of loci are needed for the LRTs to have statistical power, and we use genome sequencing of reduced representation libraries to obtain DNA sequence alignments at many loci (between 340 and 3,478; mean 1⁄4 1,678) for each triplet. We find that current species distributions are a poor predictor of whether a species pair diverged with gene flow. Interrogating the genome using the triplet method expedites the comparative study of species divergence history and the estimation of genetic parameters associated with speciation.
Data from: Going where traditional markers have not gone before: utility of and promise for RAD-sequencing in marine invertebrate phylogeography and population genomics
Characterization of large numbers of single-nucleotide polymorphisms (SNPs) throughout a genome has the power to refine the understanding of population demographic history and to identify genomic regions under selection in natural populations. To this end, population genomic approaches that harness the power of next-generation sequencing to understand the ecology and evolution of marine invertebrates represent a boon to test long-standing questions in marine biology and conservation. We employed restriction-site-associated DNA sequencing (RAD-seq) to identify SNPs in natural populations of the sea anemone Nematostella vectensis, an emerging cnidarian model with a broad geographic range in estuarine habitats in North and South America, and portions of England. We identified hundreds of SNP-containing tags in thousands of RAD loci from 30 barcoded individuals inhabiting four locations from Nova Scotia to South Carolina. Population genomic analyses using high-confidence SNPs resulted in a highly-resolved phylogeography, a result not achieved in previous studies using traditional markers. Plots of locus-specific FST against heterozygosity suggest that a majority of polymorphic sites are neutral, with a smaller proportion suggesting evidence for balancing selection. Loci inferred to be under balancing selection were mapped to the genome, where 90% were located in gene bodies, indicating potential targets of selection. The results from analyses with and without a reference genome supported similar conclusions, further highlighting RAD-seq as a method that can be efficiently applied to species lacking existing genomic resources. We discuss the utility of RAD-seq approaches in burgeoning Nematostella research as well as in other cnidarian species, particularly corals and jellyfishes, to determine phylogeographic relationships of populations and identify regions of the genome undergoing selection.
Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure
PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.
Data from: The complete sequence of the mitochondrial genome of Butomus umbellatus - a member of an early branching lineage of monocotyledons
In order to study the evolution of mitochondrial genomes in the early branching lineages of the monocotyledons, i.e., the Acorales and Alismatales, we are sequencing complete genomes from a suite of key taxa. As a starting point the present paper describes the mitochondrial genome of Butomus umbellatus (Butomaceae) based on next-generation sequencing data. The genome was assembled into a circular molecule, 450,826 bp in length. Coding sequences cover only 8.2% of the genome and include 28 protein coding genes, four rRNA genes, and 12 tRNA genes. Some of the tRNA genes and a 16S rRNA gene are transferred from the plastid genome. However, the total amount of recognized plastid sequences in the mitochondrial genome is only 1.5% and the amount of DNA transferred from the nucleus is also low. RNA editing is abundant and a total of 557 edited sites are predicted in the protein coding genes. Compared to the 40 angiosperm mitochondrial genomes sequenced to date, the GC content of the Butomus genome is uniquely high (49.1%). The overall similarity between the mitochondrial genomes of Butomus and Spirodela (Araceae), the closest relative yet sequenced, is low (less than 20%), and the two genomes differ in size by a factor 2. Gene order is also largely unconserved. However, based on its phylogenetic position within the core alismatids Butomus will serve as a good reference point for subsequent studies in the early branching lineages of the monocotyledons.
Data from: Assemblage Accumulation Curves: A framework for resolving species accumulation in biological communities using chloroplast genome sequences
The timing and tempo of the processes involved in community assembly are of substantial concern to community ecologists and conservation managers. The fossil record is a valuable source of data for studying past changes in community composition, but it is not always detailed enough to allow the process of community assembly to be resolved at regional or site scales while tracing the trajectories of known species with associated known traits. We present a three‐step framework for studying present‐day species accumulation through time: DNA sampling from multiple individuals from multiple species within a community; estimates of coalescence times for each species using molecular dating methods; and plotting the accumulation of present‐day species through time using the inferred population ages. Our approach is illustrated using whole chloroplast genomes from plants from three rainforest communities in eastern Australia. Expected times to coalescence for multiple species in each community were inferred from pooled high‐throughput sequence libraries. Local assemblage accumulation curves for each community were constructed. We also explored the variation in assemblage accumulation curves of species with different functional traits. Models of equilibrium species richness informed our null hypothesis and largely explained the shape of the assemblage accumulation curves and indicated that the complexities of the accumulation process should be explored with additional parameters, for example allowing species classes with different extinction rates. The assemblage accumulation curves for the study sites showed evidence of recent population expansions within each of the communities. This signal of recent accumulation is consistent with the increase in suitable rainforest habitat that followed the Last Glacial Maximum. Our method of constructing assemblage accumulation curves provides a simple approach for visualizing species‐accumulation data. It can be used to test hypotheses such as the relative survival potential of species‐specific ecological attributes. Although our example used single‐nucleotide polymorphisms derived from whole‐chloroplast sequencing, this framework can be applied to mitochondrial genomes and to communities of other organisms.
Data from: Transcriptome sequencing reveals both neutral and adaptive genome dynamics in a marine invader
Species invasions cause significant ecological and economic damage, and genetic information is important to understanding and managing invasive species. In the ocean, many invasive species have high dispersal and gene flow, lowering the discriminatory power of traditional genetic approaches. High-throughput sequencing holds tremendous promise for increasing resolution and illuminating the relative contributions of selection and drift in marine invasion, but has not yet been used to compare the diversity and dynamics of a high-dispersal invader in its native and invaded ranges. We test a transcriptome-based approach in the European green crab (Carcinus maenas), a widespread invasive species with high gene flow and a well-known invasion history, in two native and five invasive populations. A panel of 10 809 transcriptome-derived nuclear SNPs identified significant population structure among highly bottlenecked invasive populations that were previously undifferentiated with traditional markers. Comparing the full data set and a subset of 9246 putatively neutral SNPs strongly suggested that non-neutral processes are the primary driver of population structure within the species' native range, while neutral processes appear to dominate in the invaded range. Non-neutral native range structure coincides with significant differences in intraspecific thermal tolerance, suggesting temperature as a potential selective agent. These results underline the importance of adaptation in shaping intraspecific differences even in high geneflow marine invasive species. They also demonstrate that high-throughput approaches have broad utility in determining neutral structure in recent invasions of such species. Together, neutral and non-neutral data derived from high-throughput approaches may increase the understanding of invasion dynamics in high-dispersal species.
Data from: Whole genome-sequencing and phylogenetic analysis of a historical collection of Bacillus anthracis strains from Danish cattle
Bacillus anthracis, the causative agent of anthrax, is known as one of the most genetically monomorphic species. Canonical single-nucleotide polymorphism (SNP) typing and whole-genome sequencing were used to investigate the molecular diversity of eleven B. anthracis strains isolated from cattle in Denmark between 1935 and 1988. Danish strains were assigned into five canSNP groups or lineages, i.e. A.Br.001/002 (n = 4), A.Br.Ames (n = 2), A.Br.008/011 (n = 2), A.Br.005/006 (n = 2) and A.Br.Aust94 (n = 1). The match with the A.Br.Ames lineage is of particular interest as the occurrence of such lineage in Europe is demonstrated for the first time, filling an historical gap within the phylogeography of the lineage. Comparative genome analyses of these strains with 41 isolates from other parts of the world revealed that the two Danish A.Br.008/011 strains were related to the heroin-associated strains responsible for outbreaks of injection anthrax in drug users in Europe. Eight novel diagnostic SNPs that specifically discriminate the different sub-groups of Danish strains were identified and developed into PCR-based genotyping assays.
Data from: Genome-wide assessment of population structure and genetic diversity and development of a core germplasm set for sweet potato based on specific length amplified fragment (SLAF) sequencing
Sweet potato, Ipomoea batatas (L.) Lam., is an important food crop that is cultivated worldwide. However, no genome-wide assessment of the genetic diversity of sweet potato has been reported to date. In the present study, the population structure and genetic diversity of 197 sweet potato accessions most of which were from China were assessed using 62,363 SNPs. A model-based structure analysis divided the accessions into three groups: group 1, group 2 and group 3. The genetic relationships among the accessions were evaluated using a phylogenetic tree, which clustered all the accessions into three major groups. A principal component analysis (PCA) showed that the accessions were distributed according to their population structure. The mean genetic distance among accessions ranged from 0.290 for group 1 to 0.311 for group 3, and the mean polymorphic information content (PIC) ranged from 0.232 for group 1 to 0.251 for group 3. The mean minor allele frequency (MAF) ranged from 0.207 for group 1 to 0.222 for group 3. Analysis of molecular variance (AMOVA) showed that the maximum diversity was within accessions (89.569%). Using CoreHunter software, a core set of 39 accessions was obtained, which accounted for approximately 19.8% of the total collection. The core germplasm set of sweet potato developed will be a valuable resource for future sweet potato improvement strategies.
Data from: Genomic predictions and genome-wide association study of resistance against Piscirickettsia salmonis in coho salmon (Oncorhynchus kisutch) using ddRAD sequencing
Piscirickettsia salmonis is one of the main infectious diseases affecting coho salmon (Oncorhynchus kisutch) farming, and current treatments have been ineffective for the control of this disease. Genetic improvement for P. salmonis resistance has been proposed as a feasible alternative for the control of this infectious disease in farmed fish. Genotyping by sequencing (GBS) strategies allow genotyping of hundreds of individuals with thousands of single nucleotide polymorphisms (SNPs), which can be used to perform genome wide association studies (GWAS) and predict genetic values using genome-wide information. We used double-digest restriction-site associated DNA (ddRAD) sequencing to dissect the genetic architecture of resistance against P. salmonis in a farmed coho salmon population and to identify molecular markers associated with the trait. We also evaluated genomic selection (GS) models in order to determine the potential to accelerate the genetic improvement of this trait by means of using genome-wide molecular information. A total of 764 individuals from 33 full-sib families (17 highly resistant and 16 highly susceptible) were experimentally challenged against P. salmonis and their genotypes were assayed using ddRAD sequencing. A total of 9,389 SNPs markers were identified in the population. These markers were used to test genomic selection models and compare different GWAS methodologies for resistance measured as day of death (DD) and binary survival (BIN). Genomic selection models showed higher accuracies than the traditional pedigree-based best linear unbiased prediction (PBLUP) method, for both DD and BIN. The models showed an improvement of up to 95% and 155% respectively over PBLUP. One SNP related with B-cell development was identified as a potential functional candidate associated with resistance to P. salmonis defined as DD.
Data from: Sturgeon conservation genomics: SNP discovery and validation using RAD sequencing
Caviar-producing sturgeons belonging to the genus Acipenser are considered to be one of the most endangered species groups in the world. Continued overfishing in spite of increasing legislation, zero catch quotas and extensive aquaculture production have led to the collapse of wild stocks across Europe and Asia. The evolutionary relationships among Adriatic, Russian, Persian and Siberian sturgeons are complex because of past introgression events and remain poorly understood. Conservation management, traceability and enforcement suffer a lack of appropriate DNA markers for the genetic identification of sturgeon at the species, population and individual level. This study employed RAD sequencing to discover and characterize single nucleotide polymorphism (SNP) DNA markers for use in sturgeon conservation in these four tetraploid species over three biological levels, using a single sequencing lane. Four population meta-samples and eight individual samples from one family were barcoded separately before sequencing. Analysis of 14.4 Gb of paired-end RAD data focused on the identification of SNPs in the paired-end contig, with subsequent in silico and empirical validation of candidate markers. Thousands of putatively informative markers were identified including, for the first time, SNPs that show population-wide differentiation between Russian and Persian sturgeons, representing an important advance in our ability to manage these cryptic species. The results highlight the challenges of genotyping-by-sequencing in polyploid taxa, while establishing the potential genetic resources for developing a new range of caviar traceability and enforcement tools.
Data from: Looking into the past – the reaction of three grouse species to climate change over the last million years using whole genome sequences
Tracking past population fluctuations can give insight into current levels of genetic variation present within species. Analysing population dynamics over larger time scales can be aligned to known climatic changes to determine the response of species to varying environments. Here, we applied the Pairwise Sequentially Markovian Coalescent (PSMC) model to infer past population dynamics of three widespread grouse species; black grouse, willow grouse and rock ptarmigan. This allowed the tracking of the effective population size (Ne) of all three species beyond 1 Mya, revealing that i) early Pleistocene cooling (~2.5 Mya) caused an increase in the willow grouse and rock ptarmigan populations, ii) the mid-Brunhes event (~430 kya) and following climatic oscillations decreased the Ne of willow grouse and rock ptarmigan, but increased the Ne of black grouse and iii) all three species reacted differently to the last glacial maximum (LGM) – black grouse increased prior to it, rock ptarmigan experienced a severe bottleneck and willow grouse was maintained at large population size. We postulate that the varying PSMC signal throughout the LGM depicts only the local history of the species. Nevertheless, the large population fluctuations in willow grouse and rock ptarmigan indicate that both species are opportunistic breeders while black grouse tracks the climatic changes more slowly and is maintained at lower Ne. Our results highlight the usefulness of the PSMC approach in investigating species' reaction to climate change in the deep past, but also that caution should be taken in drawing general conclusions about the recent past.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
This study aimed to develop a set of SNP markers with high resolution and accuracy within the African buffalo. Such a set can be used, among others, to depict subtle population genetic structure for a better understanding of buffalo population dynamics. In total, 18.5 million DNA sequences of 76 bp were generated by next generation sequencing on an Illumina Genome Analyzer II from a reduced representation library using DNA from a panel of 13 African buffalo representative of the four subspecies. We identified 2534 SNPs with high confidence within the panel by aligning the short sequences to the cattle genome (Bos taurus). The average sequencing depth of the complete aligned set of reads was estimated at 5x, and at 13x when only considering the final set of putative SNPs that passed the filtering criterion. Our set of SNPs was validated by PCR amplification and Sanger sequencing of 15 SNPs. Of these 15 SNPs, 14 amplified successfully and 13 were shown to be polymorphic (success rate: 87%). The fidelity of the identified set of SNPs and potential future applications are finally discussed.
Data from: Mass production of SNP markers in a nonmodel passerine bird through RAD sequencing and contig mapping to the zebra finch genome
Here, we present an adaptation of restriction-site-associated DNA sequencing (RAD-seq) to the Illumina HiSeq2000 technology that we used to produce SNP markers in very large quantities at low cost per unit in the Réunion grey white-eye (Zosterops borbonicus), a nonmodel passerine bird species with no reference genome. We sequenced a set of six pools of 18–25 individuals using a single sequencing lane. This allowed us to build around 600 000 contigs, among which at least 386 000 could be mapped to the zebra finch (Taeniopygia guttata) genome. This yielded more than 80 000 SNPs that could be mapped unambiguously and are evenly distributed across the genome. Thus, our approach provides a good illustration of the high potential of paired-end RAD sequencing of pooled DNA samples combined with comparative assembly to the zebra finch genome to build large contigs and characterize vast numbers of informative SNPs in nonmodel passerine bird species in a very efficient and cost-effective way.
Data from: Analysis of transposable elements in the genome of Asparagus officinalis from high coverage sequence data
Asparagus officinalis is an economically and nutritionally important vegetable crop that is widely cultivated and is used as a model dioecious species to study plant sex determination and sex chromosome evolution. To improve our understanding of its genome composition, especially with respect to transposable elements (TEs), which make up the majority of the genome, we performed Illumina HiSeq2000 sequencing of both male and female asparagus genomes followed by bioinformatics analysis. We generated 17 Gb of sequence (12×coverage) and assembled them into 163,406 scaffolds with a total cumulated length of 400 Mbp, which represent about 30% of asparagus genome. Overall, TEs masked about 53% of the A. officinalis assembly. Majority of the identified TEs belonged to LTR retrotransposons, which constitute about 28% of genomic DNA, with Ty1/copia elements being more diverse and accumulated to higher copy numbers than Ty3/gypsy. Compared with LTR retrotransposons, non-LTR retrotransposons and DNA transposons were relatively rare. In addition, comparison of the abundance of the TE groups between male and female genomes showed that the overall TE composition was highly similar, with only slight differences in the abundance of several TE groups, which is consistent with the relatively recent origin of asparagus sex chromosomes. This study greatly improves our knowledge of the repetitive sequence construction of asparagus, which facilitates the identification of TEs responsible for the early evolution of plant sex chromosomes and is helpful for further studies on this dioecious plant.
scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution - Supplementary data and code
<p>Data and code to reproduce the analyses from the study: "scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution". </p>
Genome Sequence Assembly of Coffea arabica variety Geisha (UCDv1.0)
<p>Genome Sequence Assembly of <i>Coffea arabica</i> variety Geisha (UCDv1.0)</p>
High genomic plasticity and unique features of Xanthomonas translucens pv. graminis revealed through comparative analysis of complete genome sequences
<p><strong>Background:</strong> <i>Xanthomonas</i> <i>translucens</i> pv. <i>graminis</i> (<i>Xtg</i>) is a major bacterial pathogen of economically important forage grasses, causing severe yield losses. So far, genomic resources for this pathovar consisted mostly of draft genome sequences, and only one complete genome sequence was available, preventing comprehensive comparative genomic analyses. Such comparative analyses are essential in understanding the mechanisms involved in the virulence of pathogens and to identify virulence factors involved in pathogenicity.</p><p><strong>Results:</strong> In this study, we produced high-quality, complete genome sequences of four strains of <i>Xtg</i>, complementing the recently obtained complete genome sequence of the <i>Xtg </i>pathotype strain. These genomic resources allowed for a comprehensive comparative analysis, which revealed a high genomic plasticity with many chromosomal rearrangements, although the strains were highly related, with 99.9 to 100% average nucleotide identity. A high number of transposases were exclusively found in <i>Xtg </i>and corresponded to 413 to 457 insertion/excision transposable elements per strain. These mobile genetic elements are likely to be involved in the observed genomic plasticity and may play an important role in the adaptation of <i>Xtg</i>. The pathovar was found to lack a type IV secretion system, and it possessed the smallest set of type III effectors in the species. However, three XopE and XopX family effectors were found, while in the other pathovars of the species two or less were present. Additional genes that were specific to the pathovar were identified, including a unique set of minor pilins of the type IV pilus, 17 TonB-dependent receptors (TBDRs), and 11 degradative enzymes. </p><p><strong>Conclusion:</strong> These results suggest a high adaptability of <i>Xtg</i>, conferred by the abundance of mobile genetic elements, which may have led to the loss of many features. Conserved features that were specific to <i>Xtg </i>were identified, and further investigation will help to determine genes that are essential to pathogenicity and host adaptation of <i>Xtg</i>.</p>
Using a mobile Nanopore sequencing lab for end-to-end genomic surveillance of Plasmodium falciparum: a feasibility study
Open the record for dataset details and reuse information.
VCF file of whole genome sequencing data of 163 rats mapped jointly to mRatBN7.2
<p>We analyzed whole genome sequencing data of 163 rats. These data were first mapped to mRatBN7.2, followed by variant calling using deepvariant and joint analysis using GLNexus. Sites that are likely called due to base-level errors in mRatBN7.2 are removed. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.