Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Supplementary material 2 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figures S1–S13
Figure 3 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figure 3 Habitus images of NebriaAN. (Eonebria) djakonovi Semenov & Znojko BN. (Orientonebria) coreica Solsky CN. (Spelaeonebria) nudicollis Peyerimhoff DN. (Psilonebria) superna Andrewes EN. (Reductonebria) ochotica Sahlberg FN. (Catonebria) banksii Crotch. Scale bars: 1.0 mm. Photograph credits: A, B, F Kiril Makarov; C, D David Maddison; E Alexander Anischenko.
Figure 2 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figure 2 Habitus images of NebriiniALeistus (Nebrileistus) nubivagus Wollaston BL. (Leistus) ferruginosus Mannerheim CArchastes solitarius (Ledoux & Roux) DNippononebria (Vancouveria) virescens (Horn) ENebria (Oreonebria) castanea Bonelli FN. (Eurynebria) complanata (Linnaeus). Scale bars: 1.0 mm. Photograph credits: A, D–F David Maddison; B, C Alexander Anischenko.
Figure 6 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figure 6 Summary tree of nebriite phylogeny illustrating the revised classification; clade representation in Europe (including North Africa and the Middle East), Asia, and North America is indicated in the three-box bar.
Figure 5 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figure 5 Majority rule consensus tree of trees from bootstrap replicates. The first number under a branch is the percentage of bootstrap replicates with that clade, the second number is the estimate of the Bayesian posterior probability of that clade expressed as a percentage.
Figure 1 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245
Figure 1 Habitus images of NebriitaeANotiokasis chaudoiri Kavanaugh & Nègre BPelophila borealis (Paykull) COpisthius richardsoni Kirby DParopisthius indicus chinensis Bousquet & Smetana ENotiophilus palustris Duftschmid FArchileistobrius hwangtienyuni Shilenkov & Kryzhanovskij. Scale bars: 1.0 mm. Photograph credits: A, C David Maddison; B, E Kiril Makarov; D, F Alexander Anischenko.
Data from: An integrated model of phenotypic trait changes and site-specific sequence evolution
Recent years have seen a constant rise in the availability of trait data, including morphological features, ecological preferences, and life history characteristics. These phenotypic data provide means to associate genomic regions with phenotypic attributes, thus allowing the identification of phenotypic traits associated with the rate of genome and sequence evolution. However, inference methodologies that analyze sequence and phenotypic data in a unified statistical framework are still scarce. Here, we present TraitRateProp, a probabilistic method that allows testing whether the rate of sequence evolution is associated with a binary phenotypic character trait. The method further allows the detection of specific sequence sites whose evolutionary rate is most noticeably affected following the character transition, suggesting a shift in functional/structural constraints. TraitRateProp is first evaluated in simulations and then applied to study the evolutionary process of plastid plant genomes upon a transition to a heterotrophic lifestyle. To this end, we analyze 25 plastid genes across 85 orchid species, spanning different lifestyles and representing different genera in this large family of flowering plants. Our results indicate higher evolutionary rates following repeated transitions to a heterotrophic lifestyle in all but four of the loci analyzed.
Data from: PSMC (pairwise sequentially Markovian coalescent) analysis of RAD (restriction site associated DNA) sequencing data
The pairwise sequentially Markovian coalescent (PSMC) method uses the genome sequence of a single individual to estimate demographic history covering a time span of thousands of generations. Although originally designed for whole-genome data, we here use simulations to investigate its applicability to reference genome-aligned restriction site associated DNA (RAD) data. We find that RAD data can potentially be used for PSMC analysis, but at present with limitations. The key factor is the proportion (p) of the genome that the RAD data covers. In our simulations, a proportion of 10% can still retain a substantial amount of coalescent information, whereas for 1% estimation becomes unreliable. The performance depends strongly on mutation rate (μ) and recombination rate (r) and is proportional to μ*p/r. When the value of this term is low, increasing the amount of data and number of iterations helps restoring the power of the estimation. We subsequently analyse one whole-genome-sequenced and 17 RAD-sequenced three-spined sticklebacks (Gasterosteus aculeatus) from a lake in Greenland. The whole-genome sequence suggests a relatively recent expansion and decline within ca. 4000–40 000 generations ago, possibly reflecting postglacial expansion and founding of the lake population. RAD data, where chromosomes from 10 individuals are combined, identify a similar pattern. Our study provides guidance about the use of PSMC analysis and suggests measures that can improve its utility for RAD data. Finally, the study shows that RAD loci in general contain coalescent information that can be used for developing more targeted methods.
Data from: Estimation of contemporary effective population size and population declines using RAD sequence data
Large genomic datasets generated with restriction-site associated DNA sequencing (RADseq), in combination with demographic inference methods, are improving our ability to gain insights into the population history of species. We used a simulation approach to examine the potential for RADseq datasets to accurately estimate effective population size (Ne) over the course of stable and declining population trends, and we compare the ability of two methods of analysis to accurately distinguish stable from steadily declining populations over a contemporary time scale (20 generations). Using a linkage disequilibrium-based analysis, individual sampling (i.e., n ≥ 30) had the greatest effect on Ne estimation and the detection of population-size declines, with declines reliably detected across scenarios approximately 10 generations after they began. Coalescent-based inference required fewer sampled individuals (i.e., n = 15), and instead was most influenced by the size of the SNP dataset, with 25,000 to 50,000 SNPs required for accurate detection of population trends and at least 20 generations after decline began. The number of samples available and targeted number of RADseq loci are important criteria when choosing between these methods. Neither method suffered any apparent bias due to the effects of allele dropout typical of RAD data. With an understanding of the limitations and biases of these approaches, researchers can make more informed decisions when designing their sampling and analyses. Overall, our results reveal that demographic inference using RADseq data can be successfully applied to infer recent population size change and may be important tools for population monitoring and conservation biology.
Data from: Efficient detection of novel nuclear markers for Brassicaceae by transcriptome sequencing
The lack of DNA sequence information for most non-model organisms impairs the design of primers that are universally applicable for the study of molecular polymorphisms in nuclear markers. Next-generation sequencing (NGS) techniques nowadays provide a powerful approach to overcome this limitation. We present a flexible and inexpensive method to identify large numbers of nuclear primer pairs that amplify in most Brassicaceae species. We first obtained and mapped NGS transcriptome sequencing reads from two of the distantly related Brassicaceae species, Cardamine hirsuta and Arabis alpina, onto the Arabidopsis thaliana reference genome, and then identified short conserved sequence motifs among the three species bioinformatically. From these, primer pairs to amplify coding regions (nuclear protein coding loci, NPCL) and exon-primed intron-crossing sequences (EPIC) were developed. We identified 2,334 universally applicable primer pairs, targeting 1,164 genes, which provide a large pool of markers as readily usable genomic resource that will help addressing novel questions in the Brassicaceae family. Testing a subset of the newly designed nuclear primer pairs revealed that a great majority yielded a single amplicon in all of the 30 investigated Brassicaceae taxa. Sequence analysis and phylogenetic reconstruction with a subset of these markers on different levels of phylogenetic divergence in the mustard family were compared with previous studies. The results corroborate the usefulness of the newly developed primer pairs, e.g., for phylogenetic analyses or population genetic studies. Thus, our method provides a cost-effective approach for designing nuclear loci across a broad range of taxa and is compatible with current NGS technologies.
Data from: Paralogs are revealed by proportion of heterozygotes and deviations in read ratios in genotyping by sequencing data from natural populations
Whole genome duplications have occurred in the recent ancestors of many plants, fish, and amphibians, resulting in a pervasiveness of paralogous loci and the potential for both disomic and tetrasomic inheritance in the same genome. Paralogs can be difficult to reliably genotype and are often excluded from genotyping-by-sequencing (GBS) analyses; however, removal requires paralogs to be identified which is difficult without a reference genome. We present a method for identifying paralogs in natural populations by combining two properties of duplicated loci: 1) the expected frequency of heterozygotes exceeds that for singleton loci, and 2) within heterozygotes, observed read ratios for each allele in GBS data will deviate from the 1:1 expected for singleton (diploid) loci. These deviations are often not apparent within individuals, particularly when sequence coverage is low; but, we postulated that summing allele reads for each locus over all heterozygous individuals in a population would provide sufficient power to detect deviations at those loci. We identified paralogous loci in three species: Chinook salmon (Oncorhynchus tshawytscha) which retains regions with ongoing residual tetrasomy on eight chromosome arms following a recent whole genome duplication, mountain barberry (Berberis alpina) which has a large proportion of paralogs that arose through an unknown mechanism, and dusky parrotfish (Scarus niger) which has largely re-diploidized following an ancient whole genome duplication. Importantly, this approach only requires the genotype and allele-specific read counts for each individual, information which is readily obtained from most GBS analysis pipelines.
Data from: Fast and cost-effective genetic mapping in apple using next-generation sequencing
Next-generation DNA sequencing (NGS) produces vast amounts of DNA sequence data, but it is not specifically designed to generate data suitable for genetic mapping. Recently developed DNA library preparation methods for NGS have helped solve this problem, however, by combining the use of reduced representation libraries with DNA sample barcoding to generate genome-wide genotype data from a common set of genetic markers across a large number of samples. Here we use such a method, called genotyping-by-sequencing (GBS), to produce a data set for genetic mapping in an F1 population of apples (Malus x domestica) segregating for skin color. We show that GBS produces a relatively large, but extremely sparse, genotype matrix: over 270,000 SNPs were discovered, but most SNPs have too much missing data across samples to be useful for genetic mapping. After filtering for genotype quality and missing data, only 6% of the 85 million DNA sequence reads contributed to useful genotype calls. Despite this limitation, using existing software and a set of simple heuristics, we generated a final genotype matrix containing 3967 SNPs from 89 DNA samples from a single lane of Illumina HiSeq and used it to create a saturated genetic linkage map and to identify a known QTL underlying apple skin color. We therefore demonstrate that GBS is a cost effective method for generating genome-wide SNP data suitable for genetic mapping in a highly diverse and heterozygous agricultural species. We anticipate future improvements to the GBS analysis pipeline presented here that will enhance the utility of next-generation DNA sequence data for the purposes of genetic mapping across diverse species.
Data from: An evaluation of the hybrid speciation hypothesis for Xiphophorus clemenciae based on whole genome sequences
Once thought rare in animal taxa, hybridization has been increasingly recognized as an important and common force in animal evolution. In the past decade, a number of studies have suggested that hybridization has driven speciation in some animal groups. We investigate the signature of hybridization in the genome of a putative hybrid species, Xiphophorus clemenciae, through whole genome sequencing of this species and its hypothesized progenitors. Based on analysis of this data, we find that X. clemenciae is unlikely to have been derived from admixture between its proposed parental species. However, we find significant evidence for recent gene flow between Xiphophorus species. Though we detect genetic exchange in two pairs of species analyzed, the proportion of genomic regions that can be attributed to hybrid origin is small, suggesting that strong behavioral pre-mating isolation prevents frequent hybridization in Xiphophorus. The direction of gene flow between species supports a role for sexual selection in mediating hybridization.
Data from: Impacts of degraded DNA on restriction enzyme associated DNA sequencing (RADSeq)
Degraded DNA from suboptimal field sampling is common in molecular ecology. However, its impact on techniques that use restriction site associated next-generation DNA sequencing (RADSeq, GBS) is unknown. We experimentally examined the effects of in situDNA degradation on data generation for a modified double-digest RADSeq approach (3RAD). We generated libraries using genomic DNA serially extracted from the muscle tissue of 8 individual lake whitefish (Coregonus clupeaformis) following 0-, 12-, 48- and 96-h incubation at room temperature posteuthanasia. This treatment of the tissue resulted in input DNA that ranged in quality from nearly intact to highly sheared. All samples were sequenced as a multiplexed pool on an Illumina MiSeq. Libraries created from low to moderately degraded DNA (12–48 h) performed well. In contrast, the number of RADtags per individual, number of variable sites, and percentage of identical RADtags retained were all dramatically reduced when libraries were made using highly degraded DNA (96-h group). This reduction in performance was largely due to a significant and unexpected loss of raw reads as a result of poor quality scores. Our findings remained consistent after changes in restriction enzymes, modified fold coverage values (2- to 16-fold), and additional read-length trimming. We conclude that starting DNA quality is an important consideration for RADSeq; however, the approach remains robust until genomic DNA is extensively degraded.
Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis
Top predators are disappearing worldwide, significantly changing ecosystems that depend on top-down regulation. Conflict with humans remains the primary roadblock for large carnivore conservation, but for the eastern wolf (Canis lycaon), disagreement over its evolutionary origins presents a significant barrier to conservation in Canada and has impeded protection for grey wolves (Canis lupus) in the USA. Here, we use 127 235 single-nucleotide polymorphisms (SNPs) identified from restriction-site associated DNA sequencing (RAD-seq) of wolves and coyotes, in combination with genomic simulations, to test hypotheses of hybrid origins of Canis types in eastern North America. A principal components analysis revealed no evidence to support eastern wolves, or any other Canis type, as the product of grey wolf × western coyote hybridization. In contrast, simulations that included eastern wolves as a distinct taxon clarified the hybrid origins of Great Lakes-boreal wolves and eastern coyotes. Our results support the eastern wolf as a distinct genomic cluster in North America and help resolve hybrid origins of Great Lakes wolves and eastern coyotes. The data provide timely information that will shed new light on the debate over wolf conservation in eastern North America.
Data from: Sequence entropy of folding and the absolute rate of amino acid substitutions
Adequate representations of protein evolution should consider how the acceptance of mutations depends on the sequence context in which they arise. However, epistatic interactions among sites in a protein result in hererogeneities in the substitution rate, both temporal and spatial, that are beyond the capabilities of current models. Here we use parallels between amino acid substitutions and chemical reaction kinetics to develop an improved theory of protein evolution. We constructed a mechanistic framework for modelling amino acid substitution rates that uses the formalisms of statistical mechanics, with principles of population genetics underlying the analysis. Theoretical analyses and computer simulations of proteins under purifying selection for thermodynamic stability show that substitution rates and the stabilization of resident amino acids (the 'evolutionary Stokes shift') can be predicted from biophysics and the effect of sequence entropy alone. Furthermore, we demonstrate that substitutions predominantly occur when epistatic interactions result in near neutrality; substitution rates are determined by how often epistasis results in such nearly neutral conditions. This theory provides a general framework for modelling protein sequence change under purifying selection, potentially explains patterns of convergence and mutation rates in real proteins that are incompatible with previous models, and provides a better null model for the detection of adaptive changes.
Data from: Sequence capture versus restriction site associated DNA sequencing for shallow systematics
Sequence capture and restriction site associated DNA sequencing (RAD-Seq) are two genomic enrichment strategies for applying next-generation sequencing technologies to systematics studies. At shallow timescales, such as within species, RAD-Seq has been widely adopted among researchers, although there has been little discussion of the potential limitations and benefits of RAD-Seq and sequence capture. We discuss a series of issues that may impact the utility of sequence capture and RAD-Seq data for shallow systematics in non-model species. We review prior studies that used both methods, and investigate differences between the methods by re-analyzing existing RAD-Seq and sequence capture datasets from a Neotropical bird (Xenops minutus). We suggest that the strengths of RAD-Seq datasets for shallow systematics are the wide dispersion of markers across the genome, the relative ease and cost of laboratory work, the deep coverage and read overlap at recovered loci, and the high overall information that results. Sequence capture's benefits include flexibility and repeatability in the genomic regions targeted, success using low-quality samples, more straightforward read orthology assessment, and higher per-locus information content. The utility of a method in systematics, however, rests not only on its performance within a study, but on the comparability of datasets and inferences with those of prior work. In RAD-Seq datasets, comparability is compromised by low overlap of orthologous markers across species and the sensitivity of genetic diversity in a dataset to an interaction between the level of natural heterozygosity in the samples examined and the parameters used for orthology assessment. In contrast, sequence capture of conserved genomic regions permits interrogation of the same loci across divergent species, which is preferable for maintaining comparability among datasets and studies for the purpose of drawing general conclusions about the impact of historical processes across biotas. We argue that sequence capture should be given greater attention as a method of obtaining data for studies in shallow systematics and comparative phylogeography.
Data from: Genetic barcoding of dark-spored myxomycetes (Amoebozoa)—Identification, evaluation and application of a sequence similarity threshold for species differentiation in NGS studies
Unicellular, eukaryotic organisms (protists) play a key role in soil food webs as major predators of microorganisms. However, due to the polyphyletic nature of protists, no single universal barcode can be established for this group, and the structure of many protistean communities remains unresolved. Plasmodial slime moulds (Myxogastria or Myxomycetes) stand out among protists by their formation of fruit bodies, which allow for a morphological species concept. By Sanger sequencing of a large collection of morphospecies, this study presents the largest database to date of dark-spored myxomycetes and evaluate a partial 18S SSU gene marker for species annotation. We identify and discuss the use of an intraspecific sequence similarity threshold of 99.1% for species differentiation (OTU picking) in environmental PCR studies (ePCR) and estimate a hidden diversity of putative species, exceeding those of described morphospecies by 99%. When applying the identified threshold to an ePCR data set (including sequences from both NGS and cloning), we find 64 OTUs of which 21.9% had a direct match (>99.1% similarity) to the database and the remaining had on average 90.2 ± 0.8% similarity to their best match, thus thought to represent undiscovered diversity of dark-spored myxomycetes.
Data from: Genotyping-by-sequencing provides the first well-resolved phylogeny for coffee (Coffea) and insights into the evolution of caffeine content in its species: GBS coffee phylogeny and the evolution of caffeine content
A comprehensive and meaningful phylogenetic hypothesis for the commercially important coffee genus (Coffea) has long been a key objective for coffee researchers. For molecular studies, progress has been limited by low levels of sequence divergence, leading to insufficient topological resolution and statistical support in phylogenetic trees, particularly for the major lineages and for the numerous species occurring in Madagascar. We report here the first almost fully resolved, broadly sampled phylogenetic hypothesis for coffee, the result of combining genotyping-by-sequencing (GBS) technology with a newly developed, lab-based workflow to integrate short read next-generation sequencing for low numbers of additional samples. Biogeographic patterns indicate either Africa or Asia (or possibly the Arabian Peninsula) as the most likely ancestral locality for the origin of the coffee genus, with independent radiations across Africa, Asia, and the Western Indian Ocean Islands (including Madagascar and Mauritius). The evolution of caffeine, an important trait for commerce and society, was evaluated in light of our phylogeny. High and consistent caffeine content is found only in species from the equatorial, fully humid environments of West and Central Africa, possibly as an adaptive response to increased levels of pest predation. Moderate caffeine production, however, evolved at least one additional time recently (between 2 and 4 Mya) in a Madagascan lineage, which suggests that either the biosynthetic pathway was already in place during the early evolutionary history of coffee, or that caffeine synthesis within the genus is subject to convergent evolution, as is also the case for caffeine synthesis in coffee versus tea and chocolate.
Data from: "Transcriptome sequence identity between Lyme disease tick vectors, Ixodes scapularis and Ixodes ricinus" in Genomic Resources Notes accepted 1 April 2014 to 31 May 2014
Ixodes scapularis and I. ricinus transmit the Lyme disease agent Borrelia burgdorferi in the U.S. and Europe, respectively. The only tick genome sequence available is that of I. scapularis, which constitutes a limitation for tick research. Recent evidences suggest that I. ricinus and I. scapularis transcriptomes share some degree of sequence identity. However, only the global transcriptome comparison reported here demonstrated that I. ricinus and I. scapularis share a 99.232±0.005 percent sequence identity with a very low frequency of INDELs. However, due to limitations of the current I. scapularis genome assembly, the number of aligned reads was only 26-27%. These results support the use of I. scapularis genome sequence as a reference for the analysis of I. ricinus transcriptomics and proteomics data, but addressing the limitations associated with the I. scapularis genome assembly.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.