Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: A hypervariable mitochondrial protein coding sequence associated with geographical origin in a cosmopolitan bloom-forming alga, Heterosigma akashiwo

Geographic distributions of phytoplankton species can be defined by events on both evolutionary time and shorter scales, e.g., recent climate changes. Additionally, modern industrial activity, including the transport of live fish and spat for aquaculture and aquatic microorganisms in ship ballast water, may aid the spread of phytoplankton. Obtaining a reliable marker is key to gaining insight into the phylogeographic history of a species. Here, we report a hypervariable mitochondrial gene in the cosmopolitan bloom-forming alga, Heterosigma akashiwo. We compared the entire mitochondrial genome sequences of seven H. akashiwo strains from Japanese and North American coastal waters and identified a hypervariable segment. The region codes for a hypothetical protein with no defined function, and its variations between Japanese and North American isolates, were prominent, while the sequences were more conserved among Japanese strains and North American isolates. Comparison of the sequence in isolates obtained from different geographical points in the Northern Hemisphere revealed that the sequence variations largely correlated with latitude and longitude (i.e. Pacific/Atlantic oceans). Our results demonstrate the usefulness of the sequence in determining the phylogeographic history of H. akashiwo.

opencc-zeroDec 2016View details →
dryad32/100

Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013

Tree-roosting bats in the genus Lasiurus are widespread, migratory species that have not been well characterized for population genetic diversity and structure due to a lack of genetic resources. Generating genetic resources in Lasiurus is made pressing by the need for conservation genetic assessments of demographic trends in this genus, which comprise a large percentage of bat mortalities at wind turbine sites across North America. We report on marker development from whole-genome Illumina sequencing of the red bat (Lasirus borealis) and the hoary bat (L. cinereus). We generated paired-end libraries for a single individual of each species, sequenced on the Illumina HiSeq platform. We mapped a total of 46.6 million reads to the Myotis lucifigus reference genome, and used bioinformatics searches to identify tends of thousands of simple sequence repeats (SSRs) distributed across the bat genome. We selected 48 candidate microsatellite loci to develop cross-species primer sequences for Lasiurus, assembled these into multiplex combinations, and tested for amplification and polymorphism levels in a sample of 23 individuals from each of L. borealis and L. cinereus. In total, we identified 42 highly polymorphic loci that could be robustly amplified and scored, the majority of which (39) were also combinable into highly multiplexed assays of 4-8 loci each. The combination of new genomic sequence assemblies, a large set of highly polymorphic microsatellite loci, and the ability to efficiently multiplex represents a significant contribution to the genetic resources available for population and comparative genetic studies of bats.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Phylogenomics from whole genome sequences using aTRAM

Novel sequencing technologies are rapidly expanding the size of data sets that can be applied to phylogenetic studies. Currently the most commonly used phylogenomic approaches involve some form of genome reduction. While these approaches make assembling phylogenomic data sets more economical for organisms with large genomes, they reduce the genomic coverage and thereby the long-term utility of the data. Currently, for organisms with moderate to small genomes (<1000 Mbp) it is feasible to sequence the entire genome at modest coverage (10−30×). Computational challenges for handling these large data sets can be alleviated by assembling targeted reads, rather than assembling the entire genome, to produce a phylogenomic data matrix. Here we demonstrate the use of automated Target Restricted Assembly Method (aTRAM) to assemble 1107 single-copy ortholog genes from whole genome sequencing of sucking lice (Anoplura) and out-groups. We developed a pipeline to extract exon sequences from the aTRAM assemblies by annotating them with respect to the original target protein. We aligned these protein sequences with the inferred amino acids and then performed phylogenetic analyses on both the concatenated matrix of genes and on each gene separately in a coalescent analysis. Finally, we tested the limits of successful assembly in aTRAM by assembling 100 genes from close- to distantly related taxa at high to low levels of coverage.

opencc-zeroDec 2015View details →
dryad32/100

Data from: The genetic architecture of reproductive isolation during speciation-with-gene-flow in lake whitefish species pairs assessed by RAD sequencing

During speciation-with-gene-flow, effective migration varies across the genome as a function of several factors, including proximity of selected loci, recombination rate, strength of selection, and number of selected loci. Genome scans may provide better empirical understanding of the genome-wide patterns of genetic differentiation, especially if the variance due to the previously mentioned factors is partitioned. In North American lake whitefish (Coregonus clupeaformis), glacial lineages that diverged in allopatry about 60,000 years ago and came into contact 12,000 years ago have independently evolved in several lakes into two sympatric species pairs (a normal benthic and a dwarf limnetic). Variable degrees of reproductive isolation between species pairs across lakes offer a continuum of genetic and phenotypic divergence associated with adaptation to distinct ecological niches. To disentangle the complex array of genetically based barriers that locally reduce the effective migration rate between whitefish species pairs, we compared genome-wide patterns of divergence across five lakes distributed along this divergence continuum. Using restriction site associated DNA (RAD) sequencing, we combined genetic mapping and population genetics approaches to identify genomic regions resistant to introgression and derive empirical measures of the barrier strength as a function of recombination distance. We found that the size of the genomic islands of differentiation was influenced by the joint effects of linkage disequilibrium maintained by selection on many loci, the strength of ecological niche divergence, as well as demographic characteristics unique to each lake. Partial parallelism in divergent genomic regions likely reflected the combined effects of polygenic adaptation from standing variation and independent changes in the genetic architecture of postzygotic isolation. This study illustrates how integrating genetic mapping and population genomics of multiple sympatric species pairs provide a window on the speciation-with-gene-flow mechanism.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Estimation of a killer whale (Orcinus orca) population's diet using sequencing analysis of DNA from feces

Estimating diet composition is important for understanding interactions between predators and prey and thus illuminating ecosystem function. The diet of many species, however, is difficult to observe directly. Genetic analysis of fecal material collected in the field is therefore a useful tool for gaining insight into wild animal diets. In this study, we used high-throughput DNA sequencing to quantitatively estimate the diet composition of an endangered population of wild killer whales (Orcinus orca) in their summer range in the Salish Sea. We combined 175 fecal samples collected between May and September from five years between 2006 and 2011 into 13 sample groups. Two known DNA composition control groups were also created. Each group was sequenced at a ~330bp segment of the 16s gene in the mitochondrial genome using an Illumina MiSeq sequencing system. After several quality controls steps, 4,987,107 individual sequences were aligned to a custom sequence database containing 19 potential fish prey species and the most likely species of each fecal-derived sequence was determined. Based on these alignments, salmonids made up >98.6% of the total sequences and thus of the inferred diet. Of the six salmonid species, Chinook salmon made up 79.5% of the sequences, followed by coho salmon (15%). Over all years, a clear pattern emerged with Chinook salmon dominating the estimated diet early in the summer, and coho salmon contributing an average of >40% of the diet in late summer. Sockeye salmon appeared to be occasionally important, at >18% in some sample groups. Non-salmonids were rarely observed. Our results are consistent with earlier results based on surface prey remains, and confirm the importance of Chinook salmon in this population's summer diet.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Comparative species divergence across eight triplets of spiny lizards (Sceloporus) using genomic sequence data

Species divergence is typically thought to occur in the absence of gene flow, but many empirical studies are discovering that gene flow may be more pervasive during species formation. Although many examples of divergence with gene flow have been identified, only few clades have been investigated in a comparative manner, and fewer have been studied using genome-wide sequence data. We contrast species divergence genetic histories across eight triplets of North American Sceloporus lizards using a maximum likelihood implementation of the isolation–migration (IM) model. Gene flow at the time of species divergence is modeled indirectly as variation in species divergence time across the genome or explicitly using a migration rate parameter. Likelihood ratio tests (LRTs) are used to test the null model of no gene flow at speciation against these two alternative gene flow models. We also use the Akaike information criterion to rank the models. Hundreds of loci are needed for the LRTs to have statistical power, and we use genome sequencing of reduced representation libraries to obtain DNA sequence alignments at many loci (between 340 and 3,478; mean 1⁄4 1,678) for each triplet. We find that current species distributions are a poor predictor of whether a species pair diverged with gene flow. Interrogating the genome using the triplet method expedites the comparative study of species divergence history and the estimation of genetic parameters associated with speciation.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Microsatellite markers from the Ion Torrent: a multi-species contrast to 454 shotgun sequencing

The development and screening of microsatellite markers have been accelerated by next-generation sequencing (NGS) technology and in particular GS-FLX pyro-sequencing (454). More recent platforms such as the PGM semiconductor sequencer (Ion Torrent) offer potential benefits such as dramatic reductions in cost, but to date have not been well utilized. Here, we critically compare the advantages and disadvantages of microsatellite development using PGM semiconductor sequencing and GS-FLX pyro-sequencing for two gymnosperm (a conifer and a cycad) and one angiosperm species. We show that these NGS platforms differ in the quantity of returned sequence data, unique microsatellite data and primer design opportunities, mostly consistent with the differences in read length. The strength of the PGM lies in the large amount of data generated at a comparatively lower cost and time. The strength of GS-FLX lies in the return of longer average length sequences and therefore greater flexibility in producing markers with variable product length, due to longer flanking regions, which is ideal for capillary multiplexing. These differences need to be considered when choosing a NGS method for microsatellite discovery. However, the ongoing improvement in read lengths of the NGS platforms will reduce the disadvantage of the current short read lengths, particularly for the PGM platform, allowing greater flexibility in primer design coupled with the power of a larger number of sequences.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Multi-locus sequence data illuminate demographic drivers of Pleistocene speciation in semi-arid southern Australian birds (Cinclosoma spp.)

Background: During the Pleistocene, shifts of species distributions and their isolation in disjunct refugia led to varied outcomes in how taxa diversified. Some species diverged, others did not. Here, we begin to address another facet of the role of the Pleistocene in generating today's diversity. We ask which processes contributed to divergence in semi-arid southern Australian birds. We isolated 11 autosomal nuclear loci and one mitochondrial locus from a total of 29 specimens of the sister species pair, Chestnut Quail-thrush Cinclosoma castanotum and Copperback Quail-thrush C. clarum. Results: A population clustering analysis confirmed the location of the current species boundary as a well-known biogeographical barrier in southern Australia, the Eyrean Barrier. Coalescent-based analyses placed the time of species divergence to the Middle Pleistocene. Gene flow between the species since divergence has been low. The analyses suggest the effective population size of the ancestor was 54 to 178 times smaller than populations since divergence. This contrasts with recent multi-locus studies in some other Australian birds (butcherbirds, ducks) where a lack of phenotypic divergence was accompanied by larger historical population sizes. Post-divergence population size histories of C. clarum and C. castanotum were inferred using the extended Bayesian skyline model. The population size of C. clarum increased substantially during the late Pleistocene and continued to increase through the Last Glacial Maximum and Holocene. The timing of this expansion across its vast range is broadly concordant with that documented in several other Australian birds. In contrast, effective population size of C. castanotum was much more constrained and may reflect its smaller range and more restricted habitat east of the Eyrean Barrier compared with that available to C. clarum to the west. Conclusions: Our results contribute to awareness of increased population sizes, following significant contractions, as having been important in shaping diversity in Australian arid and semi-arid zones. Further, we improve knowledge of the role of Pleistocene climatic shifts in areas of the planet that were not glaciated at that time but which still experienced that period's cyclical climatic fluctuations.

opencc-zeroDec 2015View details →
dryad32/100

Data from: A pragmatic approach to the analysis of diets of generalist predators: the use of next-generation sequencing with no blocking probes

Predicting whether a predator is capable of affecting the dynamics of a prey species in the field implies the analysis of the complete diet of the predator, not simply rates of predation on a target taxon. Here, we employed the Ion Torrent next-generation sequencing technology to investigate the diet of a generalist arthropod predator. A complete dietary analysis requires the use of general primers, but these will also amplify the predator unless suppressed using a blocking probe. However, blocking probes can potentially block other species, particularly if they are phylogenetically close. Here, we aimed to demonstrate that enough prey sequence could be obtained without blocking probes. In communities with many predators, this approach obviates the need to design and test numerous blocking primers, thus making analysis of complex community food webs a viable proposition. We applied this approach to the analysis of predation by the linyphiid spider Oedothorax fuscus in an arable field. We obtained over two million raw reads. After discarding the low-quality and predator reads, the libraries still contained over 61 000 prey reads (3% of the raw reads; 6% of reads passing quality control). The libraries were rich in Collembola, Lepidoptera, Diptera and Nematoda. They also contained sequences derived from several spider species and from horticultural pests (aphids). Oedothorax fuscus is common in UK cereal fields, and the results showed that it is exploiting a wide range of prey. Next-generation sequencing using general primers but without blocking probes provided ample sequences for analysis of the prey range of this spider and proved to be a simple and inexpensive approach.

opencc-zeroDec 2012View details →
dryad32/100

Data from: RADcap: sequence capture of dual-digest RADseq libraries with identifiable duplicates and reduced missing data

Molecular ecologists seek to genotype hundreds to thousands of loci from hundreds to thousands of individuals at minimal cost per sample. Current methods, such as restriction site associated DNA sequencing (RADseq) and sequence capture, are constrained by costs associated with inefficient use of sequencing data and sample preparation. Here, we introduce RADcap, an approach that combines the major benefits of RADseq (low cost with specific start positions) with those of sequence capture (repeatable sequencing of specific loci) to significantly increase efficiency and reduce costs relative to current approaches. RADcap uses a new version of dual-digest RADseq (3RAD) to identify candidate SNP loci for capture bait design, and subsequently uses custom sequence capture baits to consistently enrich candidate SNP loci across many individuals. We combined this approach with a new library preparation method for identifying and removing PCR duplicates from 3RAD libraries, which allows researchers to process RADseq data using traditional pipelines, and we tested the RADcap method by genotyping sets of 96 to 384 Wisteria plants. Our results demonstrate that our RADcap method: (1) methodologically reduces (to <5%) and allows computational removal of PCR duplicate reads from data; (2) achieves 80-90% reads-on-target in 11 of 12 enrichments; (3) returns consistent coverage (≥4x) across >90% of individuals at up to 99.8% of the targeted loci; (4) produces consistently high occupancy matrices of genotypes across hundreds of individuals; and (5) costs significantly less than current approaches.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Allele phasing has minimal impact on phylogenetic reconstruction from targeted nuclear gene sequences in a case study of Artocarpus

Premise of the study: Untapped information about allelic diversity within populations and individuals (i.e. heterozygosity) could improve phylogenetic resolution and accuracy. Many phylogenetic reconstructions ignore heterozygosity because it is difficult to assemble allele sequences and combine allelic data across unlinked loci and it is unclear how reconstruction methods accommodate variable sequences. We review the common methods of including heterozygosity in phylogenetic studies and present a novel method for assembling allele sequences from target enriched Illumina sequencing libraries. Methods: We perform supermatrix phylogeny reconstruction and species tree estimation of Artocarpus based on three methods of accounting for heterozygous sequences: a consensus method based on de novo sequence assembly, the use of ambiguity characters, and a novel method for phasing alleles. We characterize the extent to which highly heterozygous sequences impeded phylogeny reconstruction and determine whether the use of allele sequences improves resolution or decreases topological uncertainty. Key Results: We show that it is possible to infer phased alleles from target enriched Illumina libraries. We find that highly heterozygous sequences do not contribute disproportionately to poor phylogenetic resolution and that the use of allele sequences for phylogeny reconstruction does not have a clear effect on phylogenetic resolution or topological consistency. Conclusions: We provide a framework for inferring phased alleles from target enrichment data and for assessing the contribution of allelic diversity to phylogenetic reconstruction. In our dataset, the impact of allele phasing on phylogeny is minimal compared to the impact of using phylogenetic reconstruction methods that account for gene tree incongruence.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Going where traditional markers have not gone before: utility of and promise for RAD-sequencing in marine invertebrate phylogeography and population genomics

Characterization of large numbers of single-nucleotide polymorphisms (SNPs) throughout a genome has the power to refine the understanding of population demographic history and to identify genomic regions under selection in natural populations. To this end, population genomic approaches that harness the power of next-generation sequencing to understand the ecology and evolution of marine invertebrates represent a boon to test long-standing questions in marine biology and conservation. We employed restriction-site-associated DNA sequencing (RAD-seq) to identify SNPs in natural populations of the sea anemone Nematostella vectensis, an emerging cnidarian model with a broad geographic range in estuarine habitats in North and South America, and portions of England. We identified hundreds of SNP-containing tags in thousands of RAD loci from 30 barcoded individuals inhabiting four locations from Nova Scotia to South Carolina. Population genomic analyses using high-confidence SNPs resulted in a highly-resolved phylogeography, a result not achieved in previous studies using traditional markers. Plots of locus-specific FST against heterozygosity suggest that a majority of polymorphic sites are neutral, with a smaller proportion suggesting evidence for balancing selection. Loci inferred to be under balancing selection were mapped to the genome, where 90% were located in gene bodies, indicating potential targets of selection. The results from analyses with and without a reference genome supported similar conclusions, further highlighting RAD-seq as a method that can be efficiently applied to species lacking existing genomic resources. We discuss the utility of RAD-seq approaches in burgeoning Nematostella research as well as in other cnidarian species, particularly corals and jellyfishes, to determine phylogeographic relationships of populations and identify regions of the genome undergoing selection.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Development of nuclear microsatellite loci and mitochondrial single nucleotide polymorphisms for the natterjack toad, Bufo (Epidalea) calamita (Bufonidae), using next generation sequencing and Competitive Allele Specific PCR (KASPar)

Amphibians are undergoing a major decline worldwide and the steady increase in the number of threatened species in this particular taxa highlights the need for conservation genetics studies using high-quality molecular markers. The natterjack toad, Bufo (Epidalea) calamita, is a vulnerable pioneering species confined to specialized habitats in Western Europe. To provide efficient and cost-effective genetic resources for conservation biologists, we developed and characterized 22 new nuclear microsatellite markers using next-generation sequencing. We also used sequence data acquired from Sanger sequencing to develop the first mitochondrial markers for KASPar assay genotyping. Genetic polymorphism was then analyzed for 95 toads sampled from 5 populations in France. For polymorphic microsatellite loci, number of alleles and expected heterozygosity ranged from 2 to 14 and from 0.035 to 0.720, respectively. No significant departures from panmixia were observed (mean multilocus F IS = −0.015) and population differentiation was substantial (mean multilocus F ST = 0.222, P < 0.001). From a set of 18 mitochondrial SNPs located in the 16S and D-loop region, we further developed a fast and cost-effective SNP genotyping method based on competitive allele-specific PCR amplification (KASPar). The combination of allelic states for these mitochondrial DNA SNP markers yielded 10 different haplotypes, ranging from 2 to 5 within populations. Populations were highly differentiated (G ST = 0.407, P < 0.001). These new genetic resources will facilitate future parentage, population genetics and phylogeographical studies and will be useful for both evolutionary and conservation concerns, especially for the set-up of management strategies and the definition of distinct evolutionary significant units.

opencc-zeroDec 2015View details →
dryad32/100

Data from: The evolution of heat shock protein sequences, cis-regulatory elements, and expression profiles in the eusocial Hymenoptera

Background: The eusocial Hymenoptera have radiated across a wide range of thermal environments, exposing them to significant physiological stressors. We reconstructed the evolutionary history of three families of Heat Shock Proteins (Hsp90, Hsp70, Hsp40), the primary molecular chaperones protecting against thermal damage, across 12 Hymenopteran species and four other insect orders. We also predicted and tested for thermal inducibility of eight Hsps from the presence of cis-regulatory heat shock elements (HSEs). We tested whether Hsp induction patterns in ants were associated with different thermal environments. Results: We found evidence for duplications, losses, and cis-regulatory changes in two of the three gene families. One member of the Hsp90 gene family, hsp83, duplicated basally in the Hymenoptera, with shifts in HSE motifs in the novel copy. Both copies were retained in bees, but ants retained only the novel HSE copy. For Hsp70, Hymenoptera lack the primary heat-inducible orthologue from Drosophila melanogaster and instead induce the cognate form, hsc70-4, which also underwent an early duplication. Episodic diversifying selection was detected along the branch predating the duplication of hsc70-4 and continued along one of the paralogue branches after duplication. Four out of eight Hsp genes were heat-inducible and matched the predictions based on presence of conserved HSEs. For the inducible homologues, the more thermally tolerant species, Pogonomyrmex barbatus, had greater Hsp basal expression and induction in response to heat stress than did the less thermally tolerant species, Aphaenogaster picea. Furthermore, there was no trade-off between basal expression and induction. Conclusions: Our results highlight the unique evolutionary history of Hsps in eusocial Hymenoptera, which has been shaped by gains, losses, and changes in cis-regulation. Ants, and most likely other Hymenoptera, utilize lineage-specific heat inducible Hsps, whose expression patterns are associated with adaptive variation in thermal tolerance between two ant species. Collectively, our analyses suggest that Hsp sequence and expression patterns may reflect the forces of selection acting on thermal tolerance in ants and other social Hymenoptera.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Population structure, relatedness and ploidy levels in an apple gene bank revealed through genotyping-by-sequencing

In recent years, new genome-wide marker systems have provided highly informative alternatives to low density marker systems for evaluating plant populations. To date, most apple germplasm collections have been genotyped using low-density markers such as simple sequence repeats (SSRs), whereas only a few have been explored using high-density genome-wide marker information. We explored the genetic diversity of the Pometum gene bank collection (University of Copenhagen, Denmark) of 349 apple accessions using over 15,000 genome-wide single nucleotide polymorphisms (SNPs) and 15 SSR markers, in order to compare the strength of the two approaches for describing population structure. We found that 119 accessions shared a clonal relationship with at least one other accession in the collection, resulting in the identification of 272 (78%) unique accessions. Of these unique accessions, over half (52%) share a first-degree relationship with at least one other accession. There is therefore a high degree of clonal and family relatedness in the Danish apple gene bank. We find significant genetic differentiation between Malus domestica and its supposed primary wild ancestor, M. sieversii, as well as between accessions of Danish origin and all others. Overall, we found strong concordance between analyses based on the genome-wide SNPs and the 15 SSR loci. However, we argue that GBS is superior to traditional SSR approaches because it allowed the estimation of ploidy levels that were in accordance with flow cytometry results, and can be further exploited in genome-wide association studies (GWAS). Finally, we compare GBS with SSR for the purposes of characterizing a diverse apple gene bank and discuss the advantages and constraints of the two approaches.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Detection of somatic epigenetic variation in Norway spruce via targeted bisulfite sequencing

Epigenetic mechanisms represent a possible mechanism for achieving a rapid response of long‐lived trees to changing environmental conditions. However, our knowledge on plant epigenetics is largely limited to a few model species. With increasing availability of genomic resources for many tree species, it is now possible to adopt approaches from model species that permit to obtain single‐base pair resolution data on methylation at a reasonable cost. Here, we used targeted bisulfite sequencing (TBS) to study methylation patterns in the conifer species Norway spruce (Picea abies). To circumvent the challenge of disentangling epigenetic and genetic differences, we focused on four clone pairs, where clone members were growing in different climatic conditions for 24 years. We targeted >26.000 genes using TBS and determined the performance and reproducibility of this approach. We characterized gene body methylation and compared methylation patterns between environments. We found highly comparable capture efficiency and coverage across libraries. Methylation levels were relatively constant across gene bodies, with 21.3 ± 0.3%, 11.0 ± 0.4% and 1.3 ± 0.2% in the CG, CHG, and CHH context, respectively. The variance in methylation profiles did not reveal consistent changes between environments, yet we could identify 334 differentially methylated positions (DMPs) between environments. This supports that changes in methylation patterns are a possible pathway for a plant to respond to environmental change. After this successful application of TBS in Norway spruce, we are confident that this approach can contribute to broaden our knowledge of methylation patterns in natural tree populations.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure

PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Nuclear and mitochondrial sequence data reveal and conceal different demographic histories and population genetic processes in Caribbean reef fishes

Mitochondrial and nuclear sequence data should recover historical demographic events at different temporal scales due to differences in their effective population sizes and substitution rates. This expectation was tested for two closely related coral reef fish, the tube blennies Acanthemblemaria aspera and A. spinosa. These two have similar life histories and dispersal potentials, and co-occur throughout the Caribbean. Sequence data for one mitochondrial and two nuclear markers were collected for 168 individuals across the species' Caribbean ranges. While both species shared a similar pattern of genetic subdivision, A. spinosa had 20-25-times greater nucleotide sequence divergence among populations than A. aspera at all three markers. Substitution rates estimated using a relaxed clock approach revealed that mitochondrial COI is evolving at 11.2% pairwise sequence divergence per million years. This rapid mitochondrial rate had obscured the signal of old population expansions for both species, which were only recovered using the more slowly evolving nuclear markers. However, the rapid COI rate allowed the recovery of a recent expansion in A. aspera corresponding to a period of increased habitat availability. Only by combining both nuclear and mitochondrial data were we able to recover the complex demographic history of these fishes.

opencc-zeroDec 2009View details →
dryad32/100

Data from: Genotyping-by-sequencing for estimating relatedness in non-model organisms: avoiding the trap of precise bias

There has been remarkably little attention to using the high resolution provided by genotyping-by-sequencing (i.e. RADseq and similar methods) datasets for assessing relatedness in wildlife populations. A major hurdle is the genotyping error, especially allelic dropout, often found in this type of dataset that could lead to downward-biased, yet precise, estimates of relatedness. Here we assess the applicability of genotyping-by-sequencing datasets for relatedness inferences given their relatively high genotyping error rates. Individuals of known relatedness were simulated under genotyping error, allelic dropout, and missing data scenarios based on an empirical ddRAD dataset, and their true relatedness was compared to that estimated by seven relatedness estimators. We found that an estimator chosen through such analyses can circumvent the influence of genotyping error, with the estimator of Ritland (1996) shown to be unaffected by allelic dropout and to be the most accurate when there is genotyping error. We also found that the choice of estimator should not rely solely on the strength of correlation between estimated and true relatedness as a strong correlation does not necessarily mean estimates are close to true relatedness. We also demonstrated how even a large SNP dataset with genotyping error (allelic dropout or otherwise) or missing data still performs better than a perfectly genotyped microsatellite dataset of tens of markers. The simulation-based approach used here can be easily implemented by others on their own genotyping-by-sequencing datasets to confirm the most appropriate and powerful estimator for their dataset.

opencc-zeroDec 2016View details →
dryad32/100

Data from: The complete sequence of the mitochondrial genome of Butomus umbellatus - a member of an early branching lineage of monocotyledons

In order to study the evolution of mitochondrial genomes in the early branching lineages of the monocotyledons, i.e., the Acorales and Alismatales, we are sequencing complete genomes from a suite of key taxa. As a starting point the present paper describes the mitochondrial genome of Butomus umbellatus (Butomaceae) based on next-generation sequencing data. The genome was assembled into a circular molecule, 450,826 bp in length. Coding sequences cover only 8.2% of the genome and include 28 protein coding genes, four rRNA genes, and 12 tRNA genes. Some of the tRNA genes and a 16S rRNA gene are transferred from the plastid genome. However, the total amount of recognized plastid sequences in the mitochondrial genome is only 1.5% and the amount of DNA transferred from the nucleus is also low. RNA editing is abundant and a total of 557 edited sites are predicted in the protein coding genes. Compared to the 40 angiosperm mitochondrial genomes sequenced to date, the GC content of the Butomus genome is uniquely high (49.1%). The overall similarity between the mitochondrial genomes of Butomus and Spirodela (Araceae), the closest relative yet sequenced, is low (less than 20%), and the two genomes differ in size by a factor 2. Gene order is also largely unconserved. However, based on its phylogenetic position within the core alismatids Butomus will serve as a good reference point for subsequent studies in the early branching lineages of the monocotyledons.

opencc-zeroDec 2012View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record