Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.9.0
Dataset results
44 results for “substitution rate”
Data from: Compensatory evolution in RNA secondary structures increases substitution rate variation among sites
There is growing evidence that interactions between biological molecules (e.g., RNA-RNA, protein-protein, RNA-protein) place limits on the rate and trajectory of molecular evolution. Here, by extending Kimura's model of compensatory evolution at interacting sites, we show that the ratio of transition to transversion substitutions (κ) at interacting sites should be equal to the square of the ratio at independent sites. Because transition mutations generally occur at a higher rate than transversions, the model predicts that κ should be higher at interacting sites than at independent sites. We tested this prediction in 10 RNA secondary structures by comparing phylogenetically derived estimates of κ in paired sites within stems (κ(p)) and unpaired sites within loops (κ(u)). Eight of the 10 structures showed an excellent match to the quantitative predictions of the model, and 9 of the 10 structures matched the qualitative prediction κ(p) > κ(u). Only the Rev response element from the human immunovirus (HIV) genome showed the reverse pattern, with κ(p) < κ(u). Although a variety of evolutionary forces could produce quantitative deviations from the model predictions, the reversal in magnitude of κ(p) and κ(u) could be achieved only by violating the model assumption that the underlying transition (or transversion) mutation rates were identical in paired and unpaired regions of the molecule. We explore the ability of the APOBEC3 enzymes, host defense mechanisms against retroviruses, which induce transition mutations preferentially in single-stranded regions of the HIV genome, to explain this exception to the rule. Taken as a whole, our findings suggest that kappa may have utility as a simple diagnostic to evaluate proposed secondary structures.
Data from: SNP discovery in non-model organisms: strand-bias and base-substitution errors reduce conversion rates
Single nucleotide polymorphisms (SNPs) have become the marker of choice for genetic studies in organisms of conservation, commercial or biological interest. Most SNP discovery projects in nonmodel organisms apply a strategy for identifying putative SNPs based on filtering rules that account for random sequencing errors. Here, we analyse data used to develop 4723 novel SNPs for the commercially important deep-sea fish, orange roughy (Hoplostethus atlanticus), to assess the impact of not accounting for systematic sequencing errors when filtering identified polymorphisms when discovering SNPs. We used SAMtools to identify polymorphisms in a velvet assembly of genomic DNA sequence data from seven individuals. The resulting set of polymorphisms were filtered to minimize 'bycatch'—polymorphisms caused by sequencing or assembly error. An Illumina Infinium SNP chip was used to genotype a final set of 7714 polymorphisms across 1734 individuals. Five predictors were examined for their effect on the probability of obtaining an assayable SNP: depth of coverage, number of reads that support a variant, polymorphism type (e.g. A/C), strand-bias and Illumina SNP probe design score. Our results indicate that filtering out systematic sequencing errors could substantially improve the efficiency of SNP discovery. We show that BLASTX can be used as an efficient tool to identify single-copy genomic regions in the absence of a reference genome. The results have implications for research aiming to identify assayable SNPs and build SNP genotyping assays for nonmodel organisms.
Data from: Gene tree discordance causes apparent substitution rate variation
Substitution rates are known to be variable among genes, chromosomes, species, and lineages due to multifarious biological processes. Here, we consider another source of substitution rate variation due to a technical bias associated with gene tree discordance. Discordance has been found to be rampant in genome-wide data sets, often due to incomplete lineage sorting (ILS). This apparent substitution rate variation is caused when substitutions that occur on discordant gene trees are analyzed in the context of a single, fixed species tree. Such substitutions have to be resolved by proposing multiple substitutions on the species tree, and we therefore refer to this phenomenon as Substitutions Produced by ILS (SPILS). We use simulations to demonstrate that SPILS has a larger effect with increasing levels of ILS, and on trees with larger numbers of taxa. Specific branches of the species trees are consistently, but erroneously, inferred to be longer or shorter, and we show that these branches can be predicted based on discordant tree topologies. Moreover, we observe that fixing a species tree topology when performing tests of positive selection increases the false positive rate, particularly for genes whose discordant topologies are most affected by SPILS. Finally, we use data from multiple Drosophila species to show that SPILS can be detected in nature. Although the effects of SPILS are modest per gene, it has the potential to affect substitution rate variation whenever high levels of ILS are present, particularly in rapid radiations. The problems outlined here have implications for character mapping of any type of trait, and for any biological process that causes discordance. We discuss possible solutions to these problems, and areas in which they are likely to have caused faulty inferences of convergence and accelerated evolution.
Data from: Effects of genotype on rates of substitution during experimental evolution
Rates of molecular evolution may vary widely between populations, yet the causes of this variation are still incompletely understood. Genetic differences between populations may make an important contribution to variation in rates of evolution, owing to differences in fitness, population size, mutation rates, or in the distribution of fitness effects (DFE) of available beneficial mutations. By whole genome sequencing of Escherichia coli populations experimentally evolved in the presence of a quinolone antibiotic, we found that rates of substitution varied by genotype, with evidence for a contribution from a genotype's starting fitness. Subsequent targeted sequencing showed that genotypes with high average substitution rates were more likely to undergo the simultaneous fixation of several mutations, consistent with theoretical models of multiple mutation dynamics. Moreover, patterns of substitution were indicative of epistatic relationships between known resistance mutations.
Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods
The increasing ability to extract and sequence DNA from non-contemporaneous tissue offers biologists the opportunity to analyze ancient DNA (aDNA) together with modern DNA (mDNA) to address the taxonomy of extinct species, evolutionary origins, historical phylogeography and biogeography. Perhaps more exciting are recent developments in coalescence-based Bayesian inference that offer the potential to use temporal information from aDNA and mDNA for the estimation of substitution rates and divergence dates as an alternative to fossil and geological calibration. This comes at a time of growing interest in the possibility of time dependency for molecular rate estimates. Here we provide a critical assessment of Bayesian MCMC analysis for the estimation of substitution rate using simulated samples of aDNA and mDNA. We conclude that the current models and priors employed in Bayesian MCMC analysis of heterochronous mtDNA are susceptible to an upward bias in the estimation of substitution rates due to model misspecification when the data comes from populations with less than simple demographic histories, including sudden short-lived population bottlenecks or pronounced population structure. However when model misspecification is only mild, then the 95% HPD intervals provide adequate frequentist coverage of the true rates.
Data from: Overestimation of the adaptive substitution rate in fluctuating populations
Estimating the proportion of adaptive substitutions (α) is of primary importance to uncover the determinants of adaptation in comparative genomic studies. Several methods have been proposed to estimate α from patterns polymorphism and divergence in coding sequences. However, estimators of α can be biased when the underlying assumptions are not met. Here we focus on a potential source of bias, i.e., variation through time in the long term population size (N) of the considered species. We show via simulations that ancient demographic fluctuations can generate severe overestimations of α, and this irrespective of the recent population history.
Data from: Accurate estimation of substitution rates with neighbour-dependent models in a phylogenetic context
Most models and algorithms developed to perform statistical inference from DNA data make the assumption that substitution processes affecting distinct nucleotide sites are stochastically independent. This assumption ensures both mathematical and computational tractability, but is in disagreement with observed data in many situations -- one well-known example being CpG dinucleotide hypermutability in mammalian genomes. In this paper, we consider the class of RN95+YpR substitution models, which allows neighbour-dependent effects -- including CpG hypermutability -- to be taken into account, through transitions between pyrimidine-purine dinucleotides. We show that it is possible to adapt inference methods originally developed under the assumption of independence between sites to RN95+YpR models, using a mathematically rigorous framework provided by specific structural properties of this class of models. We assess how efficient this approach is at inferring the CpG hypermutability rate from aligned DNA sequences. The method is tested on simulated data and compared against several alternatives; the results suggest that it delivers a high degree of accuracy at a low computational cost. We then apply our method to an alignment of ten DNA sequences from primate species. Model comparisons within the RN95+YpR class show the importance of taking into account neighbour-dependent effects. An application of the method to the detection of hypomethylated islands is discussed.
Data from: Sequencing of the needle transcriptome from Norway spruce (Picea abies Karst L.) reveals lower substitution rates, but similar selective constraints in gymnosperms and angiosperms
BACKGROUND: A detailed knowledge about spatial and temporal gene expression is important for understanding both the function of genes and their evolution. For the vast majority of species, transcriptomes are still largely uncharacterized and even in those where substantial information is available it is often in the form of partially sequenced transcriptomes. With the development of next generation sequencing, a single experiment can now simultaneously identify the transcribed part of a species genome and estimate levels of gene expression. RESULTS: mRNA from actively growing needles of Norway spruce (Picea abies) was sequenced using next generation sequencing technology. In total, close to 70 million fragments with a length of 76 bp were sequenced resulting in 5 Gbp of raw data. A de novo assembly of these reads, together with publicly available expressed sequence tag (EST) data from Norway spruce, was used to create a reference transcriptome. Of the 38,419 PUTs (putative unique transcripts) longer than 150 bp in this reference assembly, 83.5% show similarity to ESTs from other spruce species and of the remaining PUTs, 3,704 show similarity to protein sequences from other plant species, leaving 4,167 PUTs with limited similarity to currently available plant proteins. By predicting coding frames and comparing not only the Norway spruce PUTs, but also PUTs from the close relatives Picea glauca and Picea sitchensis to both Pinus taeda and Taxus mairei, we obtained estimates of synonymous and non-synonymous divergence among conifer species. In addition, we detected close to 15,000 SNPs of high quality and estimated gene expression differences between samples collected under dark and light conditions. CONCLUSIONS: Our study yielded a large number of single nucleotide polymorphisms as well as estimates of gene expression on transcriptome scale. In agreement with a recent study we find that the synonymous substitution rate per year (0.6 x 10-09 and 1.1 x 10-09) is an order of magnitude smaller than values reported for angiosperm herbs. However, if one takes generation time into account, most of this difference disappears. The estimates of the dN/dS ratio (non-synonymous over synonymous divergence) reported here are in general much lower than 1 and only a few genes showed a ratio larger than 1.
Data from: Sequence entropy of folding and the absolute rate of amino acid substitutions
Adequate representations of protein evolution should consider how the acceptance of mutations depends on the sequence context in which they arise. However, epistatic interactions among sites in a protein result in hererogeneities in the substitution rate, both temporal and spatial, that are beyond the capabilities of current models. Here we use parallels between amino acid substitutions and chemical reaction kinetics to develop an improved theory of protein evolution. We constructed a mechanistic framework for modelling amino acid substitution rates that uses the formalisms of statistical mechanics, with principles of population genetics underlying the analysis. Theoretical analyses and computer simulations of proteins under purifying selection for thermodynamic stability show that substitution rates and the stabilization of resident amino acids (the 'evolutionary Stokes shift') can be predicted from biophysics and the effect of sequence entropy alone. Furthermore, we demonstrate that substitutions predominantly occur when epistatic interactions result in near neutrality; substitution rates are determined by how often epistasis results in such nearly neutral conditions. This theory provides a general framework for modelling protein sequence change under purifying selection, potentially explains patterns of convergence and mutation rates in real proteins that are incompatible with previous models, and provides a better null model for the detection of adaptive changes.
Data from: Elevated substitution rates estimated from ancient DNA sequences
Open the record for dataset details and reuse information.
Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods
Open the record for dataset details and reuse information.
Data from: Compensatory evolution in RNA secondary structures increases substitution rate variation among sites
Open the record for dataset details and reuse information.
Data from: Gene tree discordance causes apparent substitution rate variation
Open the record for dataset details and reuse information.
Data from: Effects of genotype on rates of substitution during experimental evolution
Open the record for dataset details and reuse information.
Data from: The origin of modern frogs (Neobatrachia) was accompanied by acceleration in mitochondrial and nuclear substitution rates
Open the record for dataset details and reuse information.
Data from: A new hierarchy of phylogenetic models consistent with heterogeneous substitution rates
Open the record for dataset details and reuse information.
Data from: Cell tropism predicts long-term nucleotide substitution rates of mammalian RNA viruses
Open the record for dataset details and reuse information.
Data from: Accurate estimation of substitution rates with neighbour-dependent models in a phylogenetic context
Open the record for dataset details and reuse information.
Data from: Sequence entropy of folding and the absolute rate of amino acid substitutions
Open the record for dataset details and reuse information.
Data from: Overestimation of the adaptive substitution rate in fluctuating populations
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.