Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

44

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

44 results for “substitution rate”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Compensatory evolution in RNA secondary structures increases substitution rate variation among sites

There is growing evidence that interactions between biological molecules (e.g., RNA-RNA, protein-protein, RNA-protein) place limits on the rate and trajectory of molecular evolution. Here, by extending Kimura's model of compensatory evolution at interacting sites, we show that the ratio of transition to transversion substitutions (κ) at interacting sites should be equal to the square of the ratio at independent sites. Because transition mutations generally occur at a higher rate than transversions, the model predicts that κ should be higher at interacting sites than at independent sites. We tested this prediction in 10 RNA secondary structures by comparing phylogenetically derived estimates of κ in paired sites within stems (κ(p)) and unpaired sites within loops (κ(u)). Eight of the 10 structures showed an excellent match to the quantitative predictions of the model, and 9 of the 10 structures matched the qualitative prediction κ(p) > κ(u). Only the Rev response element from the human immunovirus (HIV) genome showed the reverse pattern, with κ(p) < κ(u). Although a variety of evolutionary forces could produce quantitative deviations from the model predictions, the reversal in magnitude of κ(p) and κ(u) could be achieved only by violating the model assumption that the underlying transition (or transversion) mutation rates were identical in paired and unpaired regions of the molecule. We explore the ability of the APOBEC3 enzymes, host defense mechanisms against retroviruses, which induce transition mutations preferentially in single-stranded regions of the HIV genome, to explain this exception to the rule. Taken as a whole, our findings suggest that kappa may have utility as a simple diagnostic to evaluate proposed secondary structures.

opencc-zeroDec 2007View details →
dryad28/100

Data from: SNP discovery in non-model organisms: strand-bias and base-substitution errors reduce conversion rates

Single nucleotide polymorphisms (SNPs) have become the marker of choice for genetic studies in organisms of conservation, commercial or biological interest. Most SNP discovery projects in nonmodel organisms apply a strategy for identifying putative SNPs based on filtering rules that account for random sequencing errors. Here, we analyse data used to develop 4723 novel SNPs for the commercially important deep-sea fish, orange roughy (Hoplostethus atlanticus), to assess the impact of not accounting for systematic sequencing errors when filtering identified polymorphisms when discovering SNPs. We used SAMtools to identify polymorphisms in a velvet assembly of genomic DNA sequence data from seven individuals. The resulting set of polymorphisms were filtered to minimize 'bycatch'—polymorphisms caused by sequencing or assembly error. An Illumina Infinium SNP chip was used to genotype a final set of 7714 polymorphisms across 1734 individuals. Five predictors were examined for their effect on the probability of obtaining an assayable SNP: depth of coverage, number of reads that support a variant, polymorphism type (e.g. A/C), strand-bias and Illumina SNP probe design score. Our results indicate that filtering out systematic sequencing errors could substantially improve the efficiency of SNP discovery. We show that BLASTX can be used as an efficient tool to identify single-copy genomic regions in the absence of a reference genome. The results have implications for research aiming to identify assayable SNPs and build SNP genotyping assays for nonmodel organisms.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Gene tree discordance causes apparent substitution rate variation

Substitution rates are known to be variable among genes, chromosomes, species, and lineages due to multifarious biological processes. Here, we consider another source of substitution rate variation due to a technical bias associated with gene tree discordance. Discordance has been found to be rampant in genome-wide data sets, often due to incomplete lineage sorting (ILS). This apparent substitution rate variation is caused when substitutions that occur on discordant gene trees are analyzed in the context of a single, fixed species tree. Such substitutions have to be resolved by proposing multiple substitutions on the species tree, and we therefore refer to this phenomenon as Substitutions Produced by ILS (SPILS). We use simulations to demonstrate that SPILS has a larger effect with increasing levels of ILS, and on trees with larger numbers of taxa. Specific branches of the species trees are consistently, but erroneously, inferred to be longer or shorter, and we show that these branches can be predicted based on discordant tree topologies. Moreover, we observe that fixing a species tree topology when performing tests of positive selection increases the false positive rate, particularly for genes whose discordant topologies are most affected by SPILS. Finally, we use data from multiple Drosophila species to show that SPILS can be detected in nature. Although the effects of SPILS are modest per gene, it has the potential to affect substitution rate variation whenever high levels of ILS are present, particularly in rapid radiations. The problems outlined here have implications for character mapping of any type of trait, and for any biological process that causes discordance. We discuss possible solutions to these problems, and areas in which they are likely to have caused faulty inferences of convergence and accelerated evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Effects of genotype on rates of substitution during experimental evolution

Rates of molecular evolution may vary widely between populations, yet the causes of this variation are still incompletely understood. Genetic differences between populations may make an important contribution to variation in rates of evolution, owing to differences in fitness, population size, mutation rates, or in the distribution of fitness effects (DFE) of available beneficial mutations. By whole genome sequencing of Escherichia coli populations experimentally evolved in the presence of a quinolone antibiotic, we found that rates of substitution varied by genotype, with evidence for a contribution from a genotype's starting fitness. Subsequent targeted sequencing showed that genotypes with high average substitution rates were more likely to undergo the simultaneous fixation of several mutations, consistent with theoretical models of multiple mutation dynamics. Moreover, patterns of substitution were indicative of epistatic relationships between known resistance mutations.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods

The increasing ability to extract and sequence DNA from non-contemporaneous tissue offers biologists the opportunity to analyze ancient DNA (aDNA) together with modern DNA (mDNA) to address the taxonomy of extinct species, evolutionary origins, historical phylogeography and biogeography. Perhaps more exciting are recent developments in coalescence-based Bayesian inference that offer the potential to use temporal information from aDNA and mDNA for the estimation of substitution rates and divergence dates as an alternative to fossil and geological calibration. This comes at a time of growing interest in the possibility of time dependency for molecular rate estimates. Here we provide a critical assessment of Bayesian MCMC analysis for the estimation of substitution rate using simulated samples of aDNA and mDNA. We conclude that the current models and priors employed in Bayesian MCMC analysis of heterochronous mtDNA are susceptible to an upward bias in the estimation of substitution rates due to model misspecification when the data comes from populations with less than simple demographic histories, including sudden short-lived population bottlenecks or pronounced population structure. However when model misspecification is only mild, then the 95% HPD intervals provide adequate frequentist coverage of the true rates.

opencc-zeroDec 2009View details →
dryad28/100

Data from: Overestimation of the adaptive substitution rate in fluctuating populations

Estimating the proportion of adaptive substitutions (α) is of primary importance to uncover the determinants of adaptation in comparative genomic studies. Several methods have been proposed to estimate α from patterns polymorphism and divergence in coding sequences. However, estimators of α can be biased when the underlying assumptions are not met. Here we focus on a potential source of bias, i.e., variation through time in the long term population size (N) of the considered species. We show via simulations that ancient demographic fluctuations can generate severe overestimations of α, and this irrespective of the recent population history.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Accurate estimation of substitution rates with neighbour-dependent models in a phylogenetic context

Most models and algorithms developed to perform statistical inference from DNA data make the assumption that substitution processes affecting distinct nucleotide sites are stochastically independent. This assumption ensures both mathematical and computational tractability, but is in disagreement with observed data in many situations -- one well-known example being CpG dinucleotide hypermutability in mammalian genomes. In this paper, we consider the class of RN95+YpR substitution models, which allows neighbour-dependent effects -- including CpG hypermutability -- to be taken into account, through transitions between pyrimidine-purine dinucleotides. We show that it is possible to adapt inference methods originally developed under the assumption of independence between sites to RN95+YpR models, using a mathematically rigorous framework provided by specific structural properties of this class of models. We assess how efficient this approach is at inferring the CpG hypermutability rate from aligned DNA sequences. The method is tested on simulated data and compared against several alternatives; the results suggest that it delivers a high degree of accuracy at a low computational cost. We then apply our method to an alignment of ten DNA sequences from primate species. Model comparisons within the RN95+YpR class show the importance of taking into account neighbour-dependent effects. An application of the method to the detection of hypomethylated islands is discussed.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Sequencing of the needle transcriptome from Norway spruce (Picea abies Karst L.) reveals lower substitution rates, but similar selective constraints in gymnosperms and angiosperms

BACKGROUND: A detailed knowledge about spatial and temporal gene expression is important for understanding both the function of genes and their evolution. For the vast majority of species, transcriptomes are still largely uncharacterized and even in those where substantial information is available it is often in the form of partially sequenced transcriptomes. With the development of next generation sequencing, a single experiment can now simultaneously identify the transcribed part of a species genome and estimate levels of gene expression. RESULTS: mRNA from actively growing needles of Norway spruce (Picea abies) was sequenced using next generation sequencing technology. In total, close to 70 million fragments with a length of 76 bp were sequenced resulting in 5 Gbp of raw data. A de novo assembly of these reads, together with publicly available expressed sequence tag (EST) data from Norway spruce, was used to create a reference transcriptome. Of the 38,419 PUTs (putative unique transcripts) longer than 150 bp in this reference assembly, 83.5% show similarity to ESTs from other spruce species and of the remaining PUTs, 3,704 show similarity to protein sequences from other plant species, leaving 4,167 PUTs with limited similarity to currently available plant proteins. By predicting coding frames and comparing not only the Norway spruce PUTs, but also PUTs from the close relatives Picea glauca and Picea sitchensis to both Pinus taeda and Taxus mairei, we obtained estimates of synonymous and non-synonymous divergence among conifer species. In addition, we detected close to 15,000 SNPs of high quality and estimated gene expression differences between samples collected under dark and light conditions. CONCLUSIONS: Our study yielded a large number of single nucleotide polymorphisms as well as estimates of gene expression on transcriptome scale. In agreement with a recent study we find that the synonymous substitution rate per year (0.6 x 10-09 and 1.1 x 10-09) is an order of magnitude smaller than values reported for angiosperm herbs. However, if one takes generation time into account, most of this difference disappears. The estimates of the dN/dS ratio (non-synonymous over synonymous divergence) reported here are in general much lower than 1 and only a few genes showed a ratio larger than 1.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Sequence entropy of folding and the absolute rate of amino acid substitutions

Adequate representations of protein evolution should consider how the acceptance of mutations depends on the sequence context in which they arise. However, epistatic interactions among sites in a protein result in hererogeneities in the substitution rate, both temporal and spatial, that are beyond the capabilities of current models. Here we use parallels between amino acid substitutions and chemical reaction kinetics to develop an improved theory of protein evolution. We constructed a mechanistic framework for modelling amino acid substitution rates that uses the formalisms of statistical mechanics, with principles of population genetics underlying the analysis. Theoretical analyses and computer simulations of proteins under purifying selection for thermodynamic stability show that substitution rates and the stabilization of resident amino acids (the 'evolutionary Stokes shift') can be predicted from biophysics and the effect of sequence entropy alone. Furthermore, we demonstrate that substitutions predominantly occur when epistatic interactions result in near neutrality; substitution rates are determined by how often epistasis results in such nearly neutral conditions. This theory provides a general framework for modelling protein sequence change under purifying selection, potentially explains patterns of convergence and mutation rates in real proteins that are incompatible with previous models, and provides a better null model for the detection of adaptive changes.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Elevated substitution rates estimated from ancient DNA sequences

Open the record for dataset details and reuse information.

publicApr 2008View details →
dryad28/100

Data from: Elevated substitution rate estimates from ancient DNA: model violation and bias of Bayesian methods

Open the record for dataset details and reuse information.

publicMar 2010View details →
dryad28/100

Data from: Compensatory evolution in RNA secondary structures increases substitution rate variation among sites

Open the record for dataset details and reuse information.

publicJun 2008View details →
dryad28/100

Data from: Gene tree discordance causes apparent substitution rate variation

Open the record for dataset details and reuse information.

publicFeb 2016View details →
dryad28/100

Data from: Effects of genotype on rates of substitution during experimental evolution

Open the record for dataset details and reuse information.

publicJun 2015View details →
dryad28/100

Data from: The origin of modern frogs (Neobatrachia) was accompanied by acceleration in mitochondrial and nuclear substitution rates

Open the record for dataset details and reuse information.

publicNov 2012View details →
dryad28/100

Data from: A new hierarchy of phylogenetic models consistent with heterogeneous substitution rates

Open the record for dataset details and reuse information.

publicApr 2015View details →
dryad28/100

Data from: Cell tropism predicts long-term nucleotide substitution rates of mammalian RNA viruses

Open the record for dataset details and reuse information.

publicDec 2014View details →
dryad28/100

Data from: Accurate estimation of substitution rates with neighbour-dependent models in a phylogenetic context

Open the record for dataset details and reuse information.

publicJan 2012View details →
dryad28/100

Data from: Sequence entropy of folding and the absolute rate of amino acid substitutions

Open the record for dataset details and reuse information.

publicAug 2018View details →
dryad28/100

Data from: Overestimation of the adaptive substitution rate in fluctuating populations

Open the record for dataset details and reuse information.

publicApr 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record