Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
210
datasets available to search
ShareScore release 0.7.1
Dataset results
210 results for “Bayesian inference”
Data from: From fossils to phylogenies: Exploring the integration of paleontological data into Bayesian phylogenetic inference
Open the record for dataset details and reuse information.
Data from: Phylogeographic inference using Bayesian model comparison across a fragmented chorus frog species complex
Open the record for dataset details and reuse information.
Data from: Exact Bayesian inference for animal movement in continuous time
Open the record for dataset details and reuse information.
A Multi-Type Birth-Death model for Bayesian inference of lineage-specific birth and death rates
Open the record for dataset details and reuse information.
Inferring macroevolutionary parameters from on a model of adaptive radiation by Approximate Bayesian Computation - ABC Simulations
<p>Recent advances in DNA sequencing are providing increasingly accurate phylogenetic trees to study, but understanding the evolutionary forces at play in different contexts remains a huge challenge. To tackle this issue, we applied an Bayesian approach to an existing model of phenotypic and species diversification [1] in order to retrieve 11 underlying parameters (such as basal speciation and extinction rates, but also competition strength) from phylogenetic trees with known traits values at the tips.</p> <p>This dataset corresponds to the Approximate Bayesian Computation simulations realized for the inference.</p>
FIG. 1. — Phylogenetic tree using Bayesian inference with mitochondrial 16S in A review of the genus Coccoglypta Pilsbry, 1895 (Gastropoda: Pulmonata: Camaenidae)
FIG. 1. — Phylogenetic tree using Bayesian inference with mitochondrial 16S and CO1 genes. Numbers above or below branches indicate Bayesian posterior probabilities.
[dataset] Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference
<p>Anonymized replication package of the experiments presented in the research paper: "Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference".</p> <p>See the README.md file.</p>
Bayesian inference of tree species using diffusion models: tabulated posterior statistics for SNAPP and SNAPPER analyses
<p>We describe a new and computationally efficient Bayesian methodology for inferring species trees and demographics from unlinked binary markers. Likelihood calculations are carried out using diffusion models of allele frequency dynamics combined with novel numerical algorithms. The diffusion approach allows for analysis of datasets containing hundreds or thousands of individuals. The method, which we call \snapper, has been implemented as part of the BEAST2 package. We conducted simulation experiments to assess numerical error, computational requirements and accuracy recovering known model parameters. A re-analysis of soybean SNP data demonstrates that the models implemented in \snapp and \snapper can be difficult to distinguish in practice, a characteristic which we tested with further simulations. We demonstrate the scale of analysis possible using a SNP dataset sampled from 399 fresh water turtles in 41 populations.</p>
Hierarchical Inference With Bayesian Neural Networks: An Application to Strong Gravitational Lensing - Model Weights, Chains, BNN Samples, and Simulated Datasets
<p>The model weights, chains, simulated datasets, and BNN samples used to produce the results shown in LSST DESC Collaboration paper "Hierarchical Inference With Bayesian Neural Networks: An Application to Strong Gravitational Lensing." All files presented here are meant for use in tandem with the python package "ovejero" (<a href="https://github.com/swagnercarena/ovejero">https://github.com/swagnercarena/ovejero</a>).</p>
Data from: Bayesian species delimitation can be robust to guide tree inference errors
The Bayesian method of species delimitation (Yang and Rannala, 2010) uses a so-called guide tree to reduce the number of models to be evaluated in the reversible-jump Markov chain Monte Carlo (rjMCMC) algorithm (Green, 1995). It has been pointed out that the method tends to over-split if a random population tree is used as the guide tree (Fujita and Leaché, 2011). Here we conduct a simulation study to examine the performance of the method under more realistic scenarios, that is, when the guide tree is inferred from the sequence data. We found that Bayesian species delimitation is in general robust to errors in the inferred guide tree.
Data from: Bayesian inference of selection in a heterogeneous environment from genetic time-series data
Evolutionary geneticists have sought to characterize the causes and molecular targets of selection in natural populations for many years. Although this research program has been somewhat successful, most statistical methods employed were designed to detect consistent, weak to moderate selection. In contrast, phenotypic studies in nature show that selection varies in time and that individual bouts of selection can be strong. Measurements of the genomic consequences of such fluctuating selection could help test and refine hypotheses concerning the causes of ecological specialization and the maintenance of genetic variation in populations. Herein, I proposed a Bayesian non-homogenous hidden Markov model to estimate effective population sizes and quantify variable selection in heterogeneous environments from genetic time-series data. The model is described and then evaluated using a series of simulated data, including cases where selection occurs on a trait with a simple or polygenic molecular basis. The proposed method accurately distinguished neutral loci from non-neutral loci under strong selection, but not from those under weak selection. Selection coefficients were accurately estimated when selection was constant or when the fitness values of genotypes varied linearly with the environment, but these estimates were less accurate when fitness was polygenic or the relationship between the environment and the fitness of genotypes was non-linear. Past studies of temporal evolutionary dynamics in lab populations have been remarkably successful. The proposed method makes similar analyses of genetic time-series data from natural populations more feasible, and thereby could help answer fun damental questions about the causes and consequences of evolution in the wild.
Data from: Bayesian model selection with BAMM: effects of the model prior on the inferred number of diversification shifts
1. Understanding variation in rates of speciation and extinction -- both among lineages and through time -- is critical to the testing of many hypotheses about macroevolutionary processes. BAMM is a flexible Bayesian framework for inferring the number and location of shifts in macroevolutionary rate across phylogenetic trees and has been widely used in empirical studies. BAMM requires that researchers specify a prior probability distribution on the number of diversification rate shifts before conducting an analysis. The consequences of this "model prior" for inference are poorly known but could potentially influence both the probability of accepting models that are more (high error rate) or less (low power) complex than the generating model. 2. The hierarchical Poisson process prior in BAMM reduces to a simple geometric distribution on number of rate shifts and we use this property to increase the efficiency of model selection with Bayes factors. Using BAMM v2.5, we analyzed phylogenies simulated with and without diversification heterogeneity across a broad range of prior parameterizations. We also assessed the impact of the model prior on MCMC convergence times and on diversification rate estimates. 3. For all simulation scenarios, model evidence (Bayes factor support) for the number of shifts is not sensitive to the choice of model prior over the wide range examined here. The best-supported model found using BAMM rarely includes spurious shifts (<2% of all runs) when diversification models are selected using Bayes factors. BAMM was reliably able to infer the true number of diversification rate shifts across prior expectations that varied by three orders of magnitude. However, we find a strong effect of model prior on MCMC convergence properties: a flatter prior distribution (larger expected number of shifts) can dramatically increase the efficiency of the MCMC simulation. 4. Our results support the use of a liberal model prior in BAMM, as it reduces computation time without distorting the evidence for rate heterogeneity.
Data from: Bayesian inference of reticulate phylogenies under the multispecies network coalescent
The multispecies coalescent (MSC) is a statistical framework that models how gene genealogies grow within the branches of a species tree. The field of computational phylogenetics has witnessed an explosion in the development of methods for species tree inference under MSC, owing mainly to the accumulating evidence of incomplete lineage sorting in phylogenomic analyses. However, the evolutionary history of a set of genomes, or species, could be reticulate due to the occurrence of evolutionary processes such as hybridization or horizontal gene transfer. We report on a novel method for Bayesian inference of genome and species phylogenies under the multispecies network coalescent (MSNC). This framework models gene evolution within the branches of a phylogenetic network, thus incorporating reticulate evolutionary processes, such as hybridization, in addition to incomplete lineage sorting. As phylogenetic networks with different numbers of reticulation events correspond to points of different dimensions in the space of models, we devise a reversible-jump Markov chain Monte Carlo (RJMCMC) technique for sampling the posterior distribution of phylogenetic networks under MSNC. We implemented the methods in the publicly available, open-source software package PhyloNet and studied their performance on simulated and biological data. The work extends the reach of Bayesian inference to phylogenetic networks and enables new evolutionary analyses that account for reticulation.
Data from: Modeling the perception of audiovisual distance: Bayesian causal inference and other models
Studies of audiovisual perception of distance are rare. Here, visual and auditory cue interactions in distance are tested against several multisensory models, including a modified causal inference model. In this causal inference model predictions of estimate distributions are included. In our study, the audiovisual perception of distance was overall better explained by Bayesian causal inference than by other traditional models, such as sensory dominance and mandatory integration, and no interaction. Causal inference resolved with probability matching yielded the best fit to the data. Finally, we propose that sensory weights can also be estimated from causal inference. The analysis of the sensory weights allows us to obtain windows within which there is an interaction between the audiovisual stimuli. We find that the visual stimulus always contributes by more than 80% to the perception of visual distance. The visual stimulus also contributes by more than 50% to the perception of auditory distance, but only within a mobile window of interaction, which ranges from 1 to 4 m.
FIGURE 1 in Bayesian inference reveals a complex evolutionary history of belemnites
FIGURE 1. Terminology of some of the characters applied coded for the analysis (modified after Mutterlose, 1983; Doyle, 1990; Schlegelmilch, 1998; Hoffmann and Stevens, 2020).
Figure 3. Bayesian dating tree inference performed with cytochrome b in PhylOgeOgraphy and pOtential distributiOn OF Sturnira lilium and S. giannae (ChirOptera: PhyllOstOmidae) With range eXtensiOn FOr S. giannae in the CerradO and Pantanal biOmes
Figure 3. Bayesian dating tree inference performed with cytochrome b gene for Sturnira. Numbers are the nodes ages and values of posterior probability are represented by circles in black (pp ≥ 0.9) and white (pp ≥ 0.8 <0.9). See the list of haplotypes in Table S1 and Fig. S1.
Fig. 3. Bayesian Inference tree constructed from COX1 in Ecological and geographical speciation in Lucilia bufonivora: The evolution of amphibian obligate parasitism
Fig. 3. Bayesian Inference tree constructed from COX1 (mtDNA) sequence data. Each specimen is labelled with the species name and location abbreviation as indicated in Table 1. Sequences obtained from BOLD/GenBank are also annotated with their respective accession codes. Green text corresponds to European samples of Lucilia bufonivora; red represents Lucilia elongata; purple represents Canadian L. bufonivora; orange represents Lucilia silvarum. Scale bar represents expected changes per site. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 6 Bayesian Inference relationships obtained for all the trnD1–4 in An exceptional case of mitochondrial tRNA duplication-deletion events in blood-feeding leeches
Fig. 6 Bayesian Inference relationships obtained for all the trnD1–4 sequences known to date
Data from: A new method of Bayesian causal inference in non-stationary environments
Bayesian inference is the process of narrowing down the hypotheses (causes) to the one that best explains the observational data (effects). To accurately estimate a cause, a considerable amount of data is required to be observed for as long as possible. However, the object of inference is not always constant. In this case, a method such as exponential moving average (EMA) with a discounting rate is used to improve the ability to respond to a sudden change; it is also necessary to increase the discounting rate. That is, a trade-off is established in which the followability is improved by increasing the discounting rate, but the accuracy is reduced. Here, we propose an extended Bayesian inference (EBI), wherein human-like causal inference is incorporated. We show that both the learning and forgetting effects are introduced into Bayesian inference by incorporating the causal inference. We evaluate the estimation performance of the EBI through the learning task of a dynamically changing Gaussian mixture model. In the evaluation, the EBI performance is compared with those of the EMA and a sequential discounting expectation-maximization algorithm. The EBI was shown to modify the trade-off observed in the EMA.
Data from: Inferring the origin of populations introduced from a genetically structured native range by approximate Bayesian computation: case study of the invasive ladybird Harmonia axyridis
Correct identification of the source population of an invasive species is a prerequisite for testing hypotheses concerning the factors responsible for biological invasions. The native area of invasive species may be large, poorly known and/or genetically structured. Because the actual source population may not have been sampled, studies based on molecular markers may generate incorrect conclusions about the origin of introduced populations. In this study, we characterized the genetic structure of the invasive ladybird Harmonia axyridis in its native area using various population genetic statistics and methods. We found that H. axyridis native area most likely consisted of two geographically distinct genetic clusters located in eastern and western Asia. We then performed approximate Bayesian computation (ABC) analyses on controlled simulated microsatellite data sets to evaluate: (i) the risk of selecting incorrect introduction scenarios, including admixture between sources, when the populations of the native area are genetically structured and sampling is incomplete, (ii) the ability of ABC analysis to minimize such risks by explicitly including unsampled populations in the scenarios compared. Finally, we performed additional ABC analyses on real microsatellite data sets to retrace the origin of biocontrol and invasive populations of H. axyridis, taking into account the possibility that the structured native area may have been incompletely sampled. We found that the invasive population in eastern North America, which has served as the bridgehead for worldwide invasion by H. axyridis, was probably formed by an admixture between the eastern and western native clusters. This admixture may have facilitated adaptation of the bridgehead population.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.