Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
139
datasets available to search
ShareScore release 0.7.1
Dataset results
139 results for “bayesian analysis”
Data from: Bayesian analysis of biogeography when the number of areas is large
Open the record for dataset details and reuse information.
Data from: Genetic heterogeneity underlying variation in a locally adaptive clinal trait in Pinus sylvestris revealed by a Bayesian multipopulation analysis
Open the record for dataset details and reuse information.
Data from: Reconstruction of a beech population bottleneck using archival demographic information and Bayesian analysis of genetic data
Open the record for dataset details and reuse information.
Data from: Phylogeny, macroevolutionary trends and historical biogeography of sloths: insights from a Bayesian morphological clock analysis
Open the record for dataset details and reuse information.
Data from: Comparing traditional and Bayesian approaches to ecological meta-analysis
Open the record for dataset details and reuse information.
[dataset] Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference
<p>Anonymized replication package of the experiments presented in the research paper: "Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference".</p> <p>See the README.md file.</p>
Fig. 2. Bayesian majority rule consensus tree reconstructed for 90 in Phylogenetic analysis and systematic position of two new species of the ant genus Crematogaster (Hymenoptera, Formicidae) from Southeast Asia
Fig. 2. Bayesian majority rule consensus tree reconstructed for 90 taxa using five genes (ArgK, CAD, LWRh, Top1, Wg) in a MrBayes analysis. Most of the outgroups are not shown. Above node numbers indicate posterior probability, bootstrap value for MP, and bootstrap value for ML. Data were partitioned by PartitionFinder v.1.1.1 and analyzed using a best fit model for each gene and codon position, with 10 million generations and a burn-in of 25 %.
Data from: Homoplasy-based partitioning outperforms alternatives in Bayesian analysis of discrete morphological data
Bayesian analysis of morphological data is becoming increasingly popular mainly (but not only) because it allows for time-calibrated phylogenetic inference using relaxed morphological clocks and tip dating whenever fossils are available. As with molecular data, recent studies have shown that modeling among character rate variaton (ACRV) in morphological matrices greatly improves phylogenetic inference. In a likelihood framework this may be accomplished, for instance, by employing a hidden Markov model (HMM) to assign characters to rate categories drawn from a (discretized) Γ distribution and/or by partitioning datasets according to rate heterogeneity and estimating per-partition branch lengths, conditioned on a single topology. While the first approach is available in many phylogenetic analysis software, there is still no clear consensus on how to partition data, except perhaps in the simplest cases (e.g. "by codon" partitioning of coding sequences). Additionally, there is a trade-off between improvement in likelihood scores and the number of free parameters in the analysis, which rises quickly with the number of partitions. This trade-off may be dealt with by employing statistics that penalize overfitting of complex models, such as Akaike or Bayesian information criteria (AIC and BIC), or the more recently introduced stepping-stone (SS) method for marginal likelihood approximation. We applied the latter to three distinct matrices of discrete morphological data and demonstrated that sorting characters by homoplasy scores (obtained from implied weighting parsimony analysis) outperformed other partitioning strategies (anatomically-based and PartitionFinder2). The method was in fact so efficient in segregating characters by rates of evolution that no within-partition ACRV modeling was necessary, while among partition rate variation (APRV) was adequately accommodated by rate multipliers. We conclude that partitioning by homoplasy is a powerful and easy-to-implement strategy to address ACRV in complex datasets. We provide some guidelines focusing on morphological matrices, although this approach may be also applicable to molecular datasets.
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.
Data from: Bayesian phylogenetic analysis of combined data
The recent development of Bayesian phylogenetic inference using Markov chain Monte Carlo (MCMC) techniques has facilitated the exploration of parameter-rich evolutionary models. At the same time, stochastic models have become more realistic (and complex) and have been extended to new types of data, such as morphology. Based on this foundation, we developed a Bayesian MCMC approach to the analysis of combined data sets and explored its utility in inferring relationships among gall wasps based on data from morphology and four genes (nuclear and mitochondrial, ribosomal and protein coding). Examined models range in complexity from those recognizing only a morphological and a molecular partition to those having complex substitution models with independent parameters for each gene. Bayesian MCMC analysis deals efficiently with complex models: convergence occurs faster and more predictably for complex models, mixing is adequate for all parameters even under very complex models, and the parameter update cycle is virtually unaffected by model partitioning across sites. Morphology contributed only 5% of the characters in the data set but nevertheless influenced the combined-data tree, supporting the utility of morphological data in multigene analyses. We used Bayesian criteria (Bayes factors) to show that process heterogeneity across data partitions is a significant model component, although not as important as among-site rate variation. More complex evolutionary models are associated with more topological uncertainty and less conflict between morphology and molecules. Bayes factors sometimes favor simpler models over considerably more parameter-rich models, but the best model overall is also the most complex and Bayes factors do not support exclusion of apparently weak parameters from this model. Thus, Bayes factors appear to be useful for selecting among complex models, but it is still unclear whether their use strikes a reasonable balance between model complexity and error in parameter estimates.
FIGURE 2 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences
FIGURE 2. Phylogenetic tree obtained from the Bayesian inference analysis of the 16S rDNA dataset.
FIGURE 1 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences
FIGURE 1. Phylogenetic tree obtained from the Bayesian inference analysis of the 12S rDNA dataset.
BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point Detection
<p>Artifacts for the paper titled <strong><em>BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point Detection</em></strong>.</p> <p>This artifact repository contains 3 compressed folders, as follows: </p> <table> <tbody> <tr> <td><strong>File Name</strong></td> <td><strong>Benchmark System</strong></td> </tr> <tr> <td>fse-ob.zip</td> <td>Online Boutique</td> </tr> <tr> <td>fse-ss.zip</td> <td>Sock Shop</td> </tr> <tr> <td>fse-tt.zip</td> <td>Train Ticket</td> </tr> </tbody> </table> <p>Each zip file contains the collected data from the corresponding microservice benchmark systems (e.g., fse-ob.zip contains metrics data collected from the Online Boutique system). </p> <p><strong><strong>Data description</strong></strong></p> <p>To collect the metrics data, we deploy three benchmark microservice systems: Online Boutique, Sock Shop, and Train Ticket, on a Kubernetes cluster consisting of one master node and five worker nodes. Then, we deploy a monitoring system to monitor and collect resource-level and service-level metrics. To generate traffic, we use the load generators supplied by these systems and tailor them to explore all services with a load of 40-50 requests per second. Initially, we operate the applications normally to gather metrics data under normal conditions. Then, we inject faults into the running services. We execute into the designated container using kubectl exec. For CPU hog and memory leak, we use stress-ng to stress the container resource. For network delay and packet loss, we use tc (traffic control) to manipulate the traffic of the container. Specifically, we inject faults into five targeted services of Sock Shop (carts, catalogue, orders, payment, and user), five targeted services of Online Boutique (adservice, cartservice, checkoutservice, currencyservice, and productcatalogue), and five targeted services of Train Ticket (ts-auth-service, ts-order-service, ts-route-service, ts-train-service, ts-travel-service). For each combination of fault type and targeted service, we repeat the operation (i.e., fault injection and metrics data collection) five times, resulting in 100 failure cases for each benchmark microservice system.</p> <p><strong>Code</strong></p> <p>The code to reproduce the experimental results in the paper is available at <a href="https://github.com/phamquiluan/baro">https://github.com/phamquiluan/baro</a>.</p>
Using Bayesian Spectrum Analysis to determine probabilities for LH and hot flush intervals and the probability of a match between them with simulated data
<p>To illustrate the principle of the analysis used using simulated data</p>
Data from: Conventional analysis of trial-by-trial adaptation is biased: empirical and theoretical support using a Bayesian estimator
Research on human motor adaptation has often focused on how people adapt to self-generated or externally-influenced errors. Trial-by-trial adaptation is a person's response to self-generated errors. Externally-influenced errors applied as catch-trial perturbations are used to calculate a person's perturbation adaptation rate. Although these adaptation rates are sometimes compared to one another, we show through simulation and empirical data that the two metrics are distinct. We demonstrate that the trial-by-trial adaptation rate, often calculated as a coefficient in a linear regression, is biased under typical conditions. We tested 12 able-bodied subjects moving a cursor on a screen using a computer mouse. Statistically different adaptation rates arise when sub-sets of trials from different phases of learning are analyzed from within a sequence of movement results. We propose a new approach to identify when a person's learning has stabilized in order to identify steady-state movement trials from which to calculate a more reliable trial-by-trial adaptation rate. Using a Bayesian model of human movement, we show that this analysis approach is more consistent and provides a more confident estimate than alternative approaches. Constraining analyses to steady-state conditions will allow researchers to better decouple the multiple concurrent learning processes that occur while a person makes goal-directed movements. Streamlining this analysis may help broaden the impact of motor adaptation studies, perhaps even enhancing their clinical usefulness.
The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets
<p class="CxSpFirst">The software program STRUCTURE is one of the most cited tools for determining population structure. To infer the optimal number of clusters from STRUCTURE output, the Δ<i>K</i> method is often applied. However, a recent study relying on simulated microsatellite data suggested that this method has a downward bias in its estimation of <i>K</i> and is sensitive to uneven sampling. If this finding holds for empirical datasets, conclusions about the scale of gene flow may have to be revised for a large number of studies. To determine the impact of method choice, we applied recently described estimators of <i>K</i> to re-estimate genetic structure in 41 empirical microsatellite datasets; 15 from a broad range of taxa and 26 focused on a diverse phylogenetic group, coral. We compared alternative estimates of <i>K</i> (Puechmaille statistics) with traditional (Δ<i>K</i> and posterior probability) estimates and found widespread disagreement of estimators across datasets. Thus, one estimator alone is insufficient for determining the optimal number of clusters regardless of study organism or evenness of sampling scheme. Subsequent analysis of molecular variance (AMOVA) between clustering solutions did not necessarily clarify which solution was best. To better infer population structure, we suggest a combination of visual inspection of STRUCTURE plots and calculation of the alternative estimators at various thresholds in addition to Δ<i>K</i>. Differences between estimators could reveal patterns with important biological implications, such as the potential for more population structure than previously estimated, as was the case for many studies reanalyzed here.</p>
BRACE: A Bayesian-based dimension reduction approach for single-cell alternative splicing analysis
<p>Datasets to demonstrate the utility of our BRACE method.</p>
Data from: Homoplasy-based partitioning outperforms alternatives in Bayesian analysis of discrete morphological data
Open the record for dataset details and reuse information.
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
Open the record for dataset details and reuse information.
Data from: Bayesian phylogenetic analysis of combined data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.