Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

139

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

139 results for “bayesian analysis”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Bayesian analysis of biogeography when the number of areas is large

Open the record for dataset details and reuse information.

publicMay 2013View details →
dryad32/100

Data from: Genetic heterogeneity underlying variation in a locally adaptive clinal trait in Pinus sylvestris revealed by a Bayesian multipopulation analysis

Open the record for dataset details and reuse information.

publicOct 2016View details →
dryad32/100

Data from: Reconstruction of a beech population bottleneck using archival demographic information and Bayesian analysis of genetic data

Open the record for dataset details and reuse information.

publicOct 2011View details →
dryad32/100

Data from: Phylogeny, macroevolutionary trends and historical biogeography of sloths: insights from a Bayesian morphological clock analysis

Open the record for dataset details and reuse information.

publicSep 2018View details →
dryad32/100

Data from: Comparing traditional and Bayesian approaches to ecological meta-analysis

Open the record for dataset details and reuse information.

publicJul 2020View details →
zenodo28/100

[dataset] Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference

<p>Anonymized replication package of the experiments presented in the research paper: &quot;Performance Analysis of Microservice Applications via Automated Load Testing and Bayesian Inference&quot;.</p> <p>See the README.md file.</p>

openMay 2020View details →
zenodo28/100

Fig. 2. Bayesian majority rule consensus tree reconstructed for 90 in Phylogenetic analysis and systematic position of two new species of the ant genus Crematogaster (Hymenoptera, Formicidae) from Southeast Asia

Fig. 2. Bayesian majority rule consensus tree reconstructed for 90 taxa using five genes (ArgK, CAD, LWRh, Top1, Wg) in a MrBayes analysis. Most of the outgroups are not shown. Above node numbers indicate posterior probability, bootstrap value for MP, and bootstrap value for ML. Data were partitioned by PartitionFinder v.1.1.1 and analyzed using a best fit model for each gene and codon position, with 10 million generations and a burn-in of 25 %.

opencc-by-3.0Nov 2017View details →
dryad28/100

Data from: Homoplasy-based partitioning outperforms alternatives in Bayesian analysis of discrete morphological data

Bayesian analysis of morphological data is becoming increasingly popular mainly (but not only) because it allows for time-calibrated phylogenetic inference using relaxed morphological clocks and tip dating whenever fossils are available. As with molecular data, recent studies have shown that modeling among character rate variaton (ACRV) in morphological matrices greatly improves phylogenetic inference. In a likelihood framework this may be accomplished, for instance, by employing a hidden Markov model (HMM) to assign characters to rate categories drawn from a (discretized) Γ distribution and/or by partitioning datasets according to rate heterogeneity and estimating per-partition branch lengths, conditioned on a single topology. While the first approach is available in many phylogenetic analysis software, there is still no clear consensus on how to partition data, except perhaps in the simplest cases (e.g. "by codon" partitioning of coding sequences). Additionally, there is a trade-off between improvement in likelihood scores and the number of free parameters in the analysis, which rises quickly with the number of partitions. This trade-off may be dealt with by employing statistics that penalize overfitting of complex models, such as Akaike or Bayesian information criteria (AIC and BIC), or the more recently introduced stepping-stone (SS) method for marginal likelihood approximation. We applied the latter to three distinct matrices of discrete morphological data and demonstrated that sorting characters by homoplasy scores (obtained from implied weighting parsimony analysis) outperformed other partitioning strategies (anatomically-based and PartitionFinder2). The method was in fact so efficient in segregating characters by rates of evolution that no within-partition ACRV modeling was necessary, while among partition rate variation (APRV) was adequately accommodated by rate multipliers. We conclude that partitioning by homoplasy is a powerful and easy-to-implement strategy to address ACRV in complex datasets. We provide some guidelines focusing on morphological matrices, although this approach may be also applicable to molecular datasets.

opencc-zeroDec 2018View details →
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Bayesian phylogenetic analysis of combined data

The recent development of Bayesian phylogenetic inference using Markov chain Monte Carlo (MCMC) techniques has facilitated the exploration of parameter-rich evolutionary models. At the same time, stochastic models have become more realistic (and complex) and have been extended to new types of data, such as morphology. Based on this foundation, we developed a Bayesian MCMC approach to the analysis of combined data sets and explored its utility in inferring relationships among gall wasps based on data from morphology and four genes (nuclear and mitochondrial, ribosomal and protein coding). Examined models range in complexity from those recognizing only a morphological and a molecular partition to those having complex substitution models with independent parameters for each gene. Bayesian MCMC analysis deals efficiently with complex models: convergence occurs faster and more predictably for complex models, mixing is adequate for all parameters even under very complex models, and the parameter update cycle is virtually unaffected by model partitioning across sites. Morphology contributed only 5% of the characters in the data set but nevertheless influenced the combined-data tree, supporting the utility of morphological data in multigene analyses. We used Bayesian criteria (Bayes factors) to show that process heterogeneity across data partitions is a significant model component, although not as important as among-site rate variation. More complex evolutionary models are associated with more topological uncertainty and less conflict between morphology and molecules. Bayes factors sometimes favor simpler models over considerably more parameter-rich models, but the best model overall is also the most complex and Bayes factors do not support exclusion of apparently weak parameters from this model. Thus, Bayes factors appear to be useful for selecting among complex models, but it is still unclear whether their use strikes a reasonable balance between model complexity and error in parameter estimates.

opencc-zeroDec 2017View details →
zenodo28/100

FIGURE 2 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences

FIGURE 2. Phylogenetic tree obtained from the Bayesian inference analysis of the 16S rDNA dataset.

opennotspecifiedAug 2012View details →
zenodo28/100

FIGURE 1 in Systematic position of Dinidoridae within the superfamily Pentatomoidea (Hemiptera: Heteroptera) revealed by the Bayesian phylogenetic analysis of the mitochondrial 12S and 16S rDNA sequences

FIGURE 1. Phylogenetic tree obtained from the Bayesian inference analysis of the 12S rDNA dataset.

opennotspecifiedAug 2012View details →
zenodo28/100

BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point Detection

<p>Artifacts for the paper titled <strong><em>BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point Detection</em></strong>.</p> <p>This artifact repository contains 3 compressed folders, as follows:&nbsp;</p> <table> <tbody> <tr> <td><strong>File Name</strong></td> <td><strong>Benchmark System</strong></td> </tr> <tr> <td>fse-ob.zip</td> <td>Online Boutique</td> </tr> <tr> <td>fse-ss.zip</td> <td>Sock Shop</td> </tr> <tr> <td>fse-tt.zip</td> <td>Train Ticket</td> </tr> </tbody> </table> <p>Each zip file contains the collected data from the corresponding microservice benchmark systems (e.g., fse-ob.zip contains metrics data collected from the Online Boutique system).&nbsp;</p> <p><strong><strong>Data description</strong></strong></p> <p>To collect the metrics data, we deploy three benchmark microservice systems: Online Boutique, Sock Shop, and Train Ticket, on a Kubernetes cluster consisting of one master node and five worker nodes. Then, we deploy a monitoring system to monitor and collect resource-level and service-level metrics. To generate traffic, we use the load generators supplied by these systems and tailor them to explore all services with a load of 40-50 requests per second. Initially, we operate the applications normally to gather metrics data under normal conditions. Then, we inject faults into the running services. We execute into the designated container using kubectl exec. For CPU hog and memory leak, we use stress-ng to stress the container resource. For network delay and packet loss, we use tc (traffic control) to manipulate the traffic of the container. Specifically, we inject faults into five targeted services of Sock Shop (carts, catalogue, orders, payment, and user), five targeted services of Online Boutique (adservice, cartservice, checkoutservice, currencyservice, and productcatalogue), and five targeted services of Train Ticket (ts-auth-service, ts-order-service, ts-route-service, ts-train-service, ts-travel-service). For each combination of fault type and targeted service, we repeat the operation (i.e., fault injection and metrics data collection) five times, resulting in 100 failure cases for each benchmark microservice system.</p> <p><strong>Code</strong></p> <p>The code to reproduce the experimental results in the paper is available at <a href="https://github.com/phamquiluan/baro">https://github.com/phamquiluan/baro</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo28/100

Using Bayesian Spectrum Analysis to determine probabilities for LH and hot flush intervals and the probability of a match between them with simulated data

<p>To illustrate the principle of the analysis used using&nbsp;simulated data</p>

opencc-by-4.0Mar 2019View details →
dryad28/100

Data from: Conventional analysis of trial-by-trial adaptation is biased: empirical and theoretical support using a Bayesian estimator

Research on human motor adaptation has often focused on how people adapt to self-generated or externally-influenced errors. Trial-by-trial adaptation is a person's response to self-generated errors. Externally-influenced errors applied as catch-trial perturbations are used to calculate a person's perturbation adaptation rate. Although these adaptation rates are sometimes compared to one another, we show through simulation and empirical data that the two metrics are distinct. We demonstrate that the trial-by-trial adaptation rate, often calculated as a coefficient in a linear regression, is biased under typical conditions. We tested 12 able-bodied subjects moving a cursor on a screen using a computer mouse. Statistically different adaptation rates arise when sub-sets of trials from different phases of learning are analyzed from within a sequence of movement results. We propose a new approach to identify when a person's learning has stabilized in order to identify steady-state movement trials from which to calculate a more reliable trial-by-trial adaptation rate. Using a Bayesian model of human movement, we show that this analysis approach is more consistent and provides a more confident estimate than alternative approaches. Constraining analyses to steady-state conditions will allow researchers to better decouple the multiple concurrent learning processes that occur while a person makes goal-directed movements. Streamlining this analysis may help broaden the impact of motor adaptation studies, perhaps even enhancing their clinical usefulness.

opencc-zeroDec 2017View details →
dryad28/100

The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets

<p class="CxSpFirst">The software program STRUCTURE is one of the most cited tools for determining population structure. To infer the optimal number of clusters from STRUCTURE output, the Δ<i>K</i> method is often applied. However, a recent study relying on simulated microsatellite data suggested that this method has a downward bias in its estimation of <i>K</i> and is sensitive to uneven sampling. If this finding holds for empirical datasets, conclusions about the scale of gene flow may have to be revised for a large number of studies. To determine the impact of method choice, we applied recently described estimators of <i>K</i> to re-estimate genetic structure in 41 empirical microsatellite datasets; 15 from a broad range of taxa and 26 focused on a diverse phylogenetic group, coral. We compared alternative estimates of <i>K</i> (Puechmaille statistics) with traditional (Δ<i>K</i> and posterior probability) estimates and found widespread disagreement of estimators across datasets. Thus, one estimator alone is insufficient for determining the optimal number of clusters regardless of study organism or evenness of sampling scheme. Subsequent analysis of molecular variance (AMOVA) between clustering solutions did not necessarily clarify which solution was best. To better infer population structure, we suggest a combination of visual inspection of STRUCTURE plots and calculation of the alternative estimators at various thresholds in addition to Δ<i>K</i>. Differences between estimators could reveal patterns with important biological implications, such as the potential for more population structure than previously estimated, as was the case for many studies reanalyzed here.</p>

opencc-zeroOct 2021View details →
zenodo28/100

BRACE: A Bayesian-based dimension reduction approach for single-cell alternative splicing analysis

<p>Datasets to demonstrate the utility of our BRACE method.</p>

opencc-by-4.0Jun 2023View details →
dryad28/100

Data from: Homoplasy-based partitioning outperforms alternatives in Bayesian analysis of discrete morphological data

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

Open the record for dataset details and reuse information.

publicJun 2012View details →
dryad28/100

Data from: Bayesian phylogenetic analysis of combined data

Open the record for dataset details and reuse information.

publicJul 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record