Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,066

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,066 results for “bayesian”

Learn how ShareScore rates datasets ↗
zenodo40/100

Supplementary Data to *Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data*

<p>Supplementary code and data to&nbsp;<strong>Robust adaptive distance functions for approximate Bayesian inference on outlier-corrupted data</strong> by <strong>Y. Schaelte et al., 2021</strong>.</p> <p>The archive contains&nbsp;a <strong>README.rst </strong>for information on what is where and how to execute the study and generate the figures. The underlying code without the data can be found at the repository https://github.com/yannikschaelte/study_abc_rad, of which this archive is a snapshot.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Avoiding high frequency thermoacoustic instabilities in cyrogenic rocket engines using Bayesian deep learning

<p>Destructive high-frequency thermoacoustic instabilities have afflicted liquid propellant rocket engine development for decades. The 90 MW cryogenic liquid oxygen/hydrogen multi-injector research combustor BKD operated by DLR Lampoldshausen is a platform that allows their study under realistic conditions. In this study, we use data from BKD experimental campaigns where the static chamber pressure and reactor-oxidizer ratio were varied such that the first tangential mode of the combustor is excited under some conditions. We train a Bayesian neural network to predict the occurence probability of thermoacoustic instabilities 500 ms in the future, given the power spectra of the most recent 300 ms sample of the dynamic pressure data and mass flowrate control signals as input. The Bayesian nature of our algorithms allow us to work in this &quot;small data&quot; setting where the size of our dataset is restricted by the effort and expense associated with each experimental run, without making overconfident extrapolations. We find that the network is able to accurately forecast the occurence probability of instabilities on unseen experimental runs. We envision that these algorithms will eventually be used online by rocket engine controllers to avoid regions of thermoacoustic instabilities.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
dryad40/100

Data and Scripts from: Bayesian prediction of multivariate ecology from phenotypic data yields novel insights into the diets of extant and extinct taxa

<p>Morphology often reflects ecology, enabling the prediction of ecological roles for taxa that lack direct observations such as fossils. In comparative analyses, ecological traits, like diet, are often treated as categorical, which may aid prediction and simplify analyses but ignores the multivariate nature of ecological niches. Futhermore, methods for quantifying and predicting multivariate ecology remain rare. Here, we ranked the relative importance of 13 food items for a sample of 88 extant carnivoran mammals, and then used Bayesian multilevel modeling to assess whether those rankings could be predicted from dental morphology and body size. Traditional diet categories fail to capture the true multivariate nature of carnivoran diets, but Bayesian regression models derived from living taxa have good predictive accuracy for importance ranks. Using our models to predict the importance of individual food items, the multivariate dietary niche, and the nearest extant analogs for a set of data-deficient extant and extinct carnivoran species confirms long-standing ideas for some taxa, but yields new insights about the fundamental dietary niches of others. Our approach provides a promising alternative to traditional dietary classifications. Importantly, this approach need not be limited to diet, but serves as a general framework for predicting multivariate ecology from phenotypic traits.</p>

opencc-zeroNov 2022View details →
dryad40/100

Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models

<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model.  We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate.  We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>

opencc-zeroJun 2022View details →
zenodo40/100

The thermal state of Volgo–Uralia from Bayesian inversion of surface heat flow and temperature [data set]

<p>This collection contains the dataset and the code which were used to find the thermal parameters&rsquo; lateral variations of the Volgo&ndash;Uralian subcraton through the Bayesian Markov Chain Monte Carlo (MCMC) statistical approach. The code originally was given in the analogous study of Antarctica&#39;s geothermal structure by L&ouml;sing et al. (2020) and it can be found in https://github.com/MareenLoesing/GHF-Antarctica-Bayesian. The main changes to the code of L&ouml;sing et al. (2020) are listed in the section 2 of the readme file.</p> <p>For an official use of the Bayesian inversion code please also cite: L&ouml;sing, M., Ebbing, J. &amp; Szwillus, W. (2020) Geothermal Heat Flux in Antarctica: Assessing Models and Observations by Bayesian Inversion. Front. Earth Sci., 8, 105. doi:10.3389/feart.2020.00105</p> <p>The lateral variations of the thermal parameters for the single-layer and multi-layer crust are saved in &ldquo;GHF_Volgo-Uralia_Single-layer.csv&rdquo; and &ldquo;GHF_Volgo-Uralia_Multi-layer.csv&rdquo; respectively.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Spatial confounding in Bayesian species distribution modeling

<ol> <li>Species distribution models (SDMs) are currently the main tools to derive species niche estimates and spatially explicit predictions for species geographical distribution. However, unobserved environmental conditions and ecological processes may confound the model estimates if they have a direct impact on the species and, at the same time, they are correlated with the observed environmental covariates. This, so-called spatial confounding, is a general property of spatial models but it has not been studied in the context of SDMs before.</li> <li>Here we examine how the estimation accuracy of SDMs depends on the type of spatial confounding. We construct two simulation studies where we alter spatial structures of the observed and unobserved covariates and the level of dependence between them. We fit generalized linear models with and without spatial random effects applying Bayesian inference and record the bias induced to model estimates by spatial confounding. After this, we examine spatial confounding also with real vegetation data from northern Norway.</li> <li>Our results show that model estimates for coarse-scale covariates, such as climate covariates, are likely to be biased if a species distribution depends also on an unobserved covariate operating on a finer spatial scale. Pushing higher probability for a relatively weak and spatially smoothly varying spatial random effect compared to the observed covariates improved estimation accuracy. The improvement was independent of the actual spatial structure of the unobserved covariate.</li> <li>Our study addresses the major factors of spatial confounding in SDMs and provides a list of recommendations for pre-inference assessment of spatial confounding and for inference-based methods to decrease the chance of biased model estimates.</li> </ol>

opencc-zeroAug 2022View details →
zenodo40/100

Bayesian Samples and Data Behind Figures: Comprehensive Bayesian Modeling of Tidal Circularization in Open Cluster Binaries part I

<p>Auxiliary data associated with the article <a href="https://ui.adsabs.harvard.edu/abs/2022MNRAS.516.6145P/abstract">&quot;Comprehensive Bayesian Modeling of Tidal Circularization in Open Cluster Binaries part I: M 35, NGC 6819, NGC 188&quot; by Penev, K &amp; Schussler, J</a></p> <p>The type of data corresponds to a particular filename format. Bayesian samples are in HDF5 format, directly as saved by the <a href="https://emcee.readthedocs.io/en/stable/index.html">emcee</a> sampler (see <a href="https://emcee.readthedocs.io/en/stable/user/backends/">https://emcee.readthedocs.io/en/stable/user/backends/</a>). All other files are in AAS-journal style machine readable tables format generated by <a href="https://github.com/cds-astro/cds.pyreadme">cdspyreadme</a> python library.</p> <p>Description of contents by filename format:</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_.*.h5</code></pre> <p>Bayesian analysis samples constraining the tidal dissipation efficiency of the given binary. The values of the sampled system and tidal dissipation parameters are stored as blobs (<a href="https://emcee.readthedocs.io/en/stable/user/blobs/">https://emcee.readthedocs.io/en/stable/user/blobs/)</a></p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_lgQ_period.mrt</code></pre> <p>The 2.3%, 15.9%, 84.1%, and 97.7% quantiles of <span class="math-tex">\(\log_{10}Q_\star'\)</span> for the given binary as a function of tidal period</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_burnin_period.mrt</code></pre> <p>The MCMC burn-in period before the 2.3%, 15.9%, 84.1%, and 97.7% quantiles of <span class="math-tex">\(\log_{10}Q_\star'\)</span> for the given binary are considered converged (see article text).</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_cdfstd_period.mrt</code></pre> <p>The standard deviation of the <span class="math-tex">\(CDF(\log_{10}Q_\star')\)</span> for the given binary as a function of tidal period for each of the quantiles. The maximum likelihood value is the target percentile, i.e. one of: 2.3%, 15.9%, 84.1%, and 97.7%</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains

Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 mitochondrial genomes generated in this study. The dataset was supplemented with Thyropygus sp. and Abacion magnum as outgroups, derived from GenBank. GenBank accession numbers are provided in parentheses. Blue bars indicate the 95% highest probability density intervals for node ages. Age estimation for lineage divergence was based on a general arthropod mitochondrial DNA substitution rate and should be considered with caution. *Thyropygus sp. (red font) is very likely to be a misidentification; for more information, see the Discussion.

opencc-by-4.0Sep 2022View details →
dryad40/100

Data: Applying stochastic and Bayesian integral projection modeling to amphibian population viability analysis

<p>Integral projection models (IPMs) can estimate the population dynamics of species for which both discrete life stages and continuous variables influence demographic rates. Stochastic IPMs for imperiled species, in turn, can facilitate population viability analyses (PVAs) to guide conservation decision-making. Biphasic amphibians are globally distributed, often highly imperiled, and ecologically well-suited to the IPM approach. Herein, we present the first stochastic size- and stage-structured IPM for a biphasic amphibian, the U.S. federally threatened California tiger salamander (<em>Ambystoma</em> <em>californiense</em>; CTS). This Bayesian model reveals that CTS population dynamics show the greatest elasticity to changes in juvenile and metamorph growth and that populations are likely to experience rapid growth at low density. We integrated this IPM with climatic drivers of CTS demography to develop a PVA and examined CTS extinction risk under the primary threats of habitat loss and climate change. The PVA indicates that long-term viability is possible with surprisingly high (20–50%) terrestrial mortality, but simultaneously identified likely minimum terrestrial buffer requirements of 600–1000 m while accounting for numerous parameter uncertainties through the Bayesian framework. These analyses underscore the value of stochastic and Bayesian IPMs for understanding both climate-dependent taxa and those with cryptic life histories (e.g., biphasic amphibians) in service of ecological discovery and biodiversity conservation. In addition to providing guidance for CTS recovery, the contributed IPM and PVA supply a framework for applying these tools to investigations of ecologically-similar species.</p>

opencc-zeroOct 2022View details →
dryad40/100

Cophylogeny reconstruction allowing for multiple associations through approximate Bayesian computation

<p>Phylogenetic tree reconciliation is extensively employed for the examination of coevolution between host and symbiont species. An important concern is the requirement for dependable cost values when selecting event-based parsimonious reconciliation. Although certain approaches deduce event probabilities unique to each pair of host and symbiont trees, which can subsequently be converted into cost values, a significant limitation lies in their inability to model the <em>invasion</em> of diverse host species by the same symbiont species (termed as a spread event), which is believed to occur in symbiotic relationships. Invasions lead to the observation of multiple associations between symbionts and their hosts (indicating that a symbiont is no longer exclusive to a single host), which are incompatible with the existing methods of coevolution. </p> <p>Here, we present a method called AmoCoala (an enhanced version of the tool Coala) that provides a more realistic estimation of cophylogeny event probabilities for a given pair of host and symbiont trees, even in the presence of spread events. We expand the classical 4-event coevolutionary model to include 2 additional spread events (vertical and horizontal spreads) that lead to multiple associations. In the initial step, we estimate the probabilities of spread events using heuristic frequencies. Subsequently, in the second step, we employ an approximate Bayesian computation (ABC) approach to infer the probabilities of the remaining 4 classical events (cospeciation, duplication, host switch, and loss) based on these values.</p> <p>By incorporating spread events, our reconciliation model enables a more accurate consideration of multiple associations. This improvement enhances the precision of estimated cost sets, paving the way to a more reliable reconciliation of host and symbiont trees. To validate our method, we conducted experiments on synthetic datasets and demonstrated its efficacy using real-world examples. Our results showcase that AmoCoala produces biologically plausible reconciliation scenarios, further emphasizing its effectiveness.The software is accessible at <a href="https://github.com/sinaimeri/AmoCoala" rel="noopener">https://github.com/sinaimeri/AmoCoala</a>.</p>

opencc-zeroOct 2022View details →
zenodo40/100

Fig. 1. – Bayesian 50 in Description and phylogenetic position of a new species of Nematanthus (Gesneriaceae) from Bahia, Brazil

Fig. 1. – Bayesian 50% majority rule consensus tree of Nematanthus resulting from the combined analysis of plastid loci atpB-rbcL, matK, rps16, rpl16, trnT-trnL, trnL-trnF, trnS-trnG, and the nuclear regions ncpGS and ITS. Numbers above branches are Bayesian posterior probabilities. Numbers below branches are maximum likelihood bootstrap when ≥50 %. Asterisks indicate species with funnel-shaped and laterally compressed corollas.

opencc-by-4.0Sep 2017View details →
zenodo40/100

AiNU data for Physics-based material parameters extraction from perovskite experiments via Bayesian optimization

<p>This file contains the AiNU data used for the article entitled by <em>Physics-based material parameters extraction from perovskite experiments via Bayesian optimization</em> (https://arxiv.org/abs/2402.11101).</p>

openApr 2024View details →
zenodo40/100

Association of Body Index with Fecal Microbiome in Children Cohorts with Ethnic-Geographic Factor Interaction: Accurately Using a Bayesian Zero-inflated Negative Binomial Regression Model

<p>this dataset are &ldquo;ssociation of Body Index with Fecal Microbiome in Children Cohorts with Ethnic-Geographic Factor Interaction: Accurately Using a Bayesian Zero-inflated Negative Binomial Regression Model&rdquo;&nbsp; Supplementary Material.</p>

opencc-by-4.0May 2024View details →
dryad40/100

Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria

<p>Cyanobacteria are the only prokaryotes to have evolved oxygenic photosynthesis paving the way for complex life. Studying the evolution and ecological niche of cyanobacteria and their ancestors is crucial for understanding the intricate dynamics of biosphere evolution. These organisms frequently deal with environmental stressors such as salinity and drought, and they employ compatible solutes as a mechanism to cope with these challenges. Compatible solutes are small molecules that help maintain cellular osmotic balance in high-salinity environments, such as marine waters. Their production plays a crucial role in salt tolerance, which, in turn, influences habitat preference. Among the five known compatible solutes produced by cyanobacteria (sucrose, trehalose, glucosylglycerol, glucosylglycerate, and glycine betaine), their synthesis varies between individual strains. In this study, we work in a Bayesian stochastic mapping framework, integrating multiple sources of information about compatible solute biosynthesis in order to predict the ancestral habitat preference of Cyanobacteria. Through extensive model selection analyses and statistical tests for correlation, we identify glucosylglycerol and glucosylglycerate as the most significantly correlated with habitat preference, while trehalose exhibits the weakest correlation. Additionally, glucosylglycerol, glucosylglycerate, and glycine betaine show high loss/gain rate ratios, indicating their potential role in adaptability, while sucrose and trehalose are less likely to be lost due to their additional cellular functions. Contrary to previous findings, our analyses predict that the last common ancestor of Cyanobacteria (living at around 3180 Ma) had a 97% probability of a high salinity habitat preference and was likely able to synthesize glucosylglycerol and glucosylglycerate. Nevertheless, cyanobacteria likely colonized low-salinity environments shortly after their origin, with an 89% probability of the first cyanobacterium with low-salinity habitat preference arising prior to the Great Oxygenation Event (2460 Ma). Stochastic mapping analyses provide evidence of cyanobacteria inhabiting early marine habitats, aiding in the interpretation of the geological record. Our age estimate of ~2590 Ma for the divergence of two major cyanobacterial clades (Macro- and Microcyanobacteria) suggests that these were likely significant contributors to primary productivity in marine habitats in the lead-up to the Great Oxygenation Event, and thus played a pivotal role in triggering the sudden increase in atmospheric oxygen.</p>

opencc-zeroMay 2024View details →
zenodo40/100

Figure 1. Bayesian 50 in Taxonomic revision of dragon lizards in the genus Diporiphora (Reptilia: Agamidae) from the Australian monsoonal tropics

Figure 1. Bayesian 50% majority-rules phylogenetic tree for Diporiphora based on mtDNA on ~1200 bp mitochondrial DNA (ND2). Asterisks on branches represent&gt;99% posterior probability support. Clades highlighted in green are expanded, with phylogenetic relationships within each of the species groups reviewed in the current paper: a, D. australis; b, D. bennettii; c, D. bilineata. Species reviewed in the current paper are coloured to represent the taxonomic revision that is undertaken.

opencc-by-4.0Dec 2019View details →
zenodo40/100

APPENDIX 4. — Ultrametric Bayesian tree reconstructed with the 5P in Molecular data reveal the presence of three Plocamium Lamouroux species with complex patterns of distribution in Southern Chile

APPENDIX 4. — Ultrametric Bayesian tree reconstructed with the 5P-COI marker. The dotted vertical red line indicates the maximum likelihood transition point of the switch in branching rates, as estimated by a General Mixed Yule-Coalescent (GMYC) model. The GMYC analysis was performed using a single threshold. Haplotype code as in Appendix 5.

opencc-zeroJan 2021View details →
zenodo40/100

Figure 5. Bayesian Inference tree calculated with complete cox1 in Novel phylogenetic clade of avian Haemoproteus parasites (Haemosporida, Haemoproteidae) from Accipitridae raptors, with description of a new Haemoproteus species

Figure 5. Bayesian Inference tree calculated with complete cox1 (1428 bp), cox3 (753 bp), and cytb (1127 bp) sequences of haemosporidian parasites and Klossiella equi (MH203050) and Klossia razorbacki (MT084562) as the outgroup. Bayesian posterior probabilities and Maximum Likelihood bootstrap values are indicated at most nodes. The scale bar indicates the expected number of substitutions per site according to the model of sequence evolution applied.

opencc-by-4.0Feb 2024View details →
zenodo40/100

Fig 6. A Bayesian time-tree generated from mitochondrial 16S in A New Species of Microhyla (Anurα: Microhylidαe) from Nilphαmαri, Bαnglαdesh

Fig 6. A Bayesian time-tree generated from mitochondrial 16S gene fragment for all known species in the genus Microhyla. Calibration points are indicated with arrows. Numbers are in million years, and the light blue colored bars indicate 95% confidence intervals for divergence time estimates. doi:10.1371/journal.pone.0119825.g006

opencc-by-4.0Mar 2015View details →
zenodo40/100

Fig 1. Bayesian 50 in Alphonsea glandulosa (Annonaceae), a New Species from Yunnan, China

Fig 1. Bayesian 50% majority-rule consensus tree under partitioned models (cpDNA data: matK, ndhF, psbA-trnH, rbcL, trnL-F and ycf1; 63 taxa). Numbers at the nodes indicate Bayesian posterior probabilities and maximum parsimony bootstrap values (&gt; 50%), in that order. Thick lines indicating the clades in which glands or specialized pollinator reward tissues have evolved. doi:10.1371/journal.pone.0170107.g001

opencc-by-4.0Jan 2017View details →
zenodo40/100

FIGURE 3 in Bayesian inference reveals a complex evolutionary history of belemnites

FIGURE 3. Cladogram showing the here suggested systematics of the Belemnitida based on the Bayesian tip-dated analysis. Sketches show the general outer morphological features of a typical representative of the groups in either dorsal (d), ventral (v), or lateral (l) view.

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record