Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,066
datasets available to search
ShareScore release 0.9.0
Dataset results
1,066 results for “bayesian”
Data from: Bayesian total-evidence dating revisits sloth phylogeny and biogeography: a cautionary tale on morphological clock analyses
<p>Combining morphological and molecular characters through Bayesian total-evidence dating allows inferring the phylogenetic and timescale framework of both extant and fossil taxa, while accounting for the stochasticity and incompleteness of the fossil record. Such an integrative approach is particularly needed when dealing with clades such as sloths (Mammalia: Folivora), for which developmental and biomechanical studies have shown high levels of morphological convergence whereas molecular data can only account for a limited percentage of their total species richness. Here, we propose an alternative hypothesis of sloth evolution that emphasizes the pervasiveness of morphological convergence and the importance of considering the fossil record and an adequate taxon sampling in both phylogenetic and biogeographic inferences. Regardless of different clock models and morphological datasets, the extant sloth <em>Bradypus</em> is consistently recovered as a megatherioid, and <em>Choloepus</em> as a mylodontoid, in agreement with molecular-only analyses. The recently extinct Caribbean sloths (Megalocnoidea) are found to be a monophyletic sister-clade of Megatherioidea, in contrast to previous phylogenetic hypotheses. Our results contradict previous morphological analyses and further support the polyphyly of "Megalonychidae", whose members were found in five different clades. Regardless of taxon sampling and clock models, the Caribbean colonization of sloths is compatible with the exhumation of islands along Aves Ridge and its geological time frame. Overall, our total-evidence analysis illustrates the difficulty of positioning highly incomplete fossils, although a robust phylogenetic framework was recovered by an <em>a posteriori</em> removal of taxa with high percentages of missing characters. Elimination of these taxa improved topological resolution by reducing polytomies and increasing node support. However, it introduced a systematic and geographic bias because most of these incomplete specimens are from northern South America. This is evident in biogeographic reconstructions, which suggest Patagonia as the area of origin of many clades when taxa are underrepresented, but Amazonia and/or Central and Southern Andes when all taxa are included. More generally, our analyses demonstrate the instability of topology and divergence time estimates when using different morphological datasets and clock models, and thus caution against making macroevolutionary inferences when node support is weak or when uncertainties in the fossil record are not considered.</p>
Fast Bayesian inference of phylogenies from multiple continuous characters
<p>Time-scaled phylogenetic trees are an ultimate goal of evolutionary biology and a necessary ingredient in comparative studies. The accumulation of genomic data has resolved the tree of life to a great extent, yet timing evolutionary events remains challenging if not impossible without external information such as fossil ages and morphological characters. Methods for incorporating morphology in tree estimation have lagged behind their molecular counterparts, especially in the case of continuous characters. Despite recent advances, such tools are still direly needed as we approach the limits of what molecules can teach us. Here, we implement a suite of state-of-the-art methods for leveraging continuous morphology in phylogenetics, and by conducting extensive simulation studies we thoroughly validate and explore our methods' properties. While retaining model generality and scalability, we make it possible to estimate absolute and relative divergence times from multiple continuous characters while accounting for uncertainty. We compile and analyze one of the most data-type diverse data sets to date, comprised of contemporaneous and ancient molecular sequences, and discrete and continuous characters from living and extinct Carnivora taxa. We conclude by synthesizing lessons about our method's behavior, and suggest future research venues.</p>
Data from: The fundamental role of character coding in Bayesian morphological phylogenetics
<p>Phylogenetic trees establish a historical context for the study of organismal form and function. Most phylogenetic trees are estimated using a model of evolution. For molecular data, modeling evolution is often based on biochemical observations about changes between character states. For example, there are four nucleotides, and we can make assumptions about the probability of transitions between them. By contrast, for morphological characters, we may not know a priori how many character states there are per character, as both extant sampling and the fossil record may be highly incomplete, which leads to an observer bias. For a given character, the state space may be larger than what has been observed in the sample of taxa collected by the researcher. In this case, how many evolutionary rates are needed to even describe transitions between morphological character states may not be clear, potentially leading to model misspecification. To explore the impact of this model misspecification, we simulated character data with varying numbers of character states per character. We then used the data to estimate phylogenetic trees using models of evolution with the correct number of character states and an incorrect number of character states. The results of this study indicate that this observer bias may lead to phylogenetic error, particularly in the branch lengths of trees. If the state space is wrongly assumed to be too large, then we underestimate the branch lengths, and the opposite occurs when the state space is wrongly assumed to be too small.</p>
Investigating the causal association between immune cell phenotypes and allergic diseases and non-allergic asthma using conventional Two-sample and Bayesian weighted Mendelian randomization
<p>Investigating the causal association between immune cell phenotypes and allergic diseases and non-allergic asthma using conventional Two-sample and Bayesian weighted Mendelian randomization</p>
Reduced optical data for Bayesian model trial
<p>The dataset is a subset of a few selected optical indices from the column experiment. </p> <p>The dataset contains all the observations after the column reversal, with day00 indicating the day before the reversal where all the columns are samplead and day0 indicating the day of the reversal. It does not contain any of the control measurements before day00. </p> <p>From the absorbance and fluoresence measurements, we have calculated several indices among which bix (biological index), fi (fluoresence index), hix, (humification index), a254 (decadal absorption coefficient at 254 nm) ,E2_E3 (ratio of absorbance E2 to E3) , SR (slope ratio). All of the indices are calcuated using the StaRdom package from the raw data. </p> <p>Sample name is in the samplingDay_replicate_columnNo format. </p>
Dataset: Testing for effects of growth rate on isotope trophic discrimination factors and evaluating the performance of Bayesian stable isotope mixing models experimentally: a moment of truth?
<p><span>Discerning assimilated diets of wild animals using stable isotopes is well established where potential dietary items in food webs are isotopically distinct. With the advent of mixing models, and Bayesian extensions of such models (Bayesian Stable Isotope Mixing Models, BSIMMs), statistical techniques available for these efforts have been rapidly increasing. The accuracy with which BSIMMs quantify diet, however, depends on several factors including uncertainty in tissue discrimination factors (TDFs; <em>Δ</em>) and identification of appropriate error structures. Whereas performance of BSIMMs has mostly been evaluated with simulations, here we test the efficacy of BSIMMs by raising domestic broiler chicks (<em>Gallus gallus domesticus</em>) on four isotopically distinct diets under controlled environmental conditions, ideal for evaluating factors that affect TDFs and testing how BSIMMs allocate individual birds to diets that vary in isotopic similarity. For both liver and feather tissues,<em> δ</em><sup>13</sup>C and <em>δ </em><sup>15</sup>N values differed among dietary groups. <em>Δ</em><sup>13</sup>C of liver, but not feather, was negatively related to the rate at which individuals gained body mass. For <em>Δ</em><sup>15</sup>N, we identified effects of dietary group, sex, and tissue type, as well as an interaction between sex and tissue type</span><span><span>, </span></span><span><span>with f</span></span><span>emales having higher liver <em>Δ</em><sup>15</sup>N relative to males. For both tissues, BSIMMs allocated most chicks to correct dietary groups, especially for models using combined TDFs rather than diet specific TDFs, and those applying a multiplicative error structure. These findings provide new information on how biological processes affect TDFs and confirm that adequately accounting for variability in consumer isotopes is necessary to optimize performance of BSIMMs. Moreover, they demonstrate experimentally that these types of models reliably characterize consumed diets when appropriately parameterized.<span> </span></span></p>
Supplementary data: Bayesian multi-exposure image fusion for robust high dynamic range ptychography
<p>Accompanying supplementary data for the paper. To download this data automatically and use the software, please refer to the details in the README of the linked github repository. </p> <p><strong>Github URL: </strong><a href="https://github.com/microscopic-image-analysis/bayes-mef"><strong>https://github.com/microscopic-image-analysis/bayes-mef</strong></a></p>
Fluid and kinetic studies of tokamak disruptions using Bayesian optimization
<p>The codes and the data in this directory corresponds to the code and results used in the paper [I. Ekmark et al (2024) J. Plasma Phys., Fluid and kinetic studies of tokamak disruptions using Bayesian optimization, http://arxiv.org/abs/2402.05843]. References to figures below refer to this publication. </p> <p>The optimizations have been performed using the Python package by Fernando Nogueira [https://github.com/bayesian-optimization/BayesianOptimization] and the simulations are performed using the disruptions simulation simulation tool DREAM [https://github.com/chalmersplasmatheory/DREAM, git hash: 0d786e859f6228185ef68b6b3639747e8d96172d], for more information on the latter code visit https://ft.nephy.chalmers.se/dream/.</p> <p>The codes:<br> - BayesianOptimization.py: Runs the optimization, first in fluid and then in isotropic mode. For activated simulations, use the flag "-A".<br> - BlackBox.py: Contains the functions that are run in the optimizations and sets up the simulations.<br> - utils.py: Contains the settings for the simulations as well as some other functions needed in BlackBox.py<br> - ITER.py: Contains all the ITER specific settings.<br> - Exceptions.py: Contains exceptions needed during the simulations. <br> - CostFunction.py: Contains the functions used for evaluating the cost function value for specified values of the representative runaway current, final Ohmic current, current quench time and transported heat fraction.<br> - RunCases.py: Sets up simulations for the cases of table 1 in the paper, as well as for all the optima found.</p> <p>The data:<br> - Optimization results:<br> <br> - Data/OptimizationResults/optresult_fluid.json: Contains the optimization data for the non-activated case using the fluid model. Used to produce figure 1.a. <br> - Data/OptimizationResults/optresult_isotropic.json: Contains the optimization data for the non-activated case using the isotropic model. Used to produce figure 1.b. <br> - Data/OptimizationResults/optresult_fluid_activated.json: Contains the optimization data for the activated case using the fluid model. Used to produce figure 5.a.<br> - Data/OptimizationResults/optresult_isotropic_activated.json: Contains the optimization data for the activated case using the isotropic model. Used to produce figure 5.b.<br> <br> - Data/OptimizationResults/components_fluid.json: Contains the cost function components for each sample from the optimization of the non-activated case using the fluid model. Used to produce figure 2.a. <br> - Data/OptimizationResults/components_isotropic.json: Contains the cost function components for each sample from the optimization of the non-activated case using the isotropic model. Used to produce figure 2.b. <br> - Data/OptimizationResults/components_fluid_activated.json: Contains the cost function components for each sample from the optimization of the activated case using the fluid model. Used to produce figure 6.a.<br> - Data/OptimizationResults/components_isotropic_activated.json: Contains the cost function components for each sample from the optimization of the activated case using the isotropic model. Used to produce figure 6.b.<br> <br> - Cases:<br> - Data/Cases/nonActivatedOpts/fluidOpt/: Contains outputfiles for fluid and isotropic simulations of the optimal case found for the non-activated scenario using the fluid model.<br> - Data/Cases/nonActivatedOpts/isoOpt/: Contains outputfiles for fluid and isotropic simulations of the optimal case found for the non-activated scenario using the isotropic model.<br> - Data/Cases/activatedOpts/fluidOpt/: Contains outputfiles for fluid and isotropic simulations of the optimal case found for the activated scenario using the fluid model.<br> - Data/Cases/activatedOpts/isoOpt/: Contains outputfiles for fluid and isotropic simulations of the optimal case found for the activated scenario using the isotropic model.<br> - Data/Cases/[circle, cross, square, triangle]: Contains outputfiles for fluid and isotropic simulations corresponding to the cases presented in table 1.</p>
Automatic bayesian procedure code and reply figures
<p>Bayesian automatic procedure based on an Dirichlet Multinomial Distribution.</p>
WOMBAT: A fully Bayesian global flux-inversion framework, version 1 (intermediate files)
<p>Code for reproducing the results of the <a href="https://arxiv.org/abs/2102.04004">WOMBAT v1 paper</a> is available on <a href="https://github.com/mbertolacci/wombat-paper">Github</a>. However, fully reproducing the results takes a very long time because of the need to run an atmospheric chemical transport model. To ease this, this dataset contains intermediate files comprising the outputs of the model. These can be used to run the flux inversions; instructions can be found on the code repository.</p>
Demonstrating a Bayesian Online Learning forEnergy-Aware Resource Orchestration in vRANs
<p>Radio Access Network Virtualization (vRAN) will spearhead the quest towards supple radio stacks that adapt to heterogeneous infrastructure: from energy-constrained platforms deploying cells-on-wheels (e.g., drones) or battery-powered cells to green edge clouds. We demonstrate a novel machine learning approach to solve resource orchestration problems in energy-constrained vRANs. Specifically, we demonstrate two algorithms: (i) BP-vRAN, which uses Bayesian online learning to balance performance and energy consumption, and (ii) SBP-vRAN, which augments our Bayesian optimization approach with safe controls that maximize performance while respecting hard power constraints. We show that our approaches are data-efficient, converge an order of magnitude faster than other machine learning methods-and have provably performance, which is paramount for carrier-grade vRANs. We demonstrate the advantages of our approach in a testbed comprised of fully-fledged LTE stacks and a power meter, and implemented our approach into O-RAN's non-real-time RAN Intelligent Controller (RIC).</p>
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 35 Binding Site Data
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 35 binding sites.</p>
A Clustering Approach to Improve IntraVoxel Incoherent Motion Maps from DW-MRI using Conditional Auto-Regressive Bayesian Model
<p>Simulated data generated and used in the paper "A Clustering Approach to Improve IntraVoxel Incoherent Motion Maps from DW-MRI using Conditional Auto-Regressive Bayesian Model" are here available.</p> <p>Results generated from both simulated and clinical datasets are also available on the excel tables.</p>
Bayesian machine learning analysis of single-molecule fluorescence colocalization images
<p>Data files for the "Bayesian machine learning analysis of single-molecule fluorescence colocalization images" manuscript.</p>
Data from: Mental health ecosystem of Gipuzkoa (2015) for Bayesian network modelling
<p>This dataset include data from Mental Health network of Gipuzkoa (Spain). It is included information on resources (inputs) and outcomes (outputs) of care, which are described in the manuscript: "Almeda, N., Garcia-Alonso, C. R., Gutierrez-Colosia, M. R., Salinas-Perez, J. A., Iruin-Sanz, A., & Salvador-Carulla, L. (2022). Modelling the balance of care: Impact of an evidence-informed policy on a mental health ecosystem. PLoS ONE, 17(1 January), 1–16. https://doi.org/10.1371/journal.pone.0261621". This manuscript has been published in Plos One journal.</p> <p>This research focused on developing a formal causal model based on Bayesian network prototypes which were designed by formalizing expert knowledge (by using Expertbased Cooperative Analysis) and resulting in Direct Acyclic Graphs. The best Bayesian networks and their corresponding regression models were used to estimate the statistical ranges or confidence intervals for the dependent variable (potential effect, consequence, or output) given the independent variable values. These ranges, adjusted to delimited statistical distributions (triangular, trapezoidal and gamma), were managed by a Monte Carlo simulation engine for intervention assessment. A computer-based Decision Support System (DSS) was used to assess the status of ecosystem performance: RTE, statistical stability and entropy.</p> <p>Main results of the analyses pointed out that by combining causal reasoning and statistical methods, decision makers can obtain a deep view of both pre-implementing and post-implementing situations. Knowing the causal levers, it is possible to act directly to the causes in order to potentially produce de appropriate results considering the uncertainty: to provide a more balanced and integrated MH care provision in the community. In this particular case, an improvement in the outpatient workforce increases both ecosystem performance (RTE) and stability and slightly decreases entropy.</p>
Rare and widespread: Integrating Bayesian MCMC approaches, Sanger sequencing and Hyb-Seq phylogenomics to reconstruct the origin of the enigmatic Rand Flora genus Camptoloma
<p class="MsoCommentText">Premise</p> <p class="MsoCommentText">Genera that are widespread but have a geographically discontinuous distribution and are represented by few species are intriguing. Did they achieve their disjunct distribution recently, or is it ancient in origin? Why are they species-poor? The Rand Flora is a continental-scale floristic pattern in which closely related species appear co-distributed in isolated regions over the edges of Africa and nearby archipelagos. Genus <i>Camptoloma</i> (Scrophulariaceae) is the most notable example, comprising three species isolated from each other at the ends of the African continent: <i>C. canariense </i>in the west, endemic to the Canary Islands; <i>C. lyperiiflorum </i>in the east, endemic to the Horn of Africa - Southern Arabia; and <i>C. rotundifolia</i>, restricted to Southern Africa.</p> <p class="MsoCommentText">Methods</p> <p class="MsoCommentText">Here, we employed Sanger sequencing of nuclear and plastid markers, together with genomic target sequencing of 2190 low-copy nuclear genes, to infer interspecies relationships and the position of <i>Camptoloma</i> within Scrophulariaceae, using supermatrix and multispecies-coalescent approaches. Lineage divergence times and ancestral ranges were inferred with Bayesian MCMC approaches. Population history was estimated with phylogeographic structured coalescent methods.</p> <p class="MsoCommentText">Key Results</p> <p class="MsoCommentText">Our results support <i>C. rotundifolia</i> as sister to the disjunct clade formed by <i>C. canariense</i> and <i>C. lyperiiflorum.</i> Stem divergence was dated in the Late Miocene, while the origin of extant diversification within the genus was inferred as Early Pliocene.</p> <p class="MsoCommentText">Conclusions</p> <p>We show that the current disjunct distribution of <i>Camptoloma </i>across Africa was likely the result of fragmentation and extinction/population bottlenecking events associated to historical aridification cycles, consistent with the "climatic refugia" hypothesis.</p>
Forecasting suppression of invasive Sea Lamprey in Lake Superior: data and code for Bayesian forecast model
<p>Resource managers frequently are tasked with mitigating or reversing adverse effects of invasive species through management policies and actions. In Lake Superior, of the Laurentian Great Lakes, invasive sea lamprey populations are suppressed to protect valuable fish stocks. However, the relationship between choice of long-term control strategy and the future chance of achieving the suppression target is unclear.</p> <p>Using a 60+ year time-series of suppression effort and monitoring data from 50 assessment sites located on Lake Superior tributaries, we developed a Bayesian state-space model to forecast the probability of suppressing lamprey below the suppression target.</p> <p>With annual application of lampricide (i.e., lamprey-specific pesticide) at historical mean levels, we forecasted a 15% chance of achieving the Lake Superior sea lamprey suppression target in 2040.</p> <p>Increasing lampricide effort and/or supplementing lampricide control with age-1 recruitment reduction increased suppression chance. Annual application of the maximum historical lampricide effort resulted in a 50% predicted chance of achieving the target, annual application of the mean historic lampricide effort plus a 40% reduction in recruitment resulted in a 54% chance, and the maximum amount of effort considered (maximum historic lampricide and 60% reduction in recruitment) resulted in a 94% chance.</p> <p><em><a>Policy </a>implications</em>. <a>We</a> developed a simulation model from a robust, long-term monitoring dataset that improves understanding of why long-term sea lamprey suppression objectives have been difficult to achieve in Lake Superior. Furthermore, the model provides a means to gauge efficacy of sea lamprey control policy and action scenarios based on forecasted chance of achieving the suppression target. Creating processes for iteratively refining our forecasting model with stakeholder and technical-expert input and integration with a decision analysis framework could strengthen the link between ecological knowledge obtained from long-term monitoring and invasive sea lamprey management.</p>
Dodonaphy - a Software using Hyperbolic Space for Bayesian Phylogenetic Inference
<p>Bayesian inference for phylogenetics is a gold standard for computing distributions of phylogenies. It faces the challenging problem of moving throughout the high-dimensional space of trees. However, hyperbolic space offers a low dimensional representation of tree-like data. In this paper, we embed genomic sequences into hyperbolic space and perform hyperbolic Markov Chain Monte Carlo for Bayesian inference. The posterior probability is computed by decoding a neighbour joining tree from proposed embedding locations. We empirically demonstrate the fidelity of this method on eight data sets. The sampled posterior distribution recovers the splits and branch lengths to a high degree. We investigated the effects of curvature and embedding dimension on the Markov Chain's performance. Finally, we discuss the prospects for adapting this method to navigate tree space with gradients.</p> <p>This software embeds phylogenetic taxa in hyerbolic space to perform Bayesian inference. Version 1.0.0 includes a Markov Chain Monte Carlo (MCMC) that we compared to the state-of-art on eight datasets (previously published elsewhere). This package is implemented in Python3.9 with a simle command line interface provided.</p>
Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models
<p>This is the consensus sequence repository to the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>
In Search of the Edge: A Bayesian Exploration of the Detectability of Red Edges in Exoplanet Reflection Spectra
<p>This repository contains surface albedos for a paper submitted to AAS journals under the same title. Here we include a a realistic Earth-like surface albedo and the raw albedo files used for it's calculation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.