Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “Bayesian model selection”
Code and data for Bayesian joint species distribution model selection for community-level prediction
<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction." Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript. Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>
Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models
<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model. We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate. We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>
Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria
<p>Cyanobacteria are the only prokaryotes to have evolved oxygenic photosynthesis paving the way for complex life. Studying the evolution and ecological niche of cyanobacteria and their ancestors is crucial for understanding the intricate dynamics of biosphere evolution. These organisms frequently deal with environmental stressors such as salinity and drought, and they employ compatible solutes as a mechanism to cope with these challenges. Compatible solutes are small molecules that help maintain cellular osmotic balance in high-salinity environments, such as marine waters. Their production plays a crucial role in salt tolerance, which, in turn, influences habitat preference. Among the five known compatible solutes produced by cyanobacteria (sucrose, trehalose, glucosylglycerol, glucosylglycerate, and glycine betaine), their synthesis varies between individual strains. In this study, we work in a Bayesian stochastic mapping framework, integrating multiple sources of information about compatible solute biosynthesis in order to predict the ancestral habitat preference of Cyanobacteria. Through extensive model selection analyses and statistical tests for correlation, we identify glucosylglycerol and glucosylglycerate as the most significantly correlated with habitat preference, while trehalose exhibits the weakest correlation. Additionally, glucosylglycerol, glucosylglycerate, and glycine betaine show high loss/gain rate ratios, indicating their potential role in adaptability, while sucrose and trehalose are less likely to be lost due to their additional cellular functions. Contrary to previous findings, our analyses predict that the last common ancestor of Cyanobacteria (living at around 3180 Ma) had a 97% probability of a high salinity habitat preference and was likely able to synthesize glucosylglycerol and glucosylglycerate. Nevertheless, cyanobacteria likely colonized low-salinity environments shortly after their origin, with an 89% probability of the first cyanobacterium with low-salinity habitat preference arising prior to the Great Oxygenation Event (2460 Ma). Stochastic mapping analyses provide evidence of cyanobacteria inhabiting early marine habitats, aiding in the interpretation of the geological record. Our age estimate of ~2590 Ma for the divergence of two major cyanobacterial clades (Macro- and Microcyanobacteria) suggests that these were likely significant contributors to primary productivity in marine habitats in the lead-up to the Great Oxygenation Event, and thus played a pivotal role in triggering the sudden increase in atmospheric oxygen.</p>
Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models
Open the record for dataset details and reuse information.
Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria
Open the record for dataset details and reuse information.
Code and data for Bayesian joint species distribution model selection for community-level prediction
Open the record for dataset details and reuse information.
Pre-fitted Bayesian models for "Gene panel selection for targeted spatial transcriptomics"
<p>simulation_parameters_DARTFISH_slim.rds: Bayesian model fitted on the Zhang dataset.</p> <p>simulation_parameters_MERFISH_slim.rds: Bayesian model fitted on the Moffit dataset.</p> <p>simulation_parameters_osmFISH_slim.rds: Bayesian model fitted on the Codeluppi dataset.</p> <p> </p>
Input data for Bayesian and information theoretic model selection and similarity analysis
<p>This data serves as input to the codes found in the following repository https://github.com/MariaFMoralesOreamuno/Bayesian_Information_theoretic_model_selection.git</p> <p> </p>
cellsig: a Bayesian sparse random-effect model and large-scale human bulk transcriptional catalogue support cell-type marker selection
<p><strong>Human Bulk Cell-type Catalogue (HBCC): This database harmonises 1,435 samples of 67 cell types from 58 datasets which can be utilised in the cellsig method for estimating cell-type transcriptional profiles.</strong></p>
Selecting and averaging relaxed clock models in Bayesian tip dating of Mesozoic birds
Open the record for dataset details and reuse information.
Data from: Disentangling the formation of contrasting tree-line physiognomies combining model selection and Bayesian parameterization for simulation models
Open the record for dataset details and reuse information.
Data from: Posterior predictive Bayesian phylogenetic model selection
We present two distinctly different posterior predictive approaches to Bayesian phylogenetic model selection, and illustrate these methods using examples from green algal protein-coding cpDNA sequences and flowering plant rDNA sequences. The Gelfand-Ghosh (GG) approach allows dissection of an overall measure of model fit into components due to posterior predictive variance (Pm) and goodness-of-fit (Gm), which distinguishes this method from the posterior predictive P-value approach. The conditional predictive ordinate (CPO) method provides a site-specific measure of model fit useful for exploratory analyses and can be combined over sites yielding the log pseudomarginal likelihood (LPML), which is useful as an overall measure of model fit. CPO provides a useful cross-validation approach that is computationally efficient, requiring only a sample from the posterior distribution (no additional simulation is required). Both GG and CPO add new perspectives to Bayesian phylogenetic model selection based on the predictive abilities of models, and complement the perspective provided by the marginal likelihood (including Bayes Factor comparisons) based solely on the fit of competing models to observed data.
Data from: Bayesian model selection with BAMM: effects of the model prior on the inferred number of diversification shifts
1. Understanding variation in rates of speciation and extinction -- both among lineages and through time -- is critical to the testing of many hypotheses about macroevolutionary processes. BAMM is a flexible Bayesian framework for inferring the number and location of shifts in macroevolutionary rate across phylogenetic trees and has been widely used in empirical studies. BAMM requires that researchers specify a prior probability distribution on the number of diversification rate shifts before conducting an analysis. The consequences of this "model prior" for inference are poorly known but could potentially influence both the probability of accepting models that are more (high error rate) or less (low power) complex than the generating model. 2. The hierarchical Poisson process prior in BAMM reduces to a simple geometric distribution on number of rate shifts and we use this property to increase the efficiency of model selection with Bayes factors. Using BAMM v2.5, we analyzed phylogenies simulated with and without diversification heterogeneity across a broad range of prior parameterizations. We also assessed the impact of the model prior on MCMC convergence times and on diversification rate estimates. 3. For all simulation scenarios, model evidence (Bayes factor support) for the number of shifts is not sensitive to the choice of model prior over the wide range examined here. The best-supported model found using BAMM rarely includes spurious shifts (<2% of all runs) when diversification models are selected using Bayes factors. BAMM was reliably able to infer the true number of diversification rate shifts across prior expectations that varied by three orders of magnitude. However, we find a strong effect of model prior on MCMC convergence properties: a flatter prior distribution (larger expected number of shifts) can dramatically increase the efficiency of the MCMC simulation. 4. Our results support the use of a liberal model prior in BAMM, as it reduces computation time without distorting the evidence for rate heterogeneity.
Data from: Bayesian model selection with BAMM: effects of the model prior on the inferred number of diversification shifts
Open the record for dataset details and reuse information.
Data from: Posterior predictive Bayesian phylogenetic model selection
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.