Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “Bayesian model selection”

Learn how ShareScore rates datasets ↗
dryad40/100

Code and data for Bayesian joint species distribution model selection for community-level prediction

<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction."  Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript.  Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>

opencc-zeroNov 2023View details →
dryad40/100

Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models

<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model.  We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate.  We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>

opencc-zeroJun 2022View details →
dryad40/100

Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria

<p>Cyanobacteria are the only prokaryotes to have evolved oxygenic photosynthesis paving the way for complex life. Studying the evolution and ecological niche of cyanobacteria and their ancestors is crucial for understanding the intricate dynamics of biosphere evolution. These organisms frequently deal with environmental stressors such as salinity and drought, and they employ compatible solutes as a mechanism to cope with these challenges. Compatible solutes are small molecules that help maintain cellular osmotic balance in high-salinity environments, such as marine waters. Their production plays a crucial role in salt tolerance, which, in turn, influences habitat preference. Among the five known compatible solutes produced by cyanobacteria (sucrose, trehalose, glucosylglycerol, glucosylglycerate, and glycine betaine), their synthesis varies between individual strains. In this study, we work in a Bayesian stochastic mapping framework, integrating multiple sources of information about compatible solute biosynthesis in order to predict the ancestral habitat preference of Cyanobacteria. Through extensive model selection analyses and statistical tests for correlation, we identify glucosylglycerol and glucosylglycerate as the most significantly correlated with habitat preference, while trehalose exhibits the weakest correlation. Additionally, glucosylglycerol, glucosylglycerate, and glycine betaine show high loss/gain rate ratios, indicating their potential role in adaptability, while sucrose and trehalose are less likely to be lost due to their additional cellular functions. Contrary to previous findings, our analyses predict that the last common ancestor of Cyanobacteria (living at around 3180 Ma) had a 97% probability of a high salinity habitat preference and was likely able to synthesize glucosylglycerol and glucosylglycerate. Nevertheless, cyanobacteria likely colonized low-salinity environments shortly after their origin, with an 89% probability of the first cyanobacterium with low-salinity habitat preference arising prior to the Great Oxygenation Event (2460 Ma). Stochastic mapping analyses provide evidence of cyanobacteria inhabiting early marine habitats, aiding in the interpretation of the geological record. Our age estimate of ~2590 Ma for the divergence of two major cyanobacterial clades (Macro- and Microcyanobacteria) suggests that these were likely significant contributors to primary productivity in marine habitats in the lead-up to the Great Oxygenation Event, and thus played a pivotal role in triggering the sudden increase in atmospheric oxygen.</p>

opencc-zeroMay 2024View details →
dryad40/100

Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad40/100

Data from: Stochastic character mapping, Bayesian model selection, and biosynthetic pathways shed new light on the evolution of habitat preference in cyanobacteria

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Code and data for Bayesian joint species distribution model selection for community-level prediction

Open the record for dataset details and reuse information.

publicNov 2023View details →
zenodo32/100

Pre-fitted Bayesian models for "Gene panel selection for targeted spatial transcriptomics"

<p>simulation_parameters_DARTFISH_slim.rds: Bayesian model fitted on the Zhang dataset.</p> <p>simulation_parameters_MERFISH_slim.rds: Bayesian model fitted on the Moffit dataset.</p> <p>simulation_parameters_osmFISH_slim.rds: Bayesian model fitted on the Codeluppi dataset.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

Input data for Bayesian and information theoretic model selection and similarity analysis

<p>This data serves as input to the codes found in the following repository https://github.com/MariaFMoralesOreamuno/Bayesian_Information_theoretic_model_selection.git</p> <p>&nbsp;</p>

openSep 2022View details →
zenodo32/100

cellsig: a Bayesian sparse random-effect model and large-scale human bulk transcriptional catalogue support cell-type marker selection

<p><strong>Human Bulk Cell-type Catalogue (HBCC): This database&nbsp;harmonises 1,435 samples of 67 cell types from 58 datasets which can be utilised in&nbsp;the cellsig method for estimating cell-type transcriptional profiles.</strong></p>

opencc-by-4.0Jan 2023View details →
dryad32/100

Selecting and averaging relaxed clock models in Bayesian tip dating of Mesozoic birds

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad32/100

Data from: Disentangling the formation of contrasting tree-line physiognomies combining model selection and Bayesian parameterization for simulation models

Open the record for dataset details and reuse information.

publicJan 2011View details →
dryad28/100

Data from: Posterior predictive Bayesian phylogenetic model selection

We present two distinctly different posterior predictive approaches to Bayesian phylogenetic model selection, and illustrate these methods using examples from green algal protein-coding cpDNA sequences and flowering plant rDNA sequences. The Gelfand-Ghosh (GG) approach allows dissection of an overall measure of model fit into components due to posterior predictive variance (Pm) and goodness-of-fit (Gm), which distinguishes this method from the posterior predictive P-value approach. The conditional predictive ordinate (CPO) method provides a site-specific measure of model fit useful for exploratory analyses and can be combined over sites yielding the log pseudomarginal likelihood (LPML), which is useful as an overall measure of model fit. CPO provides a useful cross-validation approach that is computationally efficient, requiring only a sample from the posterior distribution (no additional simulation is required). Both GG and CPO add new perspectives to Bayesian phylogenetic model selection based on the predictive abilities of models, and complement the perspective provided by the marginal likelihood (including Bayes Factor comparisons) based solely on the fit of competing models to observed data.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Bayesian model selection with BAMM: effects of the model prior on the inferred number of diversification shifts

1. Understanding variation in rates of speciation and extinction -- both among lineages and through time -- is critical to the testing of many hypotheses about macroevolutionary processes. BAMM is a flexible Bayesian framework for inferring the number and location of shifts in macroevolutionary rate across phylogenetic trees and has been widely used in empirical studies. BAMM requires that researchers specify a prior probability distribution on the number of diversification rate shifts before conducting an analysis. The consequences of this "model prior" for inference are poorly known but could potentially influence both the probability of accepting models that are more (high error rate) or less (low power) complex than the generating model. 2. The hierarchical Poisson process prior in BAMM reduces to a simple geometric distribution on number of rate shifts and we use this property to increase the efficiency of model selection with Bayes factors. Using BAMM v2.5, we analyzed phylogenies simulated with and without diversification heterogeneity across a broad range of prior parameterizations. We also assessed the impact of the model prior on MCMC convergence times and on diversification rate estimates. 3. For all simulation scenarios, model evidence (Bayes factor support) for the number of shifts is not sensitive to the choice of model prior over the wide range examined here. The best-supported model found using BAMM rarely includes spurious shifts (&lt;2% of all runs) when diversification models are selected using Bayes factors. BAMM was reliably able to infer the true number of diversification rate shifts across prior expectations that varied by three orders of magnitude. However, we find a strong effect of model prior on MCMC convergence properties: a flatter prior distribution (larger expected number of shifts) can dramatically increase the efficiency of the MCMC simulation. 4. Our results support the use of a liberal model prior in BAMM, as it reduces computation time without distorting the evidence for rate heterogeneity.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Bayesian model selection with BAMM: effects of the model prior on the inferred number of diversification shifts

Open the record for dataset details and reuse information.

publicAug 2017View details →
dryad28/100

Data from: Posterior predictive Bayesian phylogenetic model selection

Open the record for dataset details and reuse information.

publicOct 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record