Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
203
datasets available to search
ShareScore release 0.7.1
Dataset results
203 results for “Bayesian model”
Bayesian analysis of the equation of state of quantum chromodynamics from a holographic model
<p>Prior and posterior samples obtained from a Bayesian analysis of the equation of state of quantum chromodynamics (QCD) within a holographic Einstein-Maxwell-Dilaton model, constrained by state-of-the art lattice QCD results at a vanishing net density of baryons.</p> <p>Samples contain metadata, model parameters, and model predictions for the location of the QCD critical point.</p> <p>Supplement to <a title="Bayesian location of the QCD critical point from a holographic perspective" href="https://arxiv.org/abs/2309.00579">arXiv:2309.00579</a>.</p>
Atmospheric Halocarbon Observations at Beromünster, Switzerland, and Bayesian Inverse Modeling to assess Emissions
<p>Atmospheric halocarbon (CFCs, halons, HCFCs, HFCs, PFCs, SF<sub>6</sub>, NF<sub>3</sub>, HFOs) and carbon monoxide (CO) observations (mole fractions) from the tall tower site at Beromünster, Switzerland (47.2 °N, 8.2 °E, 797 m a.s.l., 212 m a.g.l.), covering the period September 2019 to September 2020. The halocarbon measurements were conducted using a Medusa pre-concentration unit, coupled to gas chromatography (Agilent 6890N) and mass spectrometry (Agilent 5975, GC-MS).</p> <p>For further details see: Miller, B. R., Weiss, R. F., Salameh, P. K., Tanhua, T., Greally, B. R., Mühle, J., and Simmonds, P. G.: Medusa: A Sample Preconcentration and GC/MS Detector System for in Situ Measurements of Atmospheric Trace Halocarbons, Hydrocarbons, and Sulfur Compounds, Anal. Chem., 80, 1536–1545, https://doi.org/10.1021/ac702084k, 2008).</p> <p>The data format follows that used within the AGAGE network (see AGAGE data archive: <a href="http://agage.mit.edu/data/agage-data">http://agage.mit.edu/data/agage-data</a>).</p> <p>Data results for the Bayesian inversion conducted based on the measurement data from Beromünster to assess Swiss halocarbon emissions. Files are provided in netCDF format for the 28 individual substances discussed in (Rust, D. et al., 2022, <em>Swiss halocarbon emissions for 2019 to 2020 assessed from regional atmospheric observations</em>, Atmospheric Chemistry and Physics). Each file contains the a priori and a posteriori emissions as used or calculated in the Bayesian inversion. Data are provided on the grid used in the inversion (irregular longitude/latitude). Metadata are included as netCDF attributes. The netCDF files follow the CF conventions and are readable with any netcdf interface/tool.</p>
Global Surface Ozone Concentration Dataset 1990-2017 Mapped at Fine Resolution through the Bayesian Maximum Entropy Data Fusion of Observations and Model Output
<p>This global surface ozone concentration dataset corresponds to the data developed in this paper:</p> <p>DeLang, M. N., J. S. Becker, K.-L. Chang, M. L. Serre, O. R. Cooper, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, S. Cleland, E. Collins, M. Brauer, and J. J. West (2021) Mapping yearly fine resolution global surface ozone through the Bayesian Maximum Entropy data fusion of observations and model output for 1990-2017, <em>Environmental Science & Technology</em>, 55, 4389-4398, doi: 10.1021/acs.est.0c07742.</p> <p>Ozone concentrations are estimated as described in the paper, with output shown for the Ozone Season Daily Maximum 8-hr metric (OSDMA8) for each year between 1990 and 2017, at 0.1 degree spatial resolution. Ozone is estimated through data fusion of output from several global models, with observations of ozone collected by TOAR. The data fusion involves application of the M3Fusion method to create a multi-model composite of several global models, followed by BME data fusion, as described in the paper. </p> <p>The *.nc file contains the latitude, longitude, ozone concentration estimate, and estimated variance for each 0.1 x 0.1 degree grid cell.</p> <p>Please contact Jason West (jasonwest@unc.edu) with questions about the dataset. We'd like to hear from you to know how you're using the data!</p> <p> </p> <p> </p>
GrainLearning: A Bayesian uncertainty quantification toolbox for discrete and continuum numerical models of granular materials
GrainLearning is a Bayesian uncertainty quantification and propagation toolbox for computer simulations of granular materials. The software is primarily used to infer and quantify parameter uncertainties in computational models of granular materials from observation data, also known as inverse analyses or data assimilation. Implemented in Python, GrainLearning can be loaded into a Python environment to process the simulation and observation data, or alternatively, as an independent tool where simulation runs are done separately, e.g., via a shell script.
Example code and data for ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework
<p>This repository contains an R script (grouse_example.R) and data (grouse_data.csv) used to reproduce the grouse abundance analysis described in Kellner, K. F., et al. (2021) ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework. Methods in Ecology and Evolution. The R script requires installation of the ubms R package, which can be obtained from CRAN (https://cran.r-project.org/package=ubms).</p> <p>The repository also contains an additional example occupancy analysis (occupancy_example.R) using the crossbill dataset included with the unmarked R package.</p>
SEDflow: Accelerated Bayesian SED Modeling using Amortized Neural Posterior Estimation
<p><a href="http://changhoonhahn.github.io/SEDflow">SEDflow</a> is an accelerated Bayesian SED modeling method that uses the <a href="https://ui.adsabs.harvard.edu/abs/2022arXiv220201809H/abstract">Hahn et al. (2022a)</a> PROVABGS SED model and Amortized Neural Posterior Estimation (ANPE) to derive posterior probability distributions of galaxy properties from optical photometry. SEDflow is<span class="math-tex">\(10^5\times\)</span> faster than conventional Markov Chain Monte Carlo sampling methods and takes ~1 second per galaxy to obtain posteriors. This repository includes all of the data used to train, validate, and test SEDflow.</p> <p>This repository also includes a value-added catalog with detailed physical properties of 33,884 galaxies in the NASA-Sloan Atlas (http://www.nsatlas.org/). The properties are inferred from optical photometry in the <em>u, g, r, i, z</em> bands using SEDflow. For more details on this catalog and SEDflow see the <a href="http://changhoonhahn.github.io/SEDflow">documentation</a> and Hahn & Melchior (2022). </p> <p>For each galaxy, the catalog provides posteriors of: </p> <ul> <li>log_mstar: log10 of stellar mass</li> <li>log_sfr_1gyr: log10 of average star formation rate over 1Gyr</li> <li>log_z_mw: log10 of mass-weighted metallicity</li> <li>beta1, beta2, beta3, beta4: coefficients of the non-negative matrix factorization (NMF) star formation history basis functions</li> <li>fburst: fraction of stellar mass formed by a starburst event</li> <li>tburst: time of the starburst event</li> <li>log_gamma1, log_gamma2: log10 of coefficients of the NMF metallicity history basis functions</li> <li>tau_bc: birth cloud optical depth</li> <li>tau_ism: diffuse dust optical depth</li> <li>n_dust: Calzetti (2001) dust index</li> </ul>
"Chronomodel" Bayesian chronological models for East Borneo, based on data from the Liang Abu and Kimanis sites
<p>Bayesian chronological models generated using the <a href="https://chronomodel.com"><em>ChronoModel</em></a> software East Borneo (Indonesia), based on data from the Liang Abu and Kimanis (Arifin, 2017) archaeological sites.</p> <p>Two models were generated:</p> <ul> <li> a “<strong>conservative</strong>” model, observing the Bayesian approach and the distinction between<br> prior and posterior information;</li> <li>a “<strong>restricted</strong>” model: excluding possible outliers and without application of a “Fresh-<br> water reservoir effect” correction.</li> </ul> <p>Four files are provided:</p> <ul> <li>abu-kimanis-conservative-model.chr: model specification for the “conservative” model</li> <li>abu-kimanis-conservative-model_synthetic-stats-table.csv: results for the “conservative” model</li> <li>abu-kimanis-restricted-model.chr: model specification for the “restricted” model</li> <li>abu-kimanis_restricted-model_synthetic-stats-table.csv: results for the “restricted” model</li> </ul> <p>The .chr files can be open and edited using the <em>ChronoModel</em> software.</p>
Minimal dataset for "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models"
<p>This repository contains a minimal data set to reproduce all results that don't compromise the privacy concerns for the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> <br> The repository contains the following data:</p> <ul> <li>adaptscore_acute.csv <ul> <li>A csv file that contains the estimated adaptation scores for the acute data set with HLA I model.</li> </ul> </li> <li>adaptscore_leftout.csv <ul> <li>A csv file that contains the estimated adaptation scores for the leftout data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training.csv <ul> <li>A csv file that contains the estimated adaptation scores for the traininig data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training_hla1_without_clin.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the HLA I model (via cross-validation)</li> </ul> </li> <li>adaptscore_training_seed2.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the joint HLA I and HLA II model via cross-validation with another seed</li> </ul> </li> </ul>
Waveform data for centroid moment tensor solutions presented in publication "Bayesian seismic source inversion with a 3-D Earth model of the Japanese islands"
<p>The dataset includes waveform data for centroid moment tensor solutions inferred using Hamiltonian Monte Carlo and a 3-D Earth model in the Japanese islands. The data are provided as Green's strains at the maximum-likelihood location (indicated in the title of each text file) for all study events inverted at different periods. Inversion period is also indicated in the title. All the data are filtered between 15 s and 80 s. Additionally we provide a Python code to obtain displacement from strains given a moment tensor.</p>
The molecular architecture of the yeast spindle pole body core determined by Bayesian integrative modeling
<p>This repository pertains to the molecular architecture of the yeast spindle pole body (SPB), the structural and functional equivalent of the metazoan centrosome. Data from in vivo FRET and yeast two-hybrid, along with SAXS, X-ray crystallography, and electron microscopy were integrated by a Bayesian structure modeling approach.</p> <p>For more information about how to reproduce this modeling, see the <a href="https://salilab.org/spb/">Sali lab website</a> or the README file.</p>
Research Data/Code for "Scale-bridging within a complex model hierarchy for investigation of a metal-fueled circular energy economy by use of Bayesian model calibration with model error quantification"
<p>This repository contains research data and code for supplementing the manuscript <br>"Scale-bridging within a complex model hierarchy for investigation of a metal-fueled circular energy economy by use of Bayesian model calibration with model error quantification" <br>by L. Gossel, E. Corbean, S. Dübal, P. Brand, M. Fricke, H. Nicolai, C. Hasse, S. Hartl, S. Ulbrich, and D. Bothe. </p> <p>There is a corresponding preprint available on Arxiv: https://doi.org/10.48550/arXiv.2404.13092</p> <p><br>Users are referred to the manuscript for background information. This repository shall enable reproduction of the reported results and does not stand alone. </p> <p>Please read important information in the README in the top-level directory. </p> <p>Funded by the Hessian Ministry of Higher Education, Research, Science and the Arts - cluster project Clean Circles. </p>
Dataset for Bayesian parametric models for survival prediction in medical applications
<p><strong>Data Source</strong></p> <p>The data for these experiments were derived from these sources:</p> <p>* Hosmer Jr DW, Lemeshow S, May S. Applied Survival Analysis: Regression Modeling of Time-to-Event Data. 2nd ed: John Wiley & Sons; 2008.</p> <p>* Jd K, Prentice R. The statistical analysis of failure time data. New York: John Wiley and Sons; 1980.</p> <p>* Fleming T, Harrington D. Counting Processes and Survival Analysis: John Wiley & Sons; 1991.</p> <p> </p> <p>The raw data was downloaded from web archive.</p> <p>https://web.archive.org/web/20170114043458/http://www.umass.edu/statdata/statdata/data/</p> <p><strong>Contents</strong></p> <p>Each folder contains the original data, a textfile with a description of the data, and the pre-processed version with one-hot encoded variables. An additional YAML file is included with the list of included variables, name of the time and censor variable, name of continuous variables and splitting and partitioning information.</p>
Data related to the manuscript "Bayesian Calibration and Validation of a Large-scale and Time-demanding Sediment Transport Model"
<p>1) Riverbed_Elevation_Measurements.txt<br> Description: Measured riverbed geometry of available years<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2002 [m asl], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation <br> 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 2) Hydro_FT_2D_manual.txt<br> Description: Simulation results of the manually calibrated full model<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>3.1) Hydro_FT_2D_CollocationPointBase.txt<br> Description: Parameter combinations of the collocation point base for each of the 20 simulations conducted with the full model to <br> construct the surrogate<br> Rows: Critical Shields parameter, Grain Roughness, Grain Size distribution</p> <p>3.2) Hydro_FT_2D_CollocationResults.txt<br> Description: Simulation results of the 20 simulations conducted with the full model at the collocation points<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig<br> [m asl], Elevations 2010 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig [m asl], Elevations 2013 [m asl] of<br> simulation 1 through 20<br> ----------------------------------------------------------------------------------------------------------------------------<br> 4.1) aPC_MC_N_Combinations_Weights_prior.txt<br> Description: ID of prior MC runs with tested parameter combinations and corresponding importance weights<br> Rows: ID of MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.2) aPC_MC_2005_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2005<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of MC run 1 through 100,000<br> 4.3) aPC_MC_2010_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2010<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of MC run 1 through 100,000<br> 4.4) aPC_MC_2013_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2013<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2013 [m asl] of MC run 1 through 100,000<br> <br> 4.5) aPC_MC_N_Combinations_Weights_posterior.txt<br> Description: ID of accepted (posterior) MC runs with tested parameter combinations and corresponding importance weights<br> Rows: ID of accepted MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.6) aPC_MC_2005_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2005<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of accepted MC run 1 through 857<br> 4.7) aPC_MC_2010_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2010<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of accepted MC run 1 through 857<br> 4.8) aPC_MC_2013_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2013<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2013 [m asl] of accepted MC run 1 through 857<br> ----------------------------------------------------------------------------------------------------------------------------<br> 5) aPC_MAP.txt<br> Description: Simulation results conducted with the stochastically calibrated aPC surrogate model using the MAP parameter <br> combination<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>6) Hydro_FT_2D_MAP.txt<br> Description: Simulation results conducted with the stochastically calibrated full model using the MAP parameter combination<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 7) dz.txt<br> Description: Riverbed Evolution for all nodes in the section of interest (n=1138) obtained with differently calibrated models for all <br> considered time periods<br> Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m], <br> Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p>8) dz_CalibrationNodes.txt<br> Description: Riverbed Evolution for calibration nodes (n=204) obtained with differently calibrated models for all considered time<br> periods<br> Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m], <br> Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p> </p>
Computed results for Bayesian genome scale modelling temperature effect on yeast metabolism
<p>This repository contains the computed results for reproducing the figures in the manuscript "Li G., et al. Bayesian genome scale modelling identifies thermal determinants of yeast metabolism". The scripts can be found in Github (<a href="https://github.com/Gangl2016/BayesianGEM">https://github.com/Gangl2016/BayesianGEM</a>)</p>
Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant - Datasets, Trained Models, BNN Samples, and MCMC Chains
<p>We publish the training/validation/test datasets, trained model weights, configuration files, Bayesian neural network samples, and MCMC chains used to produce the figures in the LSST DESC paper, "Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant." They are formatted to be used with the DESC package "H0rton" (<a href="https://github.com/jiwoncpark/h0rton">https://github.com/jiwoncpark/h0rton</a>). Additional descriptions can be found in the README. Please contact Ji Won Park (@jiwoncpark) on GitHub or <a href="https://github.com/jiwoncpark/h0rton/issues">make an issue</a> for any questions.</p>
Code and data for Bayesian joint species distribution model selection for community-level prediction
<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction." Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript. Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>
Data for: Reintroduced Oriental stork bayesian hierarchical model data
<p>Long-lived territorial bird populations often consist of a few territorial breeding adults and many non-breeding individuals. Some populations are threatened by anthropogenic activities, because of human conflicts for high-quality breeding habitat. Therefore, habitat restoration projects have been widely implemented to improve avian population status. In conjunction with habitat restoration, conservation translocations have been increasingly implemented. Adequate non-breeder survival can be a key factor in the success of these attempts because non-breeding birds may represent reservoirs for the replacement of breeders. The maintenance of breeding pair numbers is also influenced by the transition rate of non-breeders to breeders. The reintroduction of Oriental stork (<em>Ciconia boyciana</em>), a long-lived, territorial, endangered species, was initiated in Japan in 2005 using captive birds in hopes of increasing the population's use of restored habitat. Our objective of this study was to elucidate the factors determining reintroduced stork survival and recruitment to the breeding populations. We estimated the survival rate and breeding participation rate by sex, age, generation, wild-born or not, haplotypes, and breeding status in storks reintroduced during 2005–2022 using Bayesian hierarchical models. There was no significant difference in survival rate between non-breeders and breeders. However, the survival rate was lower in wild-born birds than released birds, which may be related to the longer-distance natal dispersal of new generations. Accelerated habitat restoration around breeding areas and preventive measures for collision with human-built structures should be implemented for the sustained growth of reintroduced populations. A low survival rate was also detected for a specific mtDNA haplotype that accounts for the majority of the reintroduced population. This phenomenon might be explained by mtDNA-encoded mutations. Moreover, captive breeding and release history might contribute to an increase in the proportion of this haplotype in the wild.</p>
Resources for: Spatio-temporal integrated Bayesian species distribution models reveal lack of broad relationships between traits and range shifts
<p><strong>Aim</strong>: Climate change and habitat loss or degradation are some of the greatest threats that species face today, often resulting in range shifts. Species traits have been discussed as important predictors of range shifts, with the identification of general trends being of great interest for conservation efforts. However, studies reviewing relationships between traits and range shifts have questioned the existence of such generalized trends, due to mixed results and weak correlations, as well as analytical shortcomings. The aim of this study was to test this relationship empirically, using analytical approaches that account for common sources of bias when assessing range trends.<br><strong>Location</strong>: Tanzania, East Africa.<br><strong>Time period</strong>: 1980-1999 and 2000-2020.<br><strong>Major taxa studied</strong>: 57 savannah specialist birds found in Tanzania, belonging to 26 families and 11 orders.<br><strong>Methods</strong>: We applied recently developed integrated spatio-temporal species distribution models in R-INLA, combining citizen science and bird atlas data to estimate ranges of species, quantify range shifts, and test the predictive power of traditional trait groups, as well as exposure-related and sensitivity traits. We based our study on 40 years of bird observations in East African savannahs, a biome that has experienced increasing climatic and non-climatic pressures over recent decades. We correlated patterns of change with species traits.<br><strong>Results</strong>: We find indications of relationships identified by previous research, but low average explanatory power of traits from an ecological perspective, confirming the lack of meaningful general associations. However, our analysis finds compelling species-specific results.<br><strong>Main conclusions</strong>: We highlight the importance of individual assessments, while demonstrating the usefulness of our analytical approach for analyses of range shifts.</p>
genomesizeR: databases and bayesian models
<p>This archive contains the reference databases as well as the bayesian models used by the R package genomesizeR.</p>
Performance of akaike information criterion and bayesian information criterion in selecting partition models and mixture models
<p>In molecular phylogenetics, partition models and mixture models provide different approaches to accommodating heterogeneity in genomic sequencing data. Both types of models generally give a superior fit to data than models that assume the process of sequence evolution is homogeneous across sites and lineages. The Akaike Information Criterion (AIC), an estimator of Kullback-Leibler divergence, and the Bayesian Information Criterion (BIC) are popular tools to select models in phylogenetics. Recent work suggests AIC should not be used for comparing mixture and partition models. In this work, we clarify that this difficulty is not fully explained by AIC misestimating the Kullback-Leibler divergence. We also investigate the performance of the AIC and BIC by comparing amongst mixture models and amongst partition models. We find that under non-standard conditions (i.e. when some edges have a small expected number of changes), AIC underestimates the expected Kullback-Leibler divergence. Under such conditions, AIC preferred the complex mixture models and BIC preferred the simpler mixture models. The mixture models selected by AIC had a better performance in estimating the edge length, while the simpler models selected by BIC performed better in estimating the base frequencies and substitution rate parameters. In contrast, AIC and BIC both prefer simpler partition models over more complex partition models under non-standard conditions, despite the fact that the more complex partition model was the generating model. We also investigated how mispartitioning (i.e. grouping sites that have not evolved under the same process) affects both the performance of partition models compared to mixture models and the model selection process. We found that as the level of mispartitioning increases, the bias of AIC in estimating the expected Kullback-Leibler divergence remains the same, and the branch lengths and evolutionary parameters estimated by partition models become less accurate. We recommend that researchers be cautious when using AIC and BIC to select among partition and mixture models; other alternatives, such as cross-validation and bootstrapping should be explored, but may suffer similar limitations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.