Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
61
datasets available to search
ShareScore release 0.7.1
Dataset results
61 results for “Bayesian method”
Bayesian Methods for Ancestral State Reconstruction in Morphosyntax
<p>Supplementary files to accompany journal submission.</p> <p>Files are:</p> <p> </p> <p>tree.pdf - pdf consensus tree, for illustration</p> <p>data.txt - coding file</p> <p>TREE_Set.t - nexus format sample of trees.</p> <p>sources.pdf - source materials used for languages</p>
Datasets for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants"
<p>All data used for "Needle in a Bayes Stack: a Hierarchical Bayesian Method for Constraining the Neutron Star Equation of State with an Ensemble of Binary Neutron Star Post-merger Remnants", Criswell, A.W., et al. (2022). The code used to create the paper results from this data can be found at <a href="https://github.com/criswellalexander/hbpm_paper">https://github.com/criswellalexander/hbpm_paper</a> and the underlying software package can be found at <a href="https://github.com/criswellalexander/bayestack">https://github.com/criswellalexander/bayestack</a>.</p>
Code + simulated + publically accessable data for "Evaluating health facility access using Bayesian spatial models and location analysis methods"
<p># README</p> <p>These files contain r data objects and R files that represent the key details of the paper, "Evaluating health facility access using Bayesian spatial models and location analysis methods".</p> <p>The following datasources are available for simulation of some of the ideas in the paper.</p> <p>- dat_grid_sim: simulated data of the grid and grid cells<br> - dat_ohca_cv_sim: simulated data containing the cross validated test/training sets of OHCA data<br> - dat_ohca_sim: simulated OHCA event data<br> - dat_aed_sim: simulated AED location data<br> - dat_bldg_sim: simulated building location data<br> - dat_municipality_sim: simulated municipality information<br> - table_1: Table 1 information containing key demographic data</p> <p>These data were produced using the code in 01-create-sim-data.R, and one of the statistical models is demonstrated in 02-demo-inla-model.R</p> <p>In terms of the paper itself, the functions and code used in the manuscript are located in:</p> <p>* 01_tidy.Rmd - analysis code used to tidy up the data</p> <p>* 02_fit_fixed_all_cv.Rmd - analysis code used to place AEDs</p> <p>* 02_model.Rmd - analysis code used to fit the model in INLA</p> <p>* 03_manuscript.Rmd - Full code and text used to create the paper</p> <p>* 04_supp_materials.Rmd - full code and text used to create the supplementary materials</p> <p>The following files are a part of an R package "swatial" that was developed along with the paper. These files are:</p> <p>* DESCRIPTION</p> <p>* NAMESPACE</p> <p>* LICENSE</p> <p>* LICENSE.md</p> <p>* decay.R</p> <p>* spherical-distance.R</p> <p>* test-figure-data-matches.R</p> <p>* test-table-data-matches.R</p> <p>* testthat.R</p> <p>* tidy-inla.R</p> <p>* tidy-posterior-coefs.R</p> <p>* tidy-predictions.R</p> <p>* utils-pipe.R</p> <p>* All files that end in .Rd are documentation files for the functions.</p> <p>## Regarding data sources</p> <p>Census information for Ticino was transcribed from the Annual Statistical Report of Canton Ticino from years 2010 to 2015. This data was taken from their publicly accessible annual reports - for example: (https://www3.ti.ch/DFE/DR/USTAT/allegati/volume/ast_2015.pdf). The raw data was extracted from these annual reports, and placed into the file: "swiss_census_popn_2010_2015.xlsx". These data are put into analysis ready format in the file “01_tidy.Rmd”</p> <p>Housing and other relevant geospatial data can be accessed via http://map.housing-stat.ch/ and https://data.geo.admin.ch/. The maps of buildings from the REA (Register of Buildings and Dwellings) can be found here: https://map.geo.admin.ch/?zoom=11&bgLayer=ch.swisstopo.pixelkarte-grau&lang=en&topic=ech&layers=ch.bfs.gebaeude_wohnungs_register,ch.swisstopo.swissboundaries3d-gemeinde-flaeche.fill,ch.bfs.volkszaehlung-gebaeudestatistik_gebaeude,ch.bfs.volkszaehlung-gebaeudestatistik_wohnungen,ch.swisstopo.swissbuildings3d_1.metadata,ch.swisstopo.swissbuildings3d_2.metadata&E=2717616.28&N=1096597.25&catalogNodes=687,696&layers_timestamp=,,2016,2016,,&layers_visibility=true,false,false,false,false,false&layers_opacity=1,1,1,1,1,0.75</p> <p>For further enquiries on this data, contact the Swiss federal Office of Statistics at the details listed here: https://www.bfs.admin.ch/bfs/en/home/services/contact.html</p> <p>The shapefiles of the Comuni can be accessed here: https://www4.ti.ch/dfe/de/ucr/documentazione/download-file/?noMobile=1</p> <p>Data from the people living in the Municipalities in Ticino can be downloaded here: https://www3.ti.ch/DFE/DR/USTAT/index.php?fuseaction=dati.home&tema=33&id2=61&id3=65&c1=01&c2=02&c3=02</p> <p>## Future work</p> <p>In the future, these functions from the paper may be generalised and put into their own package. If that happens, this repository will be updated with a link to updated functions.</p>
Integrated Probabilistic Annotation (IPA): A Bayesian-based annotation method for metabolomic profiles integrating biochemical connections, isotope patterns and adduct relationships - Supplementary Data
<ol> <li>Supplementary_data_1: data and code for standards analysis and database update</li> <li>Supplementary_data_2.zip: data and code used for the generation of the synthetic experiment</li> <li>Supplementary_data_3.zip: data, code, and results of the <em>E. coli</em> dataset analysis</li> <li>Supplementary_data_4.zip: data, code, and results of the beer dataset analysis</li> <li>Supplementary_data_5.zip: data, code, and results of the comparison with xMSannotator</li> </ol>
Dataset and Code for "Bayesian spline method for assessing extreme loads on wind turbines"
<p>Here are two files. One file contains the datasets and the other contains the computer code used to generate the results in the paper, Lee, Byon, Ntaimo, and Ding, 2013, “Bayesian spline method for assessing extreme loads on wind turbines,” <em> Annals of Applied Statistics</em>, Vol. 7, pp. 2034-2061.</p>
Combining formal methods and Bayesian approach for inferring discrete-state stochastic models from steady-state data
<p>Model, data, and a script to a paper of respective name</p>
Craton Radial Anisotropy and Low Velocity Zones imaged with Bayesian and LSQR methods
<p>******************** README for Rad_anisotropy_BOYCE_GRL *****************************</p> <p>This repository contains data files, inversion software, outputs and plotting codes to accompany the following submitted manuscript:</p> <p>Boyce, A., Bodin, T., Durand, S., Soergel, D., Debayle, E. Seismic Evidence for Craton Formation by Underplating and Development of the MLD (submitted) Geophysical Research Letters.</p> <p>The Rad_anisotropy_BOYCE_GRL_V2.tar repo contains:<br> • LSQR_inversion - Least Squares inversion of synthetic and real data sets, all input and results and plotting files are included.<br> • SRF_Forward_modelling - Codes to make Axisem synthetic models and synthetic data to be inverted with Bayesian Code. Also included are Greens functions from Axisem simulations at 5s minimum period and python codes for forward modelling synthetic S-to-p reciever functions from the Axisem outputs.<br> • Bayesian_inversion - Bayesian inversion of synthetic and real data sets, all input files, processed files and plotting codes necessary for reconstructing the posterior distributions are included. Please see https://github.com/alistairboyce11/RJ_MCMC for an up-to-date distribution.<br> • MLD_compilation - Our MLD compilation "MLD_compilation_BOYCE_2023.xlsx" and codes to combine this with Fu et al., (GRL 2022) and plot Figure 2 of main manuscript<br> • Plotting_for_manuscript<br> - Radial_anisotropy_models - codes used to extract data and make plots for Figures S1--S5. Tomographic models available to download at https://ds.iris.edu/ds/products/emc/<br> - Figures for Figures 1, 3, 4 in main manuscript.</p>
Experimental and synthetic datasets supporting FITSA: Statistical analysis of fluorescence intensity transients with Bayesian methods
Open the record for dataset details and reuse information.
Commonly used Bayesian diversification methods lead to biologically meaningful differences in branch-specific rates on empirical phylogenies
Open the record for dataset details and reuse information.
Supplementary evaluation files for the paper: Grid-Based Bayesian Filtering Methods for Pedestrian Dead Reckoning Indoor Positioning Using Smartphones
<p>This package contains evaluation supplementary files for the paper: <em>Grid-Based Bayesian Filtering Methods for Pedestrian Dead Reckoning Indoor Positioning Using Smartphones</em> by Miroslav Opiela and František Galčík.</p> <p><strong>Contents: </strong></p> <ul> <li>ground_truth - real positions of checkpoints for given input files</li> <li>input - sensor measurements recordings with initial positions (also after floor transitions) and checkpoint labels </li> <li>maps - processed map models containing positions of points and connections (e.g., walls) in custom coordinate system. Reference to GNSS and map rotation is inducted in maps-meta.xml</li> <li>output - data processed by the localization system. JSON containing the applied method, its configuration, and all estimated positions. Errors for every folder are summarized in the csv file</li> <li>visualization - trajectories visualized for selected output files</li> <li>readme.txt - describes data formats used for particular files in this dataset and summarizes output files</li> </ul> <p><strong>Venues</strong></p> <p>Data are recorded in three buildings:</p> <ul> <li>codename: SA1, SA1_rotated - recorded by the author in the faculty building (Park Angelinum 9, 04001, Košice, Slovakia) using Lenovo tablet</li> <li>codename: AtlantisR0, AtlantisR-1, AtlantisR+1, AtlantisR+2 - the shopping mall Atlantis Le Centre (Boulevard Salvador Allende, 44800 Saint-Herblain, France). Dataset is from IPIN 2018 competition and loc_20180922_160206 is recorded by the author using Xiaomi Mi 5.</li> <li>codename: CNR_0, CNR_1, CNR_2 - the research institute building CNR (Via Giuseppe Moruzzi, 56127 Pisa, Italy). Dataset is from IPIN 2019 competition. </li> </ul> <p><strong>Used datasets</strong></p> <p>A subset of input data is derivated from available logfiles provided by organizers of IPIN 2018 and IPIN 2019 competitions:</p> <ul> <li>Jimenez, A.R.; Mendoza-Silva, G.M.; Ortiz, M.; Perez-Navarro, A.; Perul, J.; Seco, F.; Torres-Sospedra, J. Datasets and Supporting Materials for the IPIN 2018 Competition Track 3 (Smartphone-based, off-site). <a href="http://dx.doi.org/10.5281/zenodo.2823964">http://dx.doi.org/10.5281/zenodo.2823964</a></li> <li>Jiménez, A. R.; Perez-Navarro, A.; Crivello, A.; Mendoza-Silva, G.; Ortiz, M.; Perul, J.; Seco, F. and Torres-Sospedra, J. Datasets and Supporting Materials for the IPIN 2019 Competition Track 3 (Smartphone-based, off-site), Zenodo 2019. <a href="http://dx.doi.org/10.5281/zenodo.3606765">http://dx.doi.org/10.5281/zenodo.3606765</a> </li> </ul> <p><strong>Funding</strong></p> <p>The work was partially supported by the Slovak Grant Agency of the Ministry of Education and Academy of Science of the Slovak Republic under grant no. 1/0056/18 and by the Slovak Research and Development Agency under the contract no. APVV-15-0091.</p> <p><strong>Contact</strong></p> <p>For any further questions, please contact:</p> <p>Miroslav Opiela, miroslav.opiela@upjs.sk Institute of Computer Science, Faculty of Science, P. J. Šafárik University (UPJS), Košice, Slovakia</p>
Data from: ClonEstiMate, a Bayesian method for quantifying rates of clonality of populations genotyped at two-time steps
Partial clonality is commonly used in Eukaryotes and has large consequences for their evolution and ecology. Assessing accurately the relative importance of clonal versus sexual reproduction matters for studying and managing such species. Here, we proposed a Bayesian approach, ClonEstiMate, to infer rates of clonality c from populations sampled twice over a short time interval, ideally one generation time. The method relies on the likelihood of the transitions between genotype frequencies of ancestral and descendent populations, using an extended Wright-Fisher model explicitly integrating reproductive modes. Our model provides posterior probability distribution of inferred c, given the assumed rates of mutation, as well as inbreeding and selfing when occurring. Tested under various conditions, this model provided accurate inferences of c, especially when the amount of information was modest, i.e. low sample sizes, few loci, low polymorphism and strong linkage disequilibrium. Inferences remained robust when mutation models and rates were misinformed. However, the method was sensitive to moderate frequencies of null alleles and when the time interval between required samplings exceeding two generations. Misinformed rates on mating modes (inbreeding and selfing) also resulted in biased inferences. Our method was tested on eleven datasets covering five partially clonal species, for which the extent of clonality was formerly deciphered. It delivered highly consistent results with previous information on the biology of those species. ClonEstiMate represents a powerful tool for detecting and inferring clonality in finite populations, genotyped with SNPs or microsatellites. It is freely available at http://https://w">https://w w w 6.rennes.inra.fr / igepp_eng/ Productions/ Software.
Data from: A Bayesian method for the joint estimation of outcrossing rate and inbreeding depression
The population outcrossing rate (t) and adult inbreeding coefficient (F) are key parameters in mating system evolution. The magnitude of inbreeding depression as expressed in the field can be estimated given t and F via the method of Ritland (1990). For a given total sample size, the optimal design for the joint estimation of t and F requires sampling large numbers of families (100-400) with fewer offspring (1-4) per family. Unfortunately, the standard inference procedure (MLTR) yields significantly biased estimates for t and F when family sizes are small and maternal genotypes are unknown (a common occurrence when sampling natural populations). Here, we present a Bayesian method implemented in the program BORICE that effectively estimates t and F when family sizes are small and maternal genotype information is lacking. BORICE should enable wider use of the Ritland approach for field-based estimates of inbreeding depression. As proof of concept, we estimate t and F in a natural population of Mimulus guttatus. In addition, we describe how individual maternal inbreeding histories inferred by BORICE may prove useful in studies of inbreeding and its consequences.
Data from: Quantifying demographic uncertainty: Bayesian methods for integral projection models (IPMs)
Integral projection models (IPMs) are a powerful and popular approach to modeling population dynamics. Generalized linear models form the statistical backbone of an IPM. These models are typically fit using a frequentist approach. We suggest that hierarchical Bayesian statistical approaches offer important advantages over frequentist methods for building and interpreting IPMs, especially given the hierarchical nature of most demographic studies. Using a stochastic IPM for a desert cactus based on a 10-year study as a worked example, we highlight the application of a Bayesian approach for translating uncertainty in the vital rates (e.g., growth, survival, fertility) to uncertainty in population-level quantities derived from them (e.g., population growth rate). The best-fit demographic model, which would have been difficult to fit under a frequentist framework, allowed for spatial and temporal variation in vital rates and correlated responses to temporal variation across vital rates. The corresponding posterior probability distribution for the stochastic population growth rate (λS) indicated that, if current vital rates continue, the study population will decline with nearly 100% probability. Interestingly, less-supported candidate models that did not include spatial variance and vital rate correlations gave similar estimates of λS. This occurred because the best-fitting model did a much better job of fitting vital rates to which the population growth rate was weakly sensitive. The cactus case study highlights several advantages of Bayesian approaches to IPM modeling, including that they: (1) provide a natural fit to demographic data, which are often collected in a hierarchical fashion (e.g., with random variance corresponding to temporal and spatial heterogeneity); (2) seamlessly combine multiple data sets or experiments; (3) readily incorporate covariance between vital rates; and, (4) easily integrate prior information, which may be particularly important for species of conservation concern where data availability may be limited. However, constructing a Bayesian IPM will often require the custom development of a statistical model tailored to the peculiarities of the sampling design and species considered; there may be circumstances under which simpler methods are adequate. Overall, Bayesian approaches provide a statistically sound way to get more information out of hard-won data, the goal of most demographic research endeavors.
Data from: Estimating age and age class of harvested hog deer from eye lens mass using frequentist and Bayesian methods
Estimation of the age or age class of harvested animals is often necessary to interpret the condition and dynamics of wildlife populations. The mammalian eye lens continues to grow until death and hence the dry mass of the eye lens has commonly been used to estimate the age of mammals. The method requires the relationship between eye lens mass and age to be parameterized using individuals of known age. However, predicting age is complicated by the curvilinear relationship between eye lens mass and age. We used frequentist and Bayesian methods to predict the ages and age classes of harvested hog deer Axis porcinus from eye lens mass. Deer were tagged as calves and harvested 4–177 months later in southeastern Australia. Lenses were extracted, fixed and oven-dried. Of the five growth models evaluated, the Lord model best described the relationship between age and eye lens dry mass (R2 = 95%). The precision of age predictions obtained using the Lord model in a Bayesian mode of inference decreased with increasing eye lens dry mass, with the size of the 95% CI equaling or exceeding predicted age for hog deer > 6 years. However, most predictions of hog deer age will have reasonable precision because few animals > 6 years are harvested. Linear discriminant analysis had high predictive power for classifying hog deer to four widely-used age classes (juvenile, yearling, prime-age and senescent). The Bayesian method is recommended for inverse non-linear prediction of age and the frequentist linear discriminant analysis method is recommended for estimating age class. We provide tables of correspondence between hog deer eye lens dry mass and predicted age and age class. Our statistical methods can be used to estimate age and age class for other mammalian species, including from other ageing techniques such as tooth eruption-wear criteria.
Data from: A rapid and scalable method for multilocus species delimitation using Bayesian model comparison and rooted triplets
Multilocus sequence data provide far greater power to resolve species limits than the single locus data typically used for broad surveys of clades. However, current statistical methods based on a multispecies coalescent framework are computationally demanding, because of the number of possible delimitations that must be compared and time-consuming likelihood calculations. New methods are therefore needed to open up the power of multilocus approaches to larger systematic surveys. Here, we present a rapid and scalable method that introduces 2 new innovations. First, the method reduces the complexity of likelihood calculations by decomposing the tree into rooted triplets. The distribution of topologies for a triplet across multiple loci has a uniform trinomial distribution when the 3 individuals belong to the same species, but a skewed distribution if they belong to separate species with a form that is specified by the multispecies coalescent. A Bayesian model comparison framework was developed and the best delimitation found by comparing the product of posterior probabilities of all triplets. The second innovation is a new dynamic programming algorithm for finding the optimum delimitation from all those compatible with a guide tree by successively analyzing subtrees defined by each node. This algorithm removes the need for heuristic searches used by current methods, and guarantees that the best solution is found and potentially could be used in other systematic applications. We assessed the performance of the method with simulated, published, and newly generated data. Analyses of simulated data demonstrate that the combined method has favorable statistical properties and scalability with increasing sample sizes. Analyses of empirical data from both eukaryotes and prokaryotes demonstrate its potential for delimiting species in real cases.
Figure 10 in Global phylogeny of Hadrosauridae (Dinosauria: Ornithopoda) using parsimony and Bayesian methods
Figure 10. Time-calibrated phylogram based on the strict reduced consensus tree derived from maximum parsimony analysis of hadrosaurid relationships. The geochronological ages are taken from Gradstein, Ogg & Smith (2004).
Figure 8 in Global phylogeny of Hadrosauridae (Dinosauria: Ornithopoda) using parsimony and Bayesian methods
Figure 8. Bayesian consensus tree of the analysis with variable rates of character change showing the phylogenetic relationships of 53 iguanodontians. At each node, the decimal number above a branch indicates a posterior probability, whereas the decimal number below represents a P value of topology-dependant permutation tail probability (T-PTP) analysis.
Figure 7 in Global phylogeny of Hadrosauridae (Dinosauria: Ornithopoda) using parsimony and Bayesian methods
Figure 7. Bayesian consensus tree of the analysis with equal rates of character change showing the phylogenetic relationships of 53 iguanodontians. At each node, the decimal number above a branch indicates a posterior probability, and the decimal number below represents a P value of topology-dependant permutation tail probability (T-PTP) analysis.
Figure 4 in Global phylogeny of Hadrosauridae (Dinosauria: Ornithopoda) using parsimony and Bayesian methods
Figure 4. Bayesian information criterion (BIC) plots resulting from the clustering model analysis implemented in Mclust (within the R statistical package). A, BIC plot of two clustering models for a one-dimensional data set (i.e. a character quantified with a linear measurement). B, BIC plot for various clustering models of NMMDS derived from a geodesic distance analysis (GDA) dissimilarity matrix. Indentifiers of clustering models indicate the type of distribution (spherical, diagonal, or ellipsoidal), volume (equal or variable), shape (equal or variable), and orientation (not applicable, coordinate axes, equal, or variable) of the clusters: EEE = ellipsoidal, equal volume, shape and orientation; EEI = diagonal, equal volume, variable shape; EEV = ellipsoidal, equal volume and shape, variable orientation; EII, spherical, equal volume and shape; EVI = diagonal, equal volume, variable shape; VEI = diagonal, variable volume, equal shape; VEV = ellipsoidal, variable volume, equal shape; VII, spherical, variable volume, equal shape; VVI = diagonal, variable volume and shape; VVV = ellipsoidal, variable volume and shape (Fraley & Raftery, 2003).
Figure 5 in Global phylogeny of Hadrosauridae (Dinosauria: Ornithopoda) using parsimony and Bayesian methods
Figure 5. Strict consensus tree of the 160 most parsimonious trees resulting from the parsimony analysis of 53 iguanodontian taxa. At each node, the pair of numbers separated by a slash above a branch represents, from left to right, a decay index and a bootstrap proportion. Bootstrap proportions lower than 50 are indicted by a hyphen. The decimal number in italics that appears below a branch represents the P value of topology-dependant permutation tail probability (T-PTP) analysis.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.