Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
70
datasets available to search
ShareScore release 0.7.1
Dataset results
70 results for “bayesian approach”
Molecular dating for phylogenies containing a mix of populations and species by using Bayesian and RelTime approaches
Open the record for dataset details and reuse information.
The optimal time to approach an unfamiliar object: A Bayesian model
Open the record for dataset details and reuse information.
MCMC data for A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection
<p>These are unprocessed Markov-chain Monte-Carlo datasets accompanying the manuscript "A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection"</p>
Data from: Comparing traditional and Bayesian approaches to ecological meta-analysis
<p>1. Despite the wide application of meta-analysis in ecology, some of the traditional methods used for meta-analysis may not perform well given the type of data characteristic of ecological meta-analyses.</p> <p>2. We reviewed published meta-analyses on the ecological impacts of global climate change, evaluating the number of replicates used in the primary studies (ni) and the number of studies or records (k) that were aggregated to calculate a mean effect size. We used the results of the review in a simulation experiment to assess the performance of conventional frequentist and Bayesian meta-analysis methods for estimating a mean effect size and its uncertainty interval.</p> <p>3. Our literature review showed that ni and k were highly variable, distributions were right-skewed, and were generally small (median ni =5, median k=44). Our simulations show that the choice of method for calculating uncertainty intervals was critical for obtaining appropriate coverage (close to the nominal value of 0.95). When k was low (<40), 95% coverage was achieved by a confidence interval based on the t-distribution that uses an adjusted standard error (the Hartung-Knapp-Sidik-Jonkman, HKSJ), or by a Bayesian credible interval, whereas bootstrap or z-distribution confidence intervals had lower coverage. Despite the importance of the method to calculate the uncertainty interval, 39% of the meta-analyses reviewed did not report the method used, and of the 61% that did, 94% used a potentially problematic method, which may be a consequence of software defaults.</p> <p>4. In general, for a simple random-effects meta-analysis, the performance of the best frequentist and Bayesian methods were similar for the same combinations of factors (k and mean replication), though the Bayesian approaches had higher than nominal (>95%) coverage for the mean effect when k was very low (k<15). Our literature review suggests that many meta-analyses that used z-distribution or bootstrapping confidence intervals may have over-estimated the statistical significance of their results when the number of studies was low; more appropriate methods need to be adopted in ecological meta-analyses.</p>
Data from: A Bayesian hierarchical approach to quantifying stakeholder attitudes toward conservation in the presence of reporting error
Stakeholder support is vital for achieving conservation success, yet there are few reliable mechanisms to monitor stakeholder attitudes towards conservation. Importantly, few approaches account for bias arising from reporting errors; that is, reporting a positive attitude towards conservation when the respondent actually does not have one (a false positive error), or not reporting a positive attitude when the respondent is positive towards conservation (a false negative error). We borrow from developments in applied conservation science to use a Bayesian hierarchical model to quantify stakeholder attitudes as the probability of having a positive attitude towards wildlife, notionally (or in abstract terms) and at localized scales. The model allows us to assess stakeholder attitudes, and factors influencing these attitudes, while accounting for false negative and false positive reporting errors. We show through simulations that this method has lower bias than naïve estimates of the proportion of respondents who are positive towards wildlife, or Likert‐scores. We demonstrate the utility of the model by applying it to questionnaire surveys on Asian elephants Elephas maximus in the Kaziranga–Karbi Anglong landscape, Northeast India. After accounting for reporting errors, we estimated the probability of being positive towards elephants notionally as 0.85; at a localized scale, however, the proportion of respondents that were positive towards elephants was 50%. In comparison, without accounting for reporting errors, the proportion of respondents professing positive attitudes towards elephants in at least one of the certain questions, was 0.69 and 0.23, notionally and at local scales, respectively. False (positive and negative) reporting probabilities were consistently non‐zero (0.22–0.68). We submit that regular and reliable assessment of stakeholder attitudes––combined with an understanding of factors contributing to variation in attitudes––can feed into participatory conservation monitoring programs, help assess the success of initiatives aimed at facilitating human behavioral change, and inform conservation decision‐making.
Data from: Movement of a Heliconius hybrid zone over 30 years: a Bayesian approach
Hybrid zones have long been of interest to biologists as natural laboratories where we can gain insight into the processes of adaptation and speciation. Repeated sampling of individual hybrid zones has been particularly useful in elucidating the dynamic balance between selection and dispersal that maintains most hybrid zones. Here, we revisit a hybrid zone between Heliconius erato butterflies in Panamá for a third time over more than 30 years. We combine a novel Bayesian extension of stepped‐cline hybrid zone models with environmental data to understand the genetic and environmental causes of cline dynamics in this species. The cline has continued to move west, likely due to dominance drive, but has slowed and broadened. Environmental analyses suggest that widespread deforestation in Panamá could be leading to decreased avian predation and relaxed selection, causing the observed changes in cline dynamics.
Data from: Empirical and Bayesian approaches to fossil-only divergence times: a study across three reptile clades
Estimating divergence times on phylogenies is critical in paleontological and neontological studies. Chronostratigraphically-constrained fossils are the only direct evidence of absolute timing of species divergence. Strict temporal calibration of fossil-only phylogenies provides minimum divergence estimates, and various methods have been proposed to estimate divergences beyond these minimum values. We explore the utility of simultaneous estimation of tree topology and divergence times using BEAST tip-dating on datasets consisting only of fossils by using relaxed morphological clocks and birth-death tree priors that include serial sampling (BDSS) at a constant rate through time. We compare BEAST results to those from the traditional maximum parsimony (MP) and undated Bayesian inference (BI) methods. Three overlapping datasets were used that span 250 million years of archosauromorph evolution leading to crocodylians. The first dataset focuses on early Sauria (31 taxa, 240 chars.), the second on early Archosauria (76 taxa, 400 chars.) and the third on Crocodyliformes (101 taxa, 340 chars.). For each dataset three time-calibrated trees (timetrees) were calculated: a minimum-age timetree with node ages based on earliest occurrences in the fossil record; a 'smoothed' timetree using a range of time added to the root that is then averaged over zero-length internodes; and a tip-dated timetree. Comparisons within datasets show that the smoothed and tip-dated timetrees provide similar estimates. Only near the root node do BEAST estimates fall outside the smoothed timetree range. The BEAST model is not able to overcome limited sampling to correctly estimate divergences considerably older than sampled fossil occurrence dates. Conversely, the smoothed timetrees consistently provide node-ages far older than the strict dates or BEAST estimates for morphologically conservative sister-taxa when they sit on long ghost lineages. In this latter case, the relaxed-clock model appears to be correctly moderating the node-age estimate based on the limited morphological divergence. Topologies are generally similar across analyses, but BEAST trees for crocodyliforms differ when clades are deeply nested but contain very old taxa. It appears that the constant-rate sampling assumption of the BDSS tree prior influences topology inference by disfavoring long, unsampled branches.
Data from: Predicting forest insect flight activity: a Bayesian network approach
Daily flight activity patterns of forest insects are influenced by temporal and meteorological conditions. Temperature and time of day are frequently cited as key drivers of activity; however, complex interactions between multiple contributing factors have also been proposed. Here, we report individual Bayesian network models to assess the probability of flight activity of three exotic insects, Hylurgus ligniperda, Hylastes ater, and Arhopalus ferus in a managed plantation forest context. Models were built from 7,144 individual hours of insect sampling, temperature, wind speed, relative humidity, photon flux density, and temporal data. Discretized meteorological and temporal variables were used to build naïve Bayes tree augmented networks. Calibration results suggested that the H. ater and A. ferus Bayesian network models had the best fit for low Type I and overall errors, and H. ligniperda had the best fit for low Type II errors. Maximum hourly temperature and time since sunrise had the largest influence on H. ligniperda flight activity predictions, whereas time of day and year had the greatest influence on H. ater and A. ferus activity. Type II model errors for the prediction of no flight activity is improved by increasing the model's predictive threshold. Improvements in model performance can be made by further sampling, increasing the sensitivity of the flight intercept traps, and replicating sampling in other regions. Predicting insect flight informs an assessment of the potential phytosanitary risks of wood exports. Quantifying this risk allows mitigation treatments to be targeted to prevent the spread of invasive species via international trade pathways.
Data from: Passerine extrapair mating dynamics: a Bayesian modeling approach comparing four species
In many socially monogamous animals, females engage in extrapair copulation (EPC), causing some broods to contain both within‐pair and extrapair young (EPY). The proportion of all young that are EPY varies across populations and species. Because an EPC that does not result in EPY leaves no forensic trace, this variation in the proportion of EPY reflects both variation in the tendency to engage in EPC and variation in the extrapair fertilization (EPF) process across populations and species. We analyzed data on the distribution of EPY in broods of four passerines (blue tit, great tit, collared flycatcher, and pied flycatcher), with 18,564 genotyped nestlings from 2,346 broods in two to nine populations per species. Our Bayesian modeling approach estimated the underlying probability function of EPC (assumed to be a Poisson function) and conditional binomial EPF probability. We used an information theoretical approach to show that the expected distribution of EPC per female varies across populations but that EPF probabilities vary on the above‐species level (tits vs. flycatchers). Hence, for these four passerines, our model suggests that the probability of an EPC mainly is determined by ecological (population‐specific) conditions, whereas EPF probabilities reflect processes that are fixed above the species level.
Data from: Modelling flight heights of lesser black-backed gulls and great skuas from GPS: a Bayesian approach
Wind energy generation is increasing globally, and associated environmental impacts must be considered. The risk of seabirds colliding with offshore wind turbines is influenced by flight height, and flight height data usually come from observers on boats, making estimates in daylight in fine weather. GPS tracking provides an alternative and generates flight height information in a range of conditions, but the raw data have associated error. Here, we present a novel analytical solution for accommodating GPS error. We use Bayesian state-space models to describe the flight height distributions and the error in altitude measured by GPS for lesser black-backed gulls and great skuas, tracked throughout the breeding season. We also examine how location and light levels influence flight height. Lesser black-backed gulls flew lower by night than by day, indicating that this species would be less likely to encounter turbine blades at night, when birds' ability to detect and avoid them might be reduced. Gulls flew highest over land and lowest near the coast. For great skuas, no significant relationships were found between flight height, time of day and location. We consider four 'collision risk windows', corresponding to the airspace swept by rotor blades for different offshore wind turbine designs. We found the highest proportion of birds at risk for a 22–250 m turbine (up to 9% for great skuas and 34% for lesser black-backed gulls) and the lowest for a 30–258 m turbine. Our results suggest lesser black-backed gulls are at greater risk of collision than great skuas, especially by day. Synthesis and applications. Our novel modelling approach is an effective way of resolving the error associated with GPS tracking data. We demonstrate its use on GPS measurements of altitude, generating important information on how breeding seabirds use their environment. This approach and the associated data also provide information to improve avian collision risk assessments for offshore wind farms. Our modelling approach could be applied to other GPS data sets to help manage the ecological needs of seabirds and other species at a time when the pressures on the marine environment are growing.
Data from: Cladogenetic and anagenetic models of chromosome number evolution: a Bayesian model averaging approach
Chromosome number is a key feature of the higher-order organization of the genome, and changes in chromosome number play a fundamental role in evolution. Dysploid gains and losses in chromosome number, as well as polyploidization events, may drive reproductive isolation and lineage diversification. The recent development of probabilistic models of chromosome number evolution in the groundbreaking work by Mayrose et al. (2010, ChromEvol) have enabled the inference of ancestral chromosome numbers over molecular phylogenies and generated new interest in studying the role of chromosome changes in evolution. However, the ChromEvol approach assumes all changes occur anagenetically (along branches), and does not model events that are specifically cladogenetic. Cladogenetic changes may be expected if chromosome changes result in reproductive isolation. Here we present a new class of models of chromosome number evolution (called ChromoSSE) that incorporate both anagenetic and cladogenetic change. The ChromoSSE models allow us to determine the mode of chromosome number evolution; is chromosome evolution occurring primarily within lineages, primarily at lineage splitting, or in clade-specific combinations of both? Furthermore, we can estimate the location and timing of possible chromosome speciation events over the phylogeny. We implemented ChromoSSE in a Bayesian statistical framework, specifically in the software RevBayes, to accommodate uncertainty in parameter estimates while leveraging the full power of likelihood based methods. We tested ChromoSSE's accuracy with simulations and re-examined chromosomal evolution in Aristolochia, Carex section Spirostachyae, Helianthus, Mimulus sensu lato (s.l.), and Primula section Aleuritia, finding evidence for clade-specific combinations of anagenetic and cladogenetic dysploid and polyploid modes of chromosome evolution.
Data from: Application of a Bayesian weighted surveillance approach for detecting chronic wasting disease in white-tailed deer
SUMMARY 1. Surveillance is critical for the early detection of emerging and re-emerging infectious diseases, and weighted surveillance uses heterogeneity in risk of infection to increase the sampling efficiency. 2. We apply a Bayesian approach to estimate weights for 16 surveillance classes of white-tailed deer in Wisconsin, USA, relative to hunter-harvested yearling males. We use these weights to conduct a surveillance program for detecting chronic wasting disease (CWD) in white-tailed deer at Shenandoah National Park (SHEN) in Virginia, USA. 3. Generally, for surveillance, risk of infection increased with age and was greater in males. Clinical suspect deer had the highest risk with weight estimates of 33.33 and 9.09, for community reported and hunter reported suspect deer, respectively, while fawns had the lowest risk with an estimated weight of 0.001. 4. We used surveillance weights for Wisconsin deer to determine sampling effort required to detect a CWD-positive case in SHEN if prevalence in yearling males ≥0.025. The sampling required to detect CWD was 37–91 adult deer, depending on the adult male:female ratio in the surveillance stream. We collected rectal biopsies from 49 and 21 adult female and male deer, respectively, and 10 additional samples from vehicle-killed deer. CWD was not detected and we concluded with 95% probability that prevalence in the reference population (yearling males) was between 0.0 to 3.6%. 5. Synthesis and applications. Our approach allows managers to estimate relative surveillance weights for different host classes and quantify limits of disease detection in real time when only a sample of animals from a population can be tested, resulting in considerable cost savings for agencies performing wildlife disease detection surveillance. Additionally, it provides a rigorous means of estimating prevalence limits when a disease/pathogen is not detected in a sample set, and is generalizable to other wildlife, domestic animal, and human disease systems which can be characterized by surveillance classes with heterogeneous probability of infection. This methodology is also extendable to other disciplines such as invasive species, environmental toxicology, and generally any ecological question seeking to efficiently use scarce financial and human resources to maximize the detection probability of a rare event.
FIGURE 1. Maximum clade credibility tree after a partitioned Bayesian analysis using 8945 in Phylogenetic analysis of the Neotropical Pristimantis leptolophus species group (Anura: Craugastoridae): molecular approach and description of a new polymorphic species
FIGURE 1. Maximum clade credibility tree after a partitioned Bayesian analysis using 8945 sites depicting the phylogenetic relationships among Pristimantis including the Pristimantis leptolophus species group. Numbers on nodes represent posterior probabilities and ultrafast bootstrap (as obtained in the ML analysis) support respectively. Asterisks represent nodal support larger than 95% in both ML and Bayesian analyses. Two dashes in ultrafast bootstrap indicate the node was not recovered in the ML analysis (see Appendix 2 for the ML tree).
Datasets for "An Evaluation of Bayesian Approaches based on Uniformly Most Powerful Tests for Detecting Violations of Local Independence"
<p>These files contain the results of the simulation studies reported in the main text. A detailed documentation is included as pdf file.</p>
FIGURE. Phylogenetic tree of specimens on Poaceae and related host plants constructed by MP method based on ITS+28S regions of rDNA. Bootstrap values of MP and ML are followed by the Bayesian posterior probabilities (Bpp) on the nodes in the topology. Asterisk (*) represents bootstrap values or Bpp less than 50% in the topology. Sample data are shown with voucher specimen number or GenBank accession number, and host plant. Sequence data determined in this study are shown in color. Teliospore shapes are shown in each clade detected, and new species are shown by asterisk (*) on clades. 0, I: Spermogonial and aecial host genus. Asterisk (*) on host plants: Spermogonial and aecial host plants. in Phylogenetic approach for identification and life cycles of Puccinia (Pucciniaceae) species on Poaceae from northeastern China
FIGURE. Phylogenetic tree of specimens on Poaceae and related host plants constructed by MP method based on ITS+28S regions of rDNA. Bootstrap values of MP and ML are followed by the Bayesian posterior probabilities (Bpp) on the nodes in the topology. Asterisk (*) represents bootstrap values or Bpp less than 50% in the topology. Sample data are shown with voucher specimen number or GenBank accession number, and host plant. Sequence data determined in this study are shown in color. Teliospore shapes are shown in each clade detected, and new species are shown by asterisk (*) on clades. 0, I: Spermogonial and aecial host genus. Asterisk (*) on host plants: Spermogonial and aecial host plants.
FIGURE 2. Overview tree for the COI gene fragment. Bayesian inference tree using MrBayes 3.2.7a in The Oracle of Delphi-a molecular phylogenetic approach to Greek Cordulegaster Leach in Brewster, 1815 (Odonata: Anisoptera: Cordulegastridae)
FIGURE 2. Overview tree for the COI gene fragment. Bayesian inference tree using MrBayes 3.2.7a using the best-fit model (GTR+I+G) identified with JModeltest 2.1.10. Bayesian posterior probabilities values are depicted at the nodes. Included are our own sequences (PCR number next to the name) and those retrieved from GenBank (accession numbers next to the name), if specimens identify different taxa in the COI and ITS analysis they are considered hybrids. Haplotype analysis (TCS-network made in PopART 1.7) is shown in Figs. 5 and 6.
FIGURE 3. Overview tree from the ITS gene fragment. Bayesian inference tree using MrBayes 3.2.7a in The Oracle of Delphi-a molecular phylogenetic approach to Greek Cordulegaster Leach in Brewster, 1815 (Odonata: Anisoptera: Cordulegastridae)
FIGURE 3. Overview tree from the ITS gene fragment. Bayesian inference tree using MrBayes 3.2.7a using the best-fit model (HKY+G) identified with JModeltest 2.1.10. Bayesian posterior probabilities values are depicted at the nodes. Included are our isolated sequences (PCR number next to the name) and those retrieved from GenBank (accession numbers next to the name), if specimens identify different taxa in the COI and ITS analysis they are indicated hybrids.
Data from: Modelling the consequence of glacier retreat on mixotrophic nanoflagellate bacterivory: a Bayesian approach
<p>This repository comprises data and statisticial analysis from the paper Schenone et al. 2020 entitled "Modelling the Consequence of Glacier Retreat on Mixotrophic Nanoflagellate Bacterivory: A Bayesian Approach" which was accepted for publication in Oikos. The main contribution of this paper is a model for mixotrophic nanoflagellate bacterivory as a function of light availability and the presence of non edible particles such as glacial clay. The dataset consists in results from eight bacterivory experiments carried out during November 2018 and January-February 2019 using natural mixotrophic nanoflagellate community of six moutain lakes from north Patagonia, Argentina. A Bayesian approach was adopted to estimate relevant parameters from the model using our experimental data. The analysis was performed using JAGS interfaced through R Studio, the script is available in this repository.</p>
Data from: A Bayesian hierarchical approach to quantifying stakeholder attitudes toward conservation in the presence of reporting error
Open the record for dataset details and reuse information.
Data from: Passerine extrapair mating dynamics: a Bayesian modeling approach comparing four species
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.