Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
89
datasets available to search
ShareScore release 0.9.0
Dataset results
89 results for “Bayesian estimation”
Supplementary data: The added value of Bayesian inference for estimating biotransformation rates of organic contaminants in aquatic invertebrates.
<p>Supporting information for the article "<strong>The added value of Bayesian inference for estimating biotransformation rates of organic contaminants in aquatic invertebrates.</strong>"</p> <p>This provides all the R script and .csv files for each dataset. </p>
SEDflow: Accelerated Bayesian SED Modeling using Amortized Neural Posterior Estimation
<p><a href="http://changhoonhahn.github.io/SEDflow">SEDflow</a> is an accelerated Bayesian SED modeling method that uses the <a href="https://ui.adsabs.harvard.edu/abs/2022arXiv220201809H/abstract">Hahn et al. (2022a)</a> PROVABGS SED model and Amortized Neural Posterior Estimation (ANPE) to derive posterior probability distributions of galaxy properties from optical photometry. SEDflow is<span class="math-tex">\(10^5\times\)</span> faster than conventional Markov Chain Monte Carlo sampling methods and takes ~1 second per galaxy to obtain posteriors. This repository includes all of the data used to train, validate, and test SEDflow.</p> <p>This repository also includes a value-added catalog with detailed physical properties of 33,884 galaxies in the NASA-Sloan Atlas (http://www.nsatlas.org/). The properties are inferred from optical photometry in the <em>u, g, r, i, z</em> bands using SEDflow. For more details on this catalog and SEDflow see the <a href="http://changhoonhahn.github.io/SEDflow">documentation</a> and Hahn & Melchior (2022). </p> <p>For each galaxy, the catalog provides posteriors of: </p> <ul> <li>log_mstar: log10 of stellar mass</li> <li>log_sfr_1gyr: log10 of average star formation rate over 1Gyr</li> <li>log_z_mw: log10 of mass-weighted metallicity</li> <li>beta1, beta2, beta3, beta4: coefficients of the non-negative matrix factorization (NMF) star formation history basis functions</li> <li>fburst: fraction of stellar mass formed by a starburst event</li> <li>tburst: time of the starburst event</li> <li>log_gamma1, log_gamma2: log10 of coefficients of the NMF metallicity history basis functions</li> <li>tau_bc: birth cloud optical depth</li> <li>tau_ism: diffuse dust optical depth</li> <li>n_dust: Calzetti (2001) dust index</li> </ul>
Minimal dataset for "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models"
<p>This repository contains a minimal data set to reproduce all results that don't compromise the privacy concerns for the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> <br> The repository contains the following data:</p> <ul> <li>adaptscore_acute.csv <ul> <li>A csv file that contains the estimated adaptation scores for the acute data set with HLA I model.</li> </ul> </li> <li>adaptscore_leftout.csv <ul> <li>A csv file that contains the estimated adaptation scores for the leftout data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training.csv <ul> <li>A csv file that contains the estimated adaptation scores for the traininig data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training_hla1_without_clin.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the HLA I model (via cross-validation)</li> </ul> </li> <li>adaptscore_training_seed2.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the joint HLA I and HLA II model via cross-validation with another seed</li> </ul> </li> </ul>
Photometric Redshifts for Cosmology: Improving accuracy and uncertainty estimates using Bayesian Neural Networks
<p><strong>This data consists of 286,401 with broad-band g,r,i,z,y photometry from the HSC DR2 survey and spectroscopic redshifts. The majority of galaxies in our sample lies between redshift of 0.01 and 2.5</strong></p>
Flux estimates for the Fermi map (Bayesian GCNN vs. NPTFit)
<p>Prediction of the Bayesian graph-convolutional neural network (GCNN) for the Fermi photon-count map. The shaded regions show the predictive (aleatoric and epistemic summed in quadrature) 1σ uncertainty. The markers with error bars (68% credible intervals) indicate the Non-Poissonian template fit (NPTFit) estimates for comparison. The NN predictions for the GCE flux in the Fermi map are similar in magnitude to those of NPTFit, but the GCE is almost entirely attributed to the smooth dark matter template.</p> <p>To view the plot, open it with a browser such as Google Chrome or Firefox. Templates can be (de-)activated by clicking on the colored rectangles. For zooming and panning, click on the buttons in the lower left corner.</p>
Biomass production at 2085 horizon for the Maurienne valley (French Alps) estimated using a Bayesian Belief Network
<p>In mountains, grasslands managed for livestock production sustain local economies, culture and identity. However, their future fodder production is highly uncertain under climate change: while an extended growing season may be beneficial, more frequent and intense summer droughts could also reduce fodder quantity and quality. Land use and land cover (LULC) changes are another major driver of regional grassland biomass production, but combined effects of future land use transitions and climate change are rarely quantified.</p> <p>We modelled combined climate and LULC scenarios for grassland production of the Maurienne Valley (French Alps) by 2100. We built a Bayesian Belief Network (BBN) from long-term grassland production monitoring data complemented with expert knowledge. We assessed the potential of two candidate adaptations, intensification as an incremental solution, and silvopastoralism as a transformative solution to compensate combined impacts of two climate scenarios and three land use change scenarios.</p> <p>Total biomass production was far more sensitive to LULC than to climate scenarios. Production losses were largest under the Conservation LULC scenario (-28% on average between 2020 and 2085), followed by the Tourism development scenario (-7%) and the Business-as-Usual scenario (+3%). Climate change under RCP 8.5 altered the seasonality of production by increasing potential production from May to July while decreasing summer regrowth. Intensification somewhat compensated effects of climate and LULC changes on biomass production, whereas silvopastoralism offered only marginal gains. The Bayesian network model explicitly captured a future increase in interannual variability in biomass production.</p> <p>Synthesis and application: Changes in LULC are more decisive for global biomass production than climate change. However, under the most extreme climate change scenario (RCP8.5), the seasonal shift in production and increased interannnual variability threaten the current grass-based Protected Designation of Origin production system. Only the intensification adaptation solution showed significant gains in total biomass production. Still, the silvopastoralism would require less investment compared to the intensification and have a similar efficiency when assessing the gains of biomass by the surface concerned with adaptation solutions. </p>
Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 in Complete mitochondrial genomes from museum specimens clarify millipede evolution in the Eastern Arc Mountains
Figure 5. Bayesian phylogeny, with species divergence age estimates reconstructed with BEAST using all the 26 mitochondrial genomes generated in this study. The dataset was supplemented with Thyropygus sp. and Abacion magnum as outgroups, derived from GenBank. GenBank accession numbers are provided in parentheses. Blue bars indicate the 95% highest probability density intervals for node ages. Age estimation for lineage divergence was based on a general arthropod mitochondrial DNA substitution rate and should be considered with caution. *Thyropygus sp. (red font) is very likely to be a misidentification; for more information, see the Discussion.
Figure 4. Phylogeny constructed through Bayesian inference estimated from the 35H in Gene Flow Patterns of the Aedes aegypti (Diptera: Culicidae) Mosquito in Colombia: a Continental Comparison Suggests Multiple Invasion Routes and Gene Exchange
Figure 4. Phylogeny constructed through Bayesian inference estimated from the 35H found of the ND4 gene for the A. aegypti populations in the American continent. The blue horizontal bars above the branches reflect the 95% CI for the branch supports. The color bars (blue, green, and red) on the tree terminals indicate which haplotypes are exclusive for a specific population. The dotted lines on the right side of the tree and numbers I or II indicate to what clade each of the terminals belong. H1-Col (Colombia (Sucre and Quindio), Venezuela, Peru, M-NA, Brasil (MA-O, RBPV, SEBr, BE-L)), H4 (Venezuela, M-NA, Brasil (MA-O, RBPV, BEL)), H3 (Venezuela, M-NA, Brazil (RBPV, SEBr, BE-L)), H2-Col (Colombia (Sucre), Venezuela, Peru, M-NA, Brazil (RBPV, SEBr, BE-L)), H13 (M-NA, Brazil (SEBr)), H8 (Venezuela, Brazil (SEBr)).
A Bayesian Estimation of the Milky Way's Circular Velocity Curve using Gaia DR3
<p>The derived input dataset of approximately 0.6 million RGB stars used in the work "A Bayesian Estimation of the Milky Way’s Circular Velocity Curve using Gaia DR3" . The authors kindly ask to cite the original work described in (<a href="https://doi.org/10.1051/0004-6361/202346474">https://doi.org/10.1051/0004-6361/202346474</a>) should one make use of this catalogue.</p> <p> </p> <p> </p>
Data from: Bayesian estimation of muscle mechanisms and therapeutic targets using variational autoencoders
Open the record for dataset details and reuse information.
Biomass production at 2085 horizon for the Maurienne valley (French Alps) estimated using a Bayesian Belief Network
Open the record for dataset details and reuse information.
All simulation results, figures and code regarding the manuscript: Calibrating models of cancer invasion: parameter estimation using Approximate Bayesian Computation and gradient matching
<p>We present two different methods to estimate parameters within a partial differential equation (PDE) model of cancer invasion. The model describes the spatio-temporal evolution of three variables -- tumour cell density, extracellular matrix density and matrix degrading enzyme concentration -- in a one-dimensional tissue domain. The first method is a likelihood-free approach associated with Approximate Bayesian Computation (ABC); the second is a two-stage gradient matching method based on smoothing the data with a Generalized Additive Model (GAM) and matching gradients from the GAM to those from the model. Both methods performed well on simulated data. To increase realism, additionally we tested the gradient matching scheme with simulated measurement error and found that the ability to estimate some model parameters deteriorated rapidly as measurement error increased.</p>
Data from: Bayesian estimation of the global biogeographical history of the Solanaceae
Aim: The tomato family Solanaceae is distributed on all major continents except Antarctica and has its centre of diversity in South America. Its worldwide distribution suggests multiple long-distance dispersals within and between the New and Old Worlds. Here, we apply maximum likelihood (ML) methods and newly developed biogeographical stochastic mapping (BSM) to infer the ancestral range of the family and to estimate the frequency of dispersal and vicariance events resulting in its present-day distribution. Location: Worldwide. Methods: Building on a recently inferred megaphylogeny of Solanaceae, we conducted ML model fitting of a range of biogeographical models with the program 'BioGeoBEARS'. We used the parameters from the best fitting model to estimate ancestral range probabilities and conduct stochastic mapping, from which we estimated the number and type of biogeographical events. Results: Our best model supported South America as the ancestral area for the Solanaceae and its major clades. The BSM analyses showed that dispersal events, particularly range expansions, are the principal mode by which members of the family have spread beyond South America. Main conclusions: For Solanaceae, South America is not only the family's current centre of diversity but also its ancestral range, and dispersal was the principal driver of range evolution. The most common dispersal patterns involved range expansions from South America into North and Central America, while dispersal in the reverse direction was less common. This directionality may be due to the early build-up of species richness in South America, resulting in large pool of potential migrants. These results demonstrate the utility of BSM not only for estimating ancestral ranges but also in inferring the frequency, direction and timing of biogeographical events in a statistically rigorous framework.
Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models
<p>This is the consensus sequence repository to the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>
Estimates of tidal-marsh bird densities using Bayesian networks
<p>This data set was generated by Wiest, WA, MD Correll, BG Marcot, BJ Olsen, CS Elphick, TP Hodgman, GR Gunterspergen, and WG Shriver. 2019. Estimates of tidal marsh bird densities using Bayesian networks. Journal of Wildlife Management 83:109-120 (paper available here <a href="https://doi.org/10.1002/jwmg.21567">https://doi.org/10.1002/jwmg.21567</a>) and was developed by the Saltmarsh Habitat and Avian Research Program (https://www.tidalmarshbirds.org/). The same data set has subsequently been used in other papers produced by our group.</p>
Bayesian Estimation on Microbial Electrochemical Time Profiles obtained by a High-Throughput Bioelectrochemical Device
<p>Bioelectrochemical systems (BESs) are attracting much attention, but the mechanisms are not yet fully clarified. One of the issues for the BESs research is the lack of a high-quality database that is developed through well-controlled experiments with a high-throughput data-collecting system. Here, we developed a new high-throughput potentiostat with 96 well plates with silk-screen-printed electrodes. With this potentiostat, we obtained 576 time-profiles of microbial current production for viewing the complex landscape of the parameters, namely, the redox mediator concentration and the electrode potential. This dataset contains the following materials:</p> <ul> <li>Current production profiles for <em>Shewanella </em>(Data with average of four data sets for impact of mediators.zip)</li> <li>Summary of the calculated values (statistics summary.csv)</li> <li>Python code for calculating the maximum slope (Slope EF.py)</li> <li>Python codes for Bayesian estimation (2D Bayesian Optimization mapping.py and 2D Bayesian Optimization_slice.py)</li> </ul>
Improved estimation of the prevalence of bovine cysticercosis and the diagnostic test characteristics in the absence of a reference standard using Bayesian Latent Class models, the example of Jimma and Ambo Abattoirs, Ethiopia
<p>Bovine cysticercosis is an infection of cattle musculature with the cestode parasite of humans known as Taenia saginata. This bovine cysticercosis data was collected from two Ambattoirs in Ethiopia namely Ambo and Jimma. Dissection of the predilection site, Ag-ELISA, and meat inspection were the diagnostic methods employed. Cysticerci collected during dissection of the predilection site were also confirmed using multiplex PCR. </p>
Estimated parameters for Bayesian Multilevel Models of KM and kcat values
<p>RData (.rds) files containing brmsfit model objects estimated with the brms R package from KM and kcat values reported in BRENDA and SABIO-RK.</p> <p>These models are used by the ENKIE python package to predict kinetic parameter values and uncertainties.</p>
Scalable Bayesian divergence time estimation with ratio transformations
<div class="page"> <div class="layoutArea"> <div class="column"> <p><span>Divergence time estimation is crucial to provide temporal signals for dating bio</span><span>logically important events, from species divergence to viral transmissions in space and </span><span>time. With the advent of high-throughput sequencing, recent Bayesian phylogenetic </span><span>studies have analyzed hundreds to thousands of sequences. Such large-scale analyses</span><span> </span><span>challenge divergence time reconstruction by requiring inference on highly-correlated</span><span> </span><span>internal node heights that often become computationally infeasible. To overcome this</span><span> </span><span>limitation, we explore a ratio transformation that maps the original </span><span>N - </span><span>1 internal</span><span> </span><span>node heights into a space of one height parameter and </span><span>N - </span><span>2 ratio parameters. To</span><span> </span><span>make the analyses scalable, we develop a collection of linear-time algorithms to com</span><span>pute the gradient and Jacobian-associated terms of the log-likelihood with respect to </span><span>these ratios. We then apply Hamiltonian Monte Carlo sampling with the ratio trans</span><span>form in a Bayesian framework to learn the divergence times in four pathogenic viruses</span><span> </span><span>(West Nile virus, rabies virus, Lassa virus and Ebola virus) and the coralline red algae.</span><span> </span><span>Our method both resolves a mixing issue in the West Nile virus example and improves</span><span> </span><span>inference efficiency by at least 5-fold for the Lassa and rabies virus examples as well</span><span> </span><span>as for the algae example. Our method now also makes it computationally feasible to</span><span> </span><span>incorporate mixed-effects molecular clock models for the Ebola virus example, confirms</span><span> </span><span>the findings from the original study and reveals clearer multimodal distributions of the</span><span> </span><span>divergence times of some clades of interest.</span></p> </div> </div> </div>
Integrating multiple field measurements in a Bayesian parallel regression framework to estimate Tasmanian devil age
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.