Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,066
datasets available to search
ShareScore release 0.7.1
Dataset results
1,066 results for “Bayesian”
Bayesian Analysis of Tree Distributions Across Space and Time in Eastern North America 2010-2011
The distributions of many organisms are spatially autocorrelated, but it is unclear whether including spatial terms in species distribution models (SDMs) improves projections of future species distributions. We provide the first comparative test of a purely spatial SDM, a purely non-spatial SDM, and an SDM that combines spatial and environmental information. Spatial SDMs provided better fits to the calibration data, more accurate predictions of a hold-out validation data set of modern trees, and lower false positive rates at all time periods than non-spatial SDMs. Hindcasted projection of spatial SDMs had higher variance than those of non-spatial SDMs. Overall predictive performance of non-spatial and spatial SDMs varied temporally and as a function of niche overlap. Ecological modelers should include spatial terms in SDMs used for projecting future distributions of species.
Bayesian analysis of the equation of state of quantum chromodynamics from a holographic model
<p>Prior and posterior samples obtained from a Bayesian analysis of the equation of state of quantum chromodynamics (QCD) within a holographic Einstein-Maxwell-Dilaton model, constrained by state-of-the art lattice QCD results at a vanishing net density of baryons.</p> <p>Samples contain metadata, model parameters, and model predictions for the location of the QCD critical point.</p> <p>Supplement to <a title="Bayesian location of the QCD critical point from a holographic perspective" href="https://arxiv.org/abs/2309.00579">arXiv:2309.00579</a>.</p>
Dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area"
<p>This repository archives the dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area". The cognate annotation of Tshangla, Kho-Bwa, Hrusish, Mishmic, and Tani languages were done by us. The cognate decision on the other languages was annotated by Sagart et al. (2019). Please use the following information to cite our work: <br> Wu, M.-S, Bodt, T. A, Tresoldi, T. (2022). Bayesian phylogenetics illuminate shallower relationships Trans-Himalayan languages in the Tibet-Arunachal area. Linguistics of the Tibeto-Burman Area. [forthcoming]</p>
A Bayesian Machine Learning Framework for Animal Telemetry Data
<p>The data and tutorial in this repository are intended to be used in conjunction with the tutorial with our manuscript titled "A Bayesian Machine Learning Framework for Animal Telemetry Data." Telemetry data for three lesser prairie-chickens are provided here as .csv files. For more information about the data, please refer to our manuscript or contact Andrew Whetten or David Haukos for more information.</p>
Atmospheric Halocarbon Observations at Beromünster, Switzerland, and Bayesian Inverse Modeling to assess Emissions
<p>Atmospheric halocarbon (CFCs, halons, HCFCs, HFCs, PFCs, SF<sub>6</sub>, NF<sub>3</sub>, HFOs) and carbon monoxide (CO) observations (mole fractions) from the tall tower site at Beromünster, Switzerland (47.2 °N, 8.2 °E, 797 m a.s.l., 212 m a.g.l.), covering the period September 2019 to September 2020. The halocarbon measurements were conducted using a Medusa pre-concentration unit, coupled to gas chromatography (Agilent 6890N) and mass spectrometry (Agilent 5975, GC-MS).</p> <p>For further details see: Miller, B. R., Weiss, R. F., Salameh, P. K., Tanhua, T., Greally, B. R., Mühle, J., and Simmonds, P. G.: Medusa: A Sample Preconcentration and GC/MS Detector System for in Situ Measurements of Atmospheric Trace Halocarbons, Hydrocarbons, and Sulfur Compounds, Anal. Chem., 80, 1536–1545, https://doi.org/10.1021/ac702084k, 2008).</p> <p>The data format follows that used within the AGAGE network (see AGAGE data archive: <a href="http://agage.mit.edu/data/agage-data">http://agage.mit.edu/data/agage-data</a>).</p> <p>Data results for the Bayesian inversion conducted based on the measurement data from Beromünster to assess Swiss halocarbon emissions. Files are provided in netCDF format for the 28 individual substances discussed in (Rust, D. et al., 2022, <em>Swiss halocarbon emissions for 2019 to 2020 assessed from regional atmospheric observations</em>, Atmospheric Chemistry and Physics). Each file contains the a priori and a posteriori emissions as used or calculated in the Bayesian inversion. Data are provided on the grid used in the inversion (irregular longitude/latitude). Metadata are included as netCDF attributes. The netCDF files follow the CF conventions and are readable with any netcdf interface/tool.</p>
Relativistic description of dense matter equation of state and compatibility with neutron star observables: a Bayesian approach
<p>The general behavior of the nuclear equation of state (EOS), relevant for the description of neutron stars (NS), is studied within a Bayesian approach applied to a set of models based on a density-dependent relativistic mean-field description of nuclear matter <a href="https://arxiv.org/abs/2201.12552">Malik et al 2022</a>. The EOS is subjected to a minimal number of constraints based on nuclear saturation properties and the low-density pure neutron matter EOS obtained from a precise next-to-next-to-next-to-leading order (N$^{3}$LO) calculation in chiral effective field theory ($\chi$EFT). The number of final sample parameters corresponding to the posterior sets is around fourteen thousand. We present five EOSs among them, namely DDBl, DDBm, DDBu1, DDBu2, and DDBx. The DDBl, DDBm, DDBu2 were chosen so that the radius of the 1.4$M_\odot$ star has the lower limit, a medium value, and the upper limit of the 90% CI for the conditional probabilities $P(R|M)$. We have also included DDBu1 that has a slightly lower $R_{1.4}$ than the upper limit but lies completely inside the 90% CI for the conditional probabilities $P(R|M)$. The DDBx is the one that predicts a maximum mass of 2.5$M_\odot$ and has the following nuclear matter properties, $K_0=300$ MeV, $J_{sym,0}=30$ MeV and $L_{sym,0}=39$ MeV.</p> <p>We also release our entire sets of ~14K NS matter EOS. All the EOSs are for NS core and starting baryon density is 0.04 fm$^{-3}$. One needs to add their own choice of crust EOS for the star properties calculation. The uncertainty in star properties for the choice of the different crust has been discussed in Section 2.1 of the manuscript (arxiv: 2201.12552). </p> <pre> To extract the entire sets of ~14K NS matter EOS files, one needs to follow the steps, 1) unzip DDB_EOS_14K.zip ----------------------------Note------------------------------------- All the eos files have three columns baryon density (fm-3), energy density (MeV.fm-3), and pressure (MeV.fm-3). The starting density is 0.04 fm-3, as it is NS core eos. One needs to add their own choice of crust eos in order to calculate NS properties. ---------------------------------------------------------------</pre> <p> </p> <p> </p>
Bayesian Methods for Ancestral State Reconstruction in Morphosyntax
<p>Supplementary files to accompany journal submission.</p> <p>Files are:</p> <p> </p> <p>tree.pdf - pdf consensus tree, for illustration</p> <p>data.txt - coding file</p> <p>TREE_Set.t - nexus format sample of trees.</p> <p>sources.pdf - source materials used for languages</p>
Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO2 Capture Technologies
<p>Dataset of process simulations results of the natural gas sweetening and flue gas treatment (first and second sheet, respectively as indicated by the sheet name in the .xlsx file). The dataset refers to the publication <em>Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO<sub>2</sub> Capture Technologies </em>by V. Negri, Vàzquey D., Sales-Pardo, Marta, Guimerà, R. and Guillén-Gosàlbez, G. The training and testing dataset are used to generate the figures in the main manuscript and supplementary information. </p> <p> </p>
Supplementary data: The added value of Bayesian inference for estimating biotransformation rates of organic contaminants in aquatic invertebrates.
<p>Supporting information for the article "<strong>The added value of Bayesian inference for estimating biotransformation rates of organic contaminants in aquatic invertebrates.</strong>"</p> <p>This provides all the R script and .csv files for each dataset. </p>
A Bayesian Approach to Detect Pedestrian Destination-Sequences from WiFi Signatures: Data (Transp. Res. Part C, 2014)
<p>This dataset contains and describes the data used in</p> <p>Danalet, A., Farooq, B., & Bierlaire, M. (2014). A Bayesian approach to detect pedestrian destination-sequences from WiFi signatures. <em>Transportation Research Part C: Emerging Technologies</em>, <strong>44</strong>, 146-170. doi:10.1016/j.trc.2014.03.015</p> <p>Specifically it contains WiFi traces, pedestrian Semantically-Enriched Routing Graph (SERG), and Potential Attractivity measure (PAM).</p>
Phlorest phylogeny derived from Walker & Ribeiro 2011 'Bayesian phylogeography of the Arawak expansion in lowland South America'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Walker, R. S., & Ribeiro, L. A. (2011). Bayesian phylogeography of the Arawak expansion in lowland South America. Proceedings of the Royal Society B: Biological Sciences, 278(1718), 2562–2567.</p> </blockquote>
Phlorest phylogeny derived from Kolipakam et al. 2018 'A Bayesian phylogenetic study of the Dravidian language family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Kolipakam V, Jordan FM, Dunn M, Greenhill SJ, Bouckaert R, Gray RD & Verkerk A. 2018. A Bayesian phylogenetic study of the Dravidian language family. R. Soc. Open Sci. 5: 171504.</p> </blockquote>
Phlorest phylogeny derived from Kitchen et al. 2009 'Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Kitchen A, Ehret C, Assefa S & Mulligan CJ. 2009. Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Proceedings of the Royal Society B: Biological Sciences, 270(1668), 2703-2710.</p> </blockquote>
Phlorest phylogeny derived from Lee & Hasegawa 2011 'Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee S, Hasegawa T (2011) Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proceedings of the Royal Society B: Biological Sciences, 278(1725):3662–9.</p> </blockquote>
Data for: Bayesian Analysis for Remote Biosignature Identification on exoEarths (BARBIE) 2: Using Grid-Based Nested Sampling in Coronagraphy Observation Simulations for O2 and O3
<p>We present all of the data across our SNR and abundance study for the molecules O2 and O3 for an exoEarth twin. The wavelength range is from 0.515-1 micron, with 25 evenly spaced 20% bandpasses in this range. The SNR ranges from 3-20, and the abundance values range in log space in steps of 0.5 and/or 0.25 (all presented in VMR in the associated table). We present the lower and upper wavelength per bandpass, the input O2 and O3 values (abundance case), the retrieved O2 and O3 values (presented as the log10(VMR)), the lower and upper limits of the 68% credible region (presented as the log10(VMR)), and the log-Bayes factor for O2 and O3. For more information about how these were calculated, please see Bayesian Analysis for Remote Biosignature Identification on exoEarths (BARBIE) 2: Using Grid-Based Nested Sampling in Coronagraphy Observation Simulations for O2 and O3, accepted and currently available on arXiv. </p> <p>To open this csv as a Pandas dataframe, use the following command:</p> <p>your_dataframe_name = pd.read_csv(f'zenodo_table.csv', dtype={'Input O2': str, {'Input O3': str}})</p>
Global Surface Ozone Concentration Dataset 1990-2017 Mapped at Fine Resolution through the Bayesian Maximum Entropy Data Fusion of Observations and Model Output
<p>This global surface ozone concentration dataset corresponds to the data developed in this paper:</p> <p>DeLang, M. N., J. S. Becker, K.-L. Chang, M. L. Serre, O. R. Cooper, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, S. Cleland, E. Collins, M. Brauer, and J. J. West (2021) Mapping yearly fine resolution global surface ozone through the Bayesian Maximum Entropy data fusion of observations and model output for 1990-2017, <em>Environmental Science & Technology</em>, 55, 4389-4398, doi: 10.1021/acs.est.0c07742.</p> <p>Ozone concentrations are estimated as described in the paper, with output shown for the Ozone Season Daily Maximum 8-hr metric (OSDMA8) for each year between 1990 and 2017, at 0.1 degree spatial resolution. Ozone is estimated through data fusion of output from several global models, with observations of ozone collected by TOAR. The data fusion involves application of the M3Fusion method to create a multi-model composite of several global models, followed by BME data fusion, as described in the paper. </p> <p>The *.nc file contains the latitude, longitude, ozone concentration estimate, and estimated variance for each 0.1 x 0.1 degree grid cell.</p> <p>Please contact Jason West (jasonwest@unc.edu) with questions about the dataset. We'd like to hear from you to know how you're using the data!</p> <p> </p> <p> </p>
Identifying the interplay between protective measures and settings on the SARS-CoV-2 transmission using a Bayesian network [Dataset]
<p>data07B.csv: dataset for the study of the SARS-CoV-2 transmission.</p> <p>CPTNetica.txt: conditional probabilities tables of each variable given through Netica once the BN obtained in R code is loaded.</p> <p>code01.R: code to learn structure and parameters of the SARS-CoV-2 BN model.</p>
GrainLearning: A Bayesian uncertainty quantification toolbox for discrete and continuum numerical models of granular materials
GrainLearning is a Bayesian uncertainty quantification and propagation toolbox for computer simulations of granular materials. The software is primarily used to infer and quantify parameter uncertainties in computational models of granular materials from observation data, also known as inverse analyses or data assimilation. Implemented in Python, GrainLearning can be loaded into a Python environment to process the simulation and observation data, or alternatively, as an independent tool where simulation runs are done separately, e.g., via a shell script.
Supplement to "Proof of concept for Bayesian inference of dynamic rating curve uncertainty" (v3)
<div>This deposit contains part of the updated supplement to “Proof of concept for Bayesian inference of dynamic rating curve uncertainty” (<a href="https://www.tandfonline.com/doi/full/10.1080/02626667.2024.2401094" target="_blank" rel="noopener">Cornelio et al. 2024, HSJ</a>). This version, in particular, contains two files in which the following changes were made from the earlier version (v2.0.1):</div> <div> <ul> <li><strong><em>250117_Lbn_RC_new.R</em></strong> is the updated R code. The argument for the random number generator (RNG) kind is defined for the set.seed() functions used in the script. </li> <li><strong><em>Lbn-DMs-csv0.csv</em></strong> is the updated input file containing the stage-discharge gaugings. The column for the stage values has been renamed to "H_rec" (instead of "H_m" as in the original CSV) to be consistent with the attribute name used throughout the R code.</li> </ul> <p>Except for the above files, all the input and output files in <a href="https://zenodo.org/records/12792513" target="_blank" rel="noopener">v2.0.1</a><span> remain unchanged. </span></p> </div> <p><u> </u></p>
Global Surface Ozone Concentration Dataset 1990-2017 Generated by Bayesian Maximum Entropy Data Fusion With RAMP Bias Correction
<p>This dataset reports estimates of surface ozone concentration at fine spatial resolution for 1990 to 2017, at 0.5 degree horizontal resolution. Also reported is the variance. Estimates correspond to this paper:</p> <p><span>Becker, J. S.</span><span>, DeLang, M. N., K.-L. Chang, M. L. Serre, O. R. Cooper, <u>H. Wang</u>, M. G. Schultz, S. Schroder, X. Lu, L. Zhang, M. Deushi, B. Josse, C. A. Keller, J.-F. Lamarque, M. Lin, J. Liu, V. Marecal, S. A. Strode, K. Sudo, S. Tilmes, L. Zhang, M. Brauer, and <span>J. J. West</span> (2023) Using Regionalized Air Quality Model Performance and Bayesian Maximum Entropy data fusion to map global surface ozone concentration, <em>Elementa Science of the Anthropocene</em>, 11: 1, doi: 10.1525/elementa.2022.00025.</span></p> <p>The dataset reports estimates of surface ozone for the OSDMA8 metric (the 6-month ozone-season average of the daily maximum 8-hr concentration), estimated through a data fusion of ozone observations from the Tropospheric Ozone Assessment Report (TOAR) database, and output from multiple global atmospheric models. Estimates are created in each year by a combination of M3Fusion to create a multi-model composite, Regional Air Quality Model Performance (RAMP) regional and nonlinear bias correction, and Bayesian Maximum Entropy (BME) data fusion in space and time. The estimates here are the final results using a weighted RAMP bias correction. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.