Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “mcmc”
Updated NICER PSR J0740+6620 Illinois-Maryland MCMC Samples
<p>Posterior Samples from "A More Precise Measurement of the Radius of PSR J0740+6620 Using Updated NICER Data." <br><br>Our primary mass and radius samples can be found in "J0740_NICERXMM_full_mr.txt" </p> <p>Files ending in "all.txt" include values for all model parameters, while files ending in "mr.txt" contain only the mass and radius values for each sample. <br>Files named "NICERonly" include samples from analyses of only NICER data, while files named "NICERXMM" hold samples from joint analyses of NICER and XMM-Newton data. <br>"full", "partial" and "helium" indicate that those samples are from analyses using fully ionized hydrogen atmosphere models, atmosphere models allowing partially ionized hydrogen, and ionized helium atmospheres respectively. <br><br></p>
Emissions-based MCMC chains for Hector emissions scenario paper
<p>These csvs contain MCMC chains and sampled subsets for emissions-based calibration of the Hector simple climate model (<a href="https://github.com/JGCRI/hector">https://github.com/JGCRI/hector</a>, DOI:10.5194/gmd-8-939-2015).</p> <p>The calibrations use a version of Hector that includes the BRICK sea-level module (<a href="https://github.com/scrim-network/BRICK">https://github.com/scrim-network/BRICK</a>, DOI:10.5194/gmd-10-2741-2017). Hector with BRICK is available on my fork of the Hector model (https://github.com/bvegawe/hector/tree/dev_slr). The calibration process is also adapted from BRICK. The code used to produce these chains can be found at https://github.com/bvegawe/hector_probabilistic, DOI:10.5281/zenodo.3236411.</p> <p>These four sets of MCMC chains were produced using hector_calib_driver.R. Inputs used to create each calibration are specified below: </p> <p>emissions_05.csv: Rscript hector_calib_driver.wideDiff.R -f *output folder* -n 1000000 --endyear 2005 --np 10</p> <p>emissions_09.csv: Rscript hector_calib_driver.wideDiff.R -f *output folder* -n 1000000 --endyear 2009 --np 10</p> <p>emissions_ohc_05.csv: Rscript hector_calib_driver.wideDiff.R -f *output folder* -n 1000000 --endyear 2005 --np 10 --obs_set noTE_obs --model_set noTE_model</p> <p>emissions_ohc_09.csv: Rscript hector_calib_driver.wideDiff.R -f *output folder* -n 1000000 --endyear 2009 --np 10 --obs_set noTE_obs --model_set noTE_model</p> <p> </p>
MCMC Traces for GitHub repo "JuBiotech/petase-paper"
<p>Additional MCMC data to re-run analyses from https://github.com/JuBiotech/petase-ts-paper. To use the existing notebooks, please clone the GitHub repository and download and unzip this folder. Afterwards, merge this dataset and its folder structure with the existing folder "data_analysis" of the repo.</p>
Author Name Disambiguation using Markov Chain Monte Carlo (MCMC)
<p>This repository contains the files processed from the <a href="https://zenodo.org/record/5675801#.Y2APcOzMJhE">Aminer-534K</a> Knowledge graph for the master's thesis on Author Name Disambiguation.</p>
TOI-332: chains from the MCMC fit
<p>The 5 chains from the MCMC analysis of the TOI-332 system (paper in review). Chain files are 50,000 steps (rows), with no burn-in removed (1,000 steps are removed in our analysis).</p> <p>Parameters are given in the column headers, and described below:</p> <ul> <li>The lightcurve means (mean_tess28, mean_tess12, mean_lco1, mean_lco2, mean_lco3, mean_lco4, mean_lco5, mean_lco6)</li> <li>Intervals for the uniform priors (u_star_quadlimbdark____0, u_star_quadlimbdark____1, m_star_interval__ , r_star_interval__, logr_interval__, b_impact__, period_interval__, t0_interval__, logK_interval__, offset_harps_interval__)</li> <li>And posterior values: <ul> <li>Log of the jitter on the HARPS data (where jitter is in m/s): log_jitter_harps </li> <li>The limb darkening: u_star__0, u_star__1 </li> <li>Stellar mass (Msun): m_star</li> <li>Stellar radius (Rsun): r_star</li> <li>Log of the planet radius (where radius is in Rsun): logr </li> <li>Radius of the planet (Rsun): r_pl</li> <li>Planet-to-star radius ratio: ror</li> <li>Impact parameter: b</li> <li>Planet period (days): period</li> <li>Reference time of midtransit (BJD-2457000): t0</li> <li>Log of the RV semi-amplitude (where semi-amplitude is in m/s): logK</li> <li>Offset on the HARPS data (m/s): offset_harps</li> </ul> </li> </ul>
MCMC chains representing stellar halo fits for Lane et al. (2023) paper published in MNRAS
<p>Results for the paper entitled "The stellar mass of the Gaia-Sausage/Enceladus accretion remnant" published in MNRAS by Lane, Bovy, & Mackereth. The results are in the form of MCMC posterior chains, which are used to generate the results recorded in Table 1 of the manuscript. The process used to generate these data is outlined in section 2 (data & sample preparation), section 3 (Density modelling framework), and section 4 (Presentation of results) of the manuscript.</p> <p>The data are contained within a single gzipped .tar file, organized in two directories: gse/ containing the results for fits to the GS/E kinematic samples, and all/ containing the results for the fits to the whole halo sample. In the gse/ directory, there are three directories: eLz/, AD/, and JRLz/, containing the results for the fits to each of the respective kinematically defined subsamples. Within these directories, as well as within all/, are a final set of directories for each of the eight density profiles considered in the work.</p> <p>Within each of these directories are a single file named "samples.npy" which is a binary file (see <a href="https://numpy.org/doc/stable/reference/generated/numpy.lib.format.html#module-numpy.lib.format">this link</a> for format information) containing a numpy array of shape (# of parameters, # of posterior samples). The number of posterior samples is the number of walkers (100) times the number of samples per walker (10,000), minus the number of burn-in steps per walker (1,000) = 900,000. The number of parameters can be inferred from the density profiles definitions in section 3 of the manuscript, and the ordering of the parameters is the same as in Table 1. The units of the parameters are as in Table 1, except for the parameters theta and phi which are scaled such that they are defined on the interval (0,1), representing the intervals (0,2pi) and (0,pi) respectively.</p>
Rare and widespread: Integrating Bayesian MCMC approaches, Sanger sequencing and Hyb-Seq phylogenomics to reconstruct the origin of the enigmatic Rand Flora genus Camptoloma
Open the record for dataset details and reuse information.
MCMC data for A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection
<p>These are unprocessed Markov-chain Monte-Carlo datasets accompanying the manuscript "A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection"</p>
MCMC Samples for Gullikson et. al. 2016
<p>This includes a tarfile with MCMC samples and everything needed to run the 'RealData.ipynb' notebook in the github repository accompanying my paper: </p> <p>https://github.com/kgullikson88/BinaryInference</p>
FIGURE. Bayesian MCMC phylogram of Gliophorus spp. based on ITS sequences. Rooted to Hygrophorus pudorinus. Bar = estimated changes/nucleotide. Support values above or below branches: Bayesian posterior probability/maximum likelihood bootstrap. in New and interesting species of Agaricomycetes from Panama
FIGURE. Bayesian MCMC phylogram of Gliophorus spp. based on ITS sequences. Rooted to Hygrophorus pudorinus. Bar = estimated changes/nucleotide. Support values above or below branches: Bayesian posterior probability/maximum likelihood bootstrap.
Model fits (MCMC samples) for two papers on perceptual confidence
<p>These are model fits for two papers on perceptual confidence. See <a href="https://github.com/wtadler/confidence/">github.com/wtadler/confidence</a> for more info.</p>
2HDM Type-II MCMC scan
<p>MCMC scan of the 2HDM Type-II parameter space, considering Higgs signal strength measurements (HiggsSignals) and the Electroweak Precision observables S and T.</p> <p>HDF5 file compressed with BLOSC compression library.</p> <p> </p> <p>To load the dataset as a pandas dataframe:</p> <pre><code class="language-python">import pandas as pd df = pd.read_hdf(path_to_file)</code></pre> <p> </p> <p>Data columns:</p> <ul> <li>cba: cos(beta-alpha)</li> <li>tb: tan(beta)</li> <li>mH: mass of the heavier higgs</li> <li>mA: mass of the pseudoscalar Higgs</li> <li>mHc: mass of the charged higgs</li> <li>Z7: quartic coupling</li> <li>li: i=1,2,...,7 potential shape parameters</li> <li>...</li> </ul>
MCMC Chains of Age Estimation for Rose et al. 2019
<p>This is the full MCMC chains from the analysis of Rose, Garnavich, Berg 2019, ApJ 874, 32 (https://doi.org/10.3847/1538-4357/ab0704). The analysis code is available on GitHub (https://github.com/benjaminrose/MC-Age) and archived at https://doi.org/10.5281/zenodo.3875494.</p>
MCMC Chains and Maximum Likelihood Parameters for a Random Walk Model of Dark Matter Halo Spins
<p>MCMC chains and the maximum likelihood model parameter file associated with the random walk dark matter halo spin model of Benson, Behrens, & Lu (2020; https://arxiv.org/abs/2001.09208). See the README file for details.</p>
MCMC Chains and Maximum Likelihood Parameters for a Random Walk Model of Dark Matter Halo Concentrations
<p>MCMC chains and the maximum likelihood model parameter file associated with the random walk dark matter halo concentration model of Johnson, Benson, & Grin (2020; https://arxiv.org/abs/2006.15231). See the README file for details.</p>
MCMC files for Inferring differential subcellular localisation in comparative spatial proteomics using BANDLE
<p>MCMC data to accompany paper </p>
Data from: Quantifying MCMC exploration of phylogenetic tree space
In order to gain an understanding of the effectiveness of phylogenetic Markov chain Monte Carlo (MCMC), it is important to understand how quickly the empirical distribution of the MCMC converges to the posterior distribution. In this paper we investigate this problem on phylogenetic tree topologies with a metric that is especially well suited to the task: the subtree prune-and-regraft (SPR) metric. This metric directly corresponds to the minimum number of MCMC rearrangements required to move between trees in common phylogenetic MCMC implementations. We develop a novel graph-based approach to analyze tree posteriors and find that the SPR metric is much more informative than simpler metrics that are unrelated to MCMC moves. In doing so we show conclusively that topological peaks do occur in Bayesian phylogenetic posteriors from real data sets as sampled with standard MCMC approaches, investigate the efficiency of Metropolis-coupled MCMC (MCMCMC) in traversing the valleys between peaks, and show that conditional clade distribution (CCD) can have systematic problems when there are multiple peaks.
MCMC output files for: Quantitative characterization of population-wide tissue- and metabolite-specific variability in perchloroethylene toxicokinetics in male mice
<p>Quantification of inter-individual variability is a continuing challenge in risk assessment, particularly for compounds with complex metabolism and multi-organ toxicity. Toxicokinetic variability for perchloroethylene (perc) was previously characterized across three mouse strains and in one mouse strain with various degrees of liver steatosis. To further characterize the role of genetic variability in toxicokinetics of perc, we applied Bayesian population physiologically-based pharmacokinetic (PBPK) modeling to the data on perc and metabolites in blood/plasma and tissues of male mice from 45 inbred strains from the Collaborative Cross (CC) mouse population. After identifying the most influential PBPK parameters based on global sensitivity analysis, we fit the model with a hierarchical Bayesian population analysis using Markov chain Monte Carlo simulation. We found that the data from three commonly used strains were not representative of the full range of variability in perc and metabolite blood/plasma and tissue concentrations across the CC population. Using inter-strain variability as a surrogate for human inter-individual variability, we calculated dose-dependent, chemical-, and tissue-specific toxicokinetic variability factors (TKVFs) as candidate science-based replacements for the default uncertainty factor for human toxicokinetic variability of 10<sup>0.5</sup>. We found that TKVFs for glutathione conjugation metabolites of perc showed the greatest variability, often exceeding the default, whereas those for oxidative metabolites and perc itself were generally less than the default. Overall, we demonstrate how a combination of a population-based mouse model such as the CC with Bayesian population PBPK modeling can reduce uncertainty associated with toxicokinetic human variability by deriving the chemical-specific adjustment factors needed to increase accuracy and precision in quantitative risk assessment.</p>
The NANOGrav Search for Signals from New Physics: MCMC chains
<p>MCMC chains for the GWB analyses performed in the paper "<em>The NANOGrav 15 yr Data Set: Search for Signals from New Physics</em>". </p> <p>The data is provided in pickle format. Each file contains a NumPy array with the MCMC chain (with burn-in already removed), and a dictionary with the model parameters' names as keys and their priors as values. You can load them as</p> <pre><code class="language-python">with open ('path/to/file.pkl', 'rb') as pick: temp = pickle.load(pick) params = temp[0] chain = temp[1]</code></pre> <p>The naming convention for the files is the following:</p> <ul> <li><strong>igw</strong>: inflationary Gravitational Waves (GWs)</li> <li>sigw: scalar-induced GWs <ul> <li><strong>sigw_box</strong>: assumes a box-like feature in the primordial power spectrum.</li> <li><strong>sigw_delta</strong>: assumes a delta-like feature in the primordial power spectrum.</li> <li><strong>sigw_gauss</strong>: assumes a Gaussian peak feature in the primordial power spectrum.</li> </ul> </li> <li>pt: cosmological phase transitions <ul> <li><strong>pt_bubble</strong>: assumes that the dominant contribution to the GW productions comes from bubble collisions.</li> <li><strong>pt_sound</strong>: assumes that the dominant contribution to the GW productions comes from sound waves.</li> </ul> </li> <li>stable: stable cosmic strings <ul> <li><strong>stable-c</strong>: stable strings emitting GWs only in the form of GW bursts from cusps on closed loops.</li> <li><strong>stable-k</strong>: stable strings emitting GWs only in the form of GW bursts from kinks on closed loops.</li> <li><strong>stable</strong>-<strong>m</strong>: stable strings emitting monochromatic GW at the fundamental frequency.</li> <li><strong>stable-n</strong>: stable strings described by numerical simulations including GWs from cusps and kinks.</li> </ul> </li> <li>meta: metastable cosmic strings <ul> <li><strong>meta</strong>-<strong>l</strong>: metastable strings with GW emission from loops only.</li> <li><strong>meta-ls</strong> metastable strings with GW emission from loops and segments.</li> </ul> </li> <li><strong>super</strong>: cosmic superstrings.</li> <li>dw: domain walls <ul> <li><strong>dw-sm</strong>: domain walls decaying into Standard Model particles.</li> <li><strong>dw-dr</strong>: domain walls decaying into dark radiation.</li> </ul> </li> </ul> <p>For each model, we provide four files. One for the run where the new-physics signal is assumed to be the only GWB source. One for the run where the new-physics signal is superimposed to the signal from Supermassive Black Hole Binaries (SMBHB), for these files "_bhb" will be appended to the model name. Then, for both these scenarios, in the "compare" folder we provide the files for the hypermodel runs that were used to derive the Bayes' factors.</p> <p>In addition to chains for the stochastic models, we also provide data for the two deterministic models considered in the paper (ULDM and DM substructures). For the ULDM model, the naming convention of the files is the following (all the ULDM signals are superimposed to the SMBHB signal, see the discussion in the paper for more details)</p> <ul> <li><strong>uldm_e</strong>: ULDM Earth signal.</li> <li>uldm_p: ULDM pulsar signal <ul> <li><strong>uldm_p_cor</strong>: correlated limit</li> <li><strong>uldm_p_unc</strong>: uncorrelated limit</li> </ul> </li> <li>uldm_c: ULDM combined Earth + pulsar signal direct coupling <ul> <li><strong>uldm_c_cor</strong>: correlated limit</li> <li><strong>uldm_c_unc</strong>: uncorrelated limit</li> </ul> </li> <li>uldm_vecB: vector ULDM coupled to the baryon number <ul> <li><strong>uldm_vecB_cor:</strong> correlated limit</li> <li><strong>uldm_vecB_unc</strong>: uncorrelated limit </li> </ul> </li> <li>uldm_vecBL: vector ULDM coupled to B-L <ul> <li><strong>uldm_vecBL_cor:</strong> correlated limit</li> <li><strong>uldm_vecBL_unc</strong>: uncorrelated limit</li> </ul> </li> <li>uldm_c_grav: ULDM combined Earth + pulsar signal for gravitational-only coupling <ul> <li>uldm_c_grav_cor: correlated limit <ul> <li><strong>uldm_c_cor_grav_low</strong>: low mass region </li> <li><strong>uldm_c_cor_grav_mon</strong>: monopole region</li> <li><strong>uldm_c_cor_grav_low</strong>: high mass region</li> </ul> </li> <li><strong>uldm_c_unc</strong>: uncorrelated limit <ul> <li><strong>uldm_c_unc_grav_low</strong>: low mass region </li> <li><strong>uldm_c_unc_grav_mon</strong>: monopole region</li> <li><strong>uldm_c_unc_grav_low</strong>: high mass region</li> </ul> </li> </ul> </li> </ul> <p>For the substructure (static) model, we provide the chain for the marginalized distribution (as for the ULDM signal, the substructure signal is always superimposed to the SMBHB signal)</p>
MCMC output files for: Quantitative characterization of population-wide tissue- and metabolite-specific variability in perchloroethylene toxicokinetics in male mice
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.