Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

31

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

31 results for “Markov chains”

Learn how ShareScore rates datasets ↗
dryad40/100

Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I

<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>

opencc-zeroNov 2020View details →
dryad40/100

Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II

<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>

opencc-zeroNov 2020View details →
zenodo40/100

Figure 9. Experimental Page Rank dependency on Markov Chain length with balanced distribution-Study of a Random Navigation on the Web Using Software Simulation

<p>This paper explored different implementation choices for analyzing the most important<br> parameters about a web. Many researchers explored the use of new search engines for studying the<br> evolution of the web (Ntoulas, Cho and Olston, 2004). Another important research is realized about<br> the Link Structure Graph (LSG). The LSG captures a complete hyperlink structure from the web<br> and models link associations reflected in the page layout (Rodrigues, Milic-Frayling and Fortuna,<br> 2007). For further works ideas like extrapolation methods for accelerating page rank calculation can<br> be developed (Kamvar et al., 2003).</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 7. Experimental Page Rank dependency on Markov Chain length-Study of a Random Navigation on the Web Using Software Simulation

<p>The next diagram proves that the values for Experimental Page Rank depend on the length<br> of the Markov Chain, while Algorithmic Page Rank remains constant.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Figure 1. Markov Chain Model&Figure 2. Transition matrix-Study of a Random Navigation on the Web Using Software Simulation

<p>For a good simulation it is very important to find methods for<br> navigating through the web (Levene and Wheeldon, 2004). John Kemeny and Laurie Snell have<br> proposed the use of Markov models for web simulations (Kemeny and Snell, 1960). Cadez et al. (2000)<br> used Markov models for classifying the sessions into different categories for browsers. Some other<br> proposed techniques choose to combine different order Markov models for obtaining low state<br> complexity and improving accuracy, as Deshpande and Karypis (2004). Dongshan and Junyi (2002)<br> used for predicting the access providing good scalability and high coverage a hybrid-order tree-like<br> Markov model. As an alternative to the Markov model Pitkow proposed a longest subsequence model<br> (Pitkow and Pirolli, 1999), also for predicting the next page accessed by the user Sarukkai chose<br> Markov models (Sarukkai, 2000).<br> Transitions are simulated using the Markov Chain nodes, Google matrix and an arbitrary initial<br> probability distribution. Examples can be seen in Figure 1 and Figure 2.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Timing data for algorithms for calculating steady state distributions of continuous time Markov chains

<p>Data showing the timings for a number of algorithms for the computation of the steady state of a continuous time Markov chain:</p> <ul> <li>Numeric integration;</li> <li>Matrix exponential;</li> <li>Eigenvector;</li> <li>Solve a linear algebraic system;</li> <li>Approximately solving a linear algebraic system.</li> </ul> <p>All calculations were done using implementations in Python for a blog post at https://vknight.org/blog/</p>

opencc-by-4.0Dec 2018View details →
dryad40/100

Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II

Open the record for dataset details and reuse information.

publicNov 2020View details →
dryad40/100

Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I

Open the record for dataset details and reuse information.

publicNov 2020View details →
zenodo36/100

Plasim dataset for manuscript titled: "Extreme heatwave sampling and prediction with analog Markov chain and comparisons with deep learning"

<p>This dataset was originally generated by running a Plasim general circulation model for 8000 years. We extract a small subset of this dataset which is only 500 years long with temperature, soil moisture and 500 hPa Geopotential extracted over the northern hemisphere above 30 North. This dataset can be used to reproduce the results in https://arxiv.org/abs/2307.09060 concerning training CNN and SWG with cross-validation length of 500 years.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Author Name Disambiguation using Markov Chain Monte Carlo (MCMC)

<p>This repository contains the files processed from the <a href="https://zenodo.org/record/5675801#.Y2APcOzMJhE">Aminer-534K</a> Knowledge graph for the master&#39;s thesis on Author Name Disambiguation.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Prioritizing test cases with Markov Chains: a Preliminary Investigation

<p>This replication package includes the inputs to replicate the results shown in a paper under revision. Data is structured in four&nbsp;folders, one for each considered case study (there is an initial example besides the three&nbsp;case studies reported in Section 5). Within each folder, it is possible to inspect the input files as well as the results, for both proposed heuristics, i.e., H1 and H2. The input formats can be consulted to understanding and replicating the study. The dataset also contains the algorithm written in Python.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Data extraction form - A Systematic Literature Review on Prioritizing Software Test Cases using Markov Chains

<p>A data extraction form was created to gather all relevant data from the identified studies and manage the selection process in this systematic literature review. Some of the main information presented in this form was followed by a protocol, identifier (id) for each returned study, bibliographic reference, and answers to research questions. This catalog helps us in the data extraction and synthesis procedures and may be used by potentially interested, for example, for updating or replication.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
dryad32/100

Age-by-maternal-age Markov chain with rewards analysis of variance in LRO in rotifers

<p>This set of data and code files is a supplement to the American Naturalist paper:<br> van Daalen et al. 2021 The contributions of maternal age heterogeneity to variance in lifetime reproductive output</p> <p>It provides the user the matrices (as first presented in Hernandez et al., 2020, PNAS) and the code to reproduce our analysis or apply the methods to their own data. Two multistate, age-by-maternal-age matrices with basic demographic information (survival and transitions, and fertility) are provided, as well as data to paramaterize a reward matrix. The code presents a Markov chain with rewards approach to calculating mean and variance in lifetime reproductive output from a multistate matrix model, and a method to decompose variance into contributions by individual heterogeneity (from maternal age) and stochasticity (due to inherent randomness in the outcomes of age-specific survival probabilities and age-specific reproductive output).</p>

opencc-zeroMar 2022View details →
zenodo32/100

A study on the Application of Markov Chains to Prioritize Test Cases

<p>The entire&nbsp;results and the inputs to run the experiment are available in this dataset.&nbsp;The data is composed of five folders one for each case study&nbsp;and one for the running example&nbsp;presented in Section 3. Within each folder, it is possible do see the input files, as well as the results, both for H1 and H2.&nbsp;The input formats can be consulted to understanding and potentially replication of the study. This dataset also contains&nbsp;the algorithm, in Python.&nbsp;The algorithm is available on GitHub, we did not put the link because of the blind review.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Computing Expected Visiting Times and Stationary Distributions in Markov Chains: Fast and Accurate (Artifact)

<p>This artifact contains the raw data of our experiments as well as scripts and benchmarks to reproduce the experiments.<br>Furthermore, the considered version of [Storm](http://stormchecker.org) is included, which contains our implementation.</p> <p>Please also consider the artifact of the conference paper available at [zenodo](https://zenodo.org/records/10438916) which has been accepted by the TACAS Artifact evaluation committee.</p> <p><br>This artifact contains:&nbsp;<br>`LICENSE`: The license document.<br>`README.md`: The instructions.<br>`raw_data.zip`: The raw data obtained during our experiments<br>`raw_data_with_results.zip`: The raw data, also including the resulting stationary distributions and evts in an explicit format. (84 GB!)<br>`reproduce.zip` contains benchmarks and scripts for reproducing the experiments<br>`storm-0b1cae2a94f06984f3cf4cecf5a5090e9bc71a56.zip` is the exact Storm version we considered.</p>

opencc-by-4.0Oct 2024View details →
dryad32/100

Data from: Markov-modulated continuous-time Markov chains to identify site- and branch-specific evolutionary variation in BEAST

<p>Markov models of character substitution on phylogenies form the foundation of phylogenetic inference frameworks. Early models made the simplifying assumption that the substitution process is homogeneous over time and across sites in the molecular sequence alignment. While standard practice adopts extensions that accommodate heterogeneity of substitution rates across sites, heterogeneity in the process over time in a site-specific manner remains frequently overlooked. This is problematic, as evolutionary processes that act at the molecular level are highly variable, subjecting different sites to different selective constraints over time, impacting their substitution behaviour. We propose incorporating time variability through Markov-modulated models (MMMs), which extend covarion-like models and allow the substitution process (including relative character exchange rates as well as the overall substitution rate) at individual sites to vary across lineages. We implement a general MMM framework in BEAST, a popular Bayesian phylogenetic inference software package, allowing researchers to compose a wide range of MMMs through flexible XML specification. Using examples from bacterial, viral and plastid genome evolution, we show that MMMs impact phylogenetic tree estimation and can substantially improve model fit compared to standard substitution models. Through simulations, we show that marginal likelihood estimation accurately identifies the generative model and does not systematically prefer the more parameter-rich MMMs. To mitigate the increased computational demands associated with MMMs, our implementation exploits recent developments in BEAGLE, a high-performance computational library for phylogenetic inference.</p>

opencc-zeroMay 2020View details →
zenodo32/100

Inferring Log-Based Behavioural System Models using Markov Chains

<p>The datasets we used for the bachelor thesis&nbsp;<em><strong>Inferring Log-Based Behavioural System Models using Markov Chains</strong></em>, consisting of log traces of the XRP Ledger Conensus Protocol&nbsp;split into 5&nbsp;different datasets.</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Affects Affect Affects: A Markov Chain

<p>Data for the article &quot;<strong>Affects Affect Affects: A Markov Chain&quot;</strong></p> <p>Data calculation for the probability vector in each step with the State transitions matrix, with initial state (S<sub>0</sub>) and the resulting calculation of the steady-state vector.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Experiments for 'Efficient Sensitivity Analysis for Parametric Robust Markov Chains'

<p>This artifact accompanies the CAV 2023&nbsp;paper with the title &#39;Efficient Sensitivity Analysis for Parametric Robust Markov Chains&#39;. The artifact contains a docker file, which can be unzipped and then loaded with:</p> <pre><code>docker load -i prmc_sensitivity_cav23_docker.tar</code></pre> <p>Depending on your permissions, you may need to run this command with sudo. Please refer to the ReadMe for more information.</p> <p>The source code of the docker container is available on GitHub:&nbsp;<a href="https://github.com/LAVA-LAB/prmc-sensitivity">https://github.com/LAVA-LAB/prmc-sensitivity</a>.</p>

openapgl-v3May 2023View details →
dryad32/100

Data from: samc: An R package for connectivity modeling with spatial absorbing Markov chains

Open the record for dataset details and reuse information.

publicDec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record