Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “Markov chains”
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Figure 9. Experimental Page Rank dependency on Markov Chain length with balanced distribution-Study of a Random Navigation on the Web Using Software Simulation
<p>This paper explored different implementation choices for analyzing the most important<br> parameters about a web. Many researchers explored the use of new search engines for studying the<br> evolution of the web (Ntoulas, Cho and Olston, 2004). Another important research is realized about<br> the Link Structure Graph (LSG). The LSG captures a complete hyperlink structure from the web<br> and models link associations reflected in the page layout (Rodrigues, Milic-Frayling and Fortuna,<br> 2007). For further works ideas like extrapolation methods for accelerating page rank calculation can<br> be developed (Kamvar et al., 2003).</p>
Figure 7. Experimental Page Rank dependency on Markov Chain length-Study of a Random Navigation on the Web Using Software Simulation
<p>The next diagram proves that the values for Experimental Page Rank depend on the length<br> of the Markov Chain, while Algorithmic Page Rank remains constant.</p>
Figure 1. Markov Chain Model&Figure 2. Transition matrix-Study of a Random Navigation on the Web Using Software Simulation
<p>For a good simulation it is very important to find methods for<br> navigating through the web (Levene and Wheeldon, 2004). John Kemeny and Laurie Snell have<br> proposed the use of Markov models for web simulations (Kemeny and Snell, 1960). Cadez et al. (2000)<br> used Markov models for classifying the sessions into different categories for browsers. Some other<br> proposed techniques choose to combine different order Markov models for obtaining low state<br> complexity and improving accuracy, as Deshpande and Karypis (2004). Dongshan and Junyi (2002)<br> used for predicting the access providing good scalability and high coverage a hybrid-order tree-like<br> Markov model. As an alternative to the Markov model Pitkow proposed a longest subsequence model<br> (Pitkow and Pirolli, 1999), also for predicting the next page accessed by the user Sarukkai chose<br> Markov models (Sarukkai, 2000).<br> Transitions are simulated using the Markov Chain nodes, Google matrix and an arbitrary initial<br> probability distribution. Examples can be seen in Figure 1 and Figure 2.</p>
Timing data for algorithms for calculating steady state distributions of continuous time Markov chains
<p>Data showing the timings for a number of algorithms for the computation of the steady state of a continuous time Markov chain:</p> <ul> <li>Numeric integration;</li> <li>Matrix exponential;</li> <li>Eigenvector;</li> <li>Solve a linear algebraic system;</li> <li>Approximately solving a linear algebraic system.</li> </ul> <p>All calculations were done using implementations in Python for a blog post at https://vknight.org/blog/</p>
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II
Open the record for dataset details and reuse information.
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I
Open the record for dataset details and reuse information.
Plasim dataset for manuscript titled: "Extreme heatwave sampling and prediction with analog Markov chain and comparisons with deep learning"
<p>This dataset was originally generated by running a Plasim general circulation model for 8000 years. We extract a small subset of this dataset which is only 500 years long with temperature, soil moisture and 500 hPa Geopotential extracted over the northern hemisphere above 30 North. This dataset can be used to reproduce the results in https://arxiv.org/abs/2307.09060 concerning training CNN and SWG with cross-validation length of 500 years. </p>
Author Name Disambiguation using Markov Chain Monte Carlo (MCMC)
<p>This repository contains the files processed from the <a href="https://zenodo.org/record/5675801#.Y2APcOzMJhE">Aminer-534K</a> Knowledge graph for the master's thesis on Author Name Disambiguation.</p>
Prioritizing test cases with Markov Chains: a Preliminary Investigation
<p>This replication package includes the inputs to replicate the results shown in a paper under revision. Data is structured in four folders, one for each considered case study (there is an initial example besides the three case studies reported in Section 5). Within each folder, it is possible to inspect the input files as well as the results, for both proposed heuristics, i.e., H1 and H2. The input formats can be consulted to understanding and replicating the study. The dataset also contains the algorithm written in Python.</p>
Data extraction form - A Systematic Literature Review on Prioritizing Software Test Cases using Markov Chains
<p>A data extraction form was created to gather all relevant data from the identified studies and manage the selection process in this systematic literature review. Some of the main information presented in this form was followed by a protocol, identifier (id) for each returned study, bibliographic reference, and answers to research questions. This catalog helps us in the data extraction and synthesis procedures and may be used by potentially interested, for example, for updating or replication. </p>
Age-by-maternal-age Markov chain with rewards analysis of variance in LRO in rotifers
<p>This set of data and code files is a supplement to the American Naturalist paper:<br> van Daalen et al. 2021 The contributions of maternal age heterogeneity to variance in lifetime reproductive output</p> <p>It provides the user the matrices (as first presented in Hernandez et al., 2020, PNAS) and the code to reproduce our analysis or apply the methods to their own data. Two multistate, age-by-maternal-age matrices with basic demographic information (survival and transitions, and fertility) are provided, as well as data to paramaterize a reward matrix. The code presents a Markov chain with rewards approach to calculating mean and variance in lifetime reproductive output from a multistate matrix model, and a method to decompose variance into contributions by individual heterogeneity (from maternal age) and stochasticity (due to inherent randomness in the outcomes of age-specific survival probabilities and age-specific reproductive output).</p>
A study on the Application of Markov Chains to Prioritize Test Cases
<p>The entire results and the inputs to run the experiment are available in this dataset. The data is composed of five folders one for each case study and one for the running example presented in Section 3. Within each folder, it is possible do see the input files, as well as the results, both for H1 and H2. The input formats can be consulted to understanding and potentially replication of the study. This dataset also contains the algorithm, in Python. The algorithm is available on GitHub, we did not put the link because of the blind review.</p>
Computing Expected Visiting Times and Stationary Distributions in Markov Chains: Fast and Accurate (Artifact)
<p>This artifact contains the raw data of our experiments as well as scripts and benchmarks to reproduce the experiments.<br>Furthermore, the considered version of [Storm](http://stormchecker.org) is included, which contains our implementation.</p> <p>Please also consider the artifact of the conference paper available at [zenodo](https://zenodo.org/records/10438916) which has been accepted by the TACAS Artifact evaluation committee.</p> <p><br>This artifact contains: <br>`LICENSE`: The license document.<br>`README.md`: The instructions.<br>`raw_data.zip`: The raw data obtained during our experiments<br>`raw_data_with_results.zip`: The raw data, also including the resulting stationary distributions and evts in an explicit format. (84 GB!)<br>`reproduce.zip` contains benchmarks and scripts for reproducing the experiments<br>`storm-0b1cae2a94f06984f3cf4cecf5a5090e9bc71a56.zip` is the exact Storm version we considered.</p>
Data from: Markov-modulated continuous-time Markov chains to identify site- and branch-specific evolutionary variation in BEAST
<p>Markov models of character substitution on phylogenies form the foundation of phylogenetic inference frameworks. Early models made the simplifying assumption that the substitution process is homogeneous over time and across sites in the molecular sequence alignment. While standard practice adopts extensions that accommodate heterogeneity of substitution rates across sites, heterogeneity in the process over time in a site-specific manner remains frequently overlooked. This is problematic, as evolutionary processes that act at the molecular level are highly variable, subjecting different sites to different selective constraints over time, impacting their substitution behaviour. We propose incorporating time variability through Markov-modulated models (MMMs), which extend covarion-like models and allow the substitution process (including relative character exchange rates as well as the overall substitution rate) at individual sites to vary across lineages. We implement a general MMM framework in BEAST, a popular Bayesian phylogenetic inference software package, allowing researchers to compose a wide range of MMMs through flexible XML specification. Using examples from bacterial, viral and plastid genome evolution, we show that MMMs impact phylogenetic tree estimation and can substantially improve model fit compared to standard substitution models. Through simulations, we show that marginal likelihood estimation accurately identifies the generative model and does not systematically prefer the more parameter-rich MMMs. To mitigate the increased computational demands associated with MMMs, our implementation exploits recent developments in BEAGLE, a high-performance computational library for phylogenetic inference.</p>
Inferring Log-Based Behavioural System Models using Markov Chains
<p>The datasets we used for the bachelor thesis <em><strong>Inferring Log-Based Behavioural System Models using Markov Chains</strong></em>, consisting of log traces of the XRP Ledger Conensus Protocol split into 5 different datasets.</p>
Affects Affect Affects: A Markov Chain
<p>Data for the article "<strong>Affects Affect Affects: A Markov Chain"</strong></p> <p>Data calculation for the probability vector in each step with the State transitions matrix, with initial state (S<sub>0</sub>) and the resulting calculation of the steady-state vector.</p>
Experiments for 'Efficient Sensitivity Analysis for Parametric Robust Markov Chains'
<p>This artifact accompanies the CAV 2023 paper with the title 'Efficient Sensitivity Analysis for Parametric Robust Markov Chains'. The artifact contains a docker file, which can be unzipped and then loaded with:</p> <pre><code>docker load -i prmc_sensitivity_cav23_docker.tar</code></pre> <p>Depending on your permissions, you may need to run this command with sudo. Please refer to the ReadMe for more information.</p> <p>The source code of the docker container is available on GitHub: <a href="https://github.com/LAVA-LAB/prmc-sensitivity">https://github.com/LAVA-LAB/prmc-sensitivity</a>.</p>
Data from: samc: An R package for connectivity modeling with spatial absorbing Markov chains
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.