Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “hidden Markov model”
Multi-parameter photon-by-photon hidden Markov modeling dataset
<p>The core jupyter notebooks demonstrating mpH<sup>2</sup>MM using real nsALEX data on DNA hairpin, the maltose binding protein MalE, and the type III secretion system effector YopO. Also included are the jupyter notebooks for generating simulated photon trajectories to test the validity of the Integrated complete likelihood (ICL) and mpH<sup>2</sup>MM on data where the ground truth is known.</p> <p>Included are all HDF5 files used by the jupyter notebooks. Those for MalE and YopO are contained in zip files, and should be unziped maintaining the directory structure. All other HDF5 files should be kept in the same folder as the jupyter notebooks. Additional figures and csv files of H2MM models are included in separate zip file.</p>
Figure 8. Exchange mutation.-Neuroevolution Mechanism for Hidden Markov Model
<p>Exchange two neighbor weights involved in a summation of 1.0. This is illustrated in Figure 8.</p>
Profile Comparer Extended: phylogeny of LPMO families using profile hidden Markov model alignments
<p>searchable pdf phylogenetic tree (Fig S1) and sequence data (Table S1) belonging to the paper "Profile Comparer Extended: phylogeny of LPMO families using profile hidden Markov model alignments"</p>
Hidden Markov models with serial correlation for identifying stock-recruitment regime shifts
Open the record for dataset details and reuse information.
From pup to predator; generalized hidden Markov models reveal rapid development of movement strategies in a naïve long‐lived vertebrate
Open the record for dataset details and reuse information.
Data from: Hidden Markov models reveal tactical adjustment of temporally-clustered courtship displays in response to the behaviors of a robotic female
Open the record for dataset details and reuse information.
Data from: A hidden Markov model to identify and adjust for selection bias: an example involving mixed migration strategies
An important assumption in observational studies is that sampled individuals are representative of some larger study population. Yet, this assumption is often unrealistic. Notable examples include online public-opinion polls, publication biases associated with statistically significant results, and in ecology, telemetry studies with significant habitat-induced probabilities of missed locations. This problem can be overcome by modeling selection probabilities simultaneously with other predictor–response relationships or by weighting observations by inverse selection probabilities. We illustrate the problem and a solution when modeling mixed migration strategies of northern white-tailed deer (Odocoileus virginianus). Captures occur on winter yards where deer migrate in response to changing environmental conditions. Yet, not all deer migrate in all years, and captures during mild years are more likely to target deer that migrate every year (i.e., obligate migrators). Characterizing deer as conditional or obligate migrators is also challenging unless deer are observed for many years and under a variety of winter conditions. We developed a hidden Markov model where the probability of capture depends on each individual's migration strategy (conditional versus obligate migrator), a partially latent variable that depends on winter severity in the year of capture. In a 15-year study, involving 168 white-tailed deer, the estimated probability of migrating for conditional migrators increased nonlinearly with an index of winter severity. We estimated a higher proportion of obligates in the study cohort than in the population, except during a span of 3 years surrounding back-to-back severe winters. These results support the hypothesis that selection biases occur as a result of capturing deer on winter yards, with the magnitude of bias depending on the severity of winter weather. Hidden Markov models offer an attractive framework for addressing selection biases due to their ability to incorporate latent variables and model direct and indirect links between state variables and capture probabilities.
Data from: Joint modelling of multi-scale animal movement data using hierarchical hidden Markov models
1. Hidden Markov models are prevalent in animal movement modelling, where they are widely used to infer behavioural modes and their drivers from various types of telemetry data. To allow for meaningful inference, observations need to be equally spaced in time, or otherwise regularly sampled, where the corresponding temporal resolution strongly affects what kind of behaviours can be inferred from the data. 2. Recent advances in biologging technology have led to a variety of novel telemetry sensors which often collect data from the same individual simultaneously at different time scales, e.g. step lengths obtained from GPS tags every hour, dive depths obtained from time-depth recorders once per dive, or accelerations obtained from accelerometers several times per second. However, to date, statistical machinery to address the corresponding complex multi-stream and multi-scale data is lacking. 3. We propose hierarchical hidden Markov models as a versatile statistical framework that naturally accounts for differing temporal resolutions across multiple variables. In these models, the observations are regarded as stemming from multiple, connected behavioural processes, each of which operates at the time scale at which the corresponding variables were observed. 4. By jointly modelling multiple data streams, collected at different temporal resolutions, corresponding models can be used to infer behavioural modes at multiple time scales, and in particular help to draw a much more comprehensive picture of an animal's movement patterns, e.g. with regard to long-term vs. short-term movement strategies. 5. The suggested approach is illustrated in two real-data applications, where we jointly model i) coarse-scale horizontal and fine-scale vertical Atlantic cod (Gadus morhua) movements throughout the English Channel, and ii) coarse-scale horizontal movements and corresponding fine-scale accelerations of a horn shark (Heterodontus francisci) tagged off the Californian coast.
Data from: Use of hidden Markov capture-recapture models to estimate abundance in presence of uncertainty: application to estimating the prevalence of hybrids in animal populations
Estimating the relative abundance (prevalence) of different population segments is a key step in addressing fundamental research questions in ecology, evolution, and conservation. The raw percentage of individuals in the sample (naive prevalence) is generally used for this purpose, but it is likely to be subject to two main sources of bias. First, the detectability of individuals is ignored; second, classification errors may occur due to some inherent limits of the diagnostic methods. We developed a hidden Markov (also known as multievent) capture–recapture model to estimate prevalence in free‐ranging populations accounting for imperfect detectability and uncertainty in individual's classification. We carried out a simulation study to compare naive and model‐based estimates of prevalence and assess the performance of our model under different sampling scenarios. We then illustrate our method with a real‐world case study of estimating the prevalence of wolf (Canis lupus) and dog (Canis lupus familiaris) hybrids in a wolf population in northern Italy. We showed that the prevalence of hybrids could be estimated while accounting for both detectability and classification uncertainty. Model‐based prevalence consistently had better performance than naive prevalence in the presence of differential detectability and assignment probability and was unbiased for sampling scenarios with high detectability. We also showed that ignoring detectability and uncertainty in the wolf case study would lead to underestimating the prevalence of hybrids. Our results underline the importance of a model‐based approach to obtain unbiased estimates of prevalence of different population segments. Our model can be adapted to any taxa, and it can be used to estimate absolute abundance and prevalence in a variety of cases involving imperfect detection and uncertainty in classification of individuals (e.g., sex ratio, proportion of breeders, and prevalence of infected individuals).
Data from: Analysis of animal accelerometer data using hidden Markov models
Use of accelerometers is now widespread within animal biologging as they provide a means of measuring an animal's activity in a meaningful and quantitative way where direct observation is not possible. In sequential acceleration data, there is a natural dependence between observations of behaviour, a fact that has been largely ignored in most analyses. Analyses of acceleration data where serial dependence has been explicitly modelled have largely relied on hidden Markov models (HMMs). Depending on the aim of an analysis, an HMM can be used for state prediction or to make inferences about drivers of behaviour. For state prediction, a supervised learning approach can be applied. That is, an HMM is trained to classify unlabelled acceleration data into a finite set of pre-specified categories. An unsupervised learning approach can be used to infer new aspects of animal behaviour when biologically meaningful response variables are used, with the caveat that the states may not map to specific behaviours. We provide the details necessary to implement and assess an HMM in both the supervised and unsupervised learning context and discuss the data requirements of each case. We outline two applications to marine and aerial systems (shark and eagle) taking the unsupervised learning approach, which is more readily applicable to animal activity measured in the field. HMMs were used to infer the effects of temporal, atmospheric and tidal inputs on animal behaviour. Animal accelerometer data allow ecologists to identify important correlates and drivers of animal activity (and hence behaviour). The HMM framework is well suited to deal with the main features commonly observed in accelerometer data and can easily be extended to suit a wide range of types of animal activity data. The ability to combine direct observations of animal activity with statistical models, which account for the features of accelerometer data, offers a new way to quantify animal behaviour and energetic expenditure and to deepen our insights into individual behaviour as a constituent of populations and ecosystems.
Dataset from "Employing hidden Markov models to assess the genetic content of genome assemblies"
<p>Dataset used to reach the conclusions in "Employing hidden Markov models to assess the genetic content of genome assemblies".</p>
Accompanying simulated data for "Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity"
<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated for the manuscript: "Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity". It comprehends: (1) model outputs (maximum a posteriori estimates) for each repetition (n=100) of each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for 4000 MCMC iterations) for two chains of each repetition (n=3) of each scenario (n=324). Please note that the empirical data used in the manuscript is not available as part of this repository. A subsample of the data used in the empirical example are openly available as an example data set in the R package <a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set is available on request from the authors.</p>
Datasets for optical tweezer autoregressive hidden Markov modeling
<p>Data to regenerate figures and tables analyzing optical tweezer data with arHMM. OT_arHMM.zip contains the ot_arhmm library needed to analyze these files. scripts.zip contains all the jupyter notebooks used to analyze data and make figures and tables. data.zip contains the raw data and processed_data.zip contains data that has already been put through the HMMs for analysis. These processed files can be generated from the raw data by running the analyze_all.ipynb notebook. All other notebooks require processed files to be present before running.</p>
Data from: A hidden Markov model to identify and adjust for selection bias: an example involving mixed migration strategies
Open the record for dataset details and reuse information.
Data from: Analysis of animal accelerometer data using hidden Markov models
Open the record for dataset details and reuse information.
Data from: Use of hidden Markov capture-recapture models to estimate abundance in presence of uncertainty: application to estimating the prevalence of hybrids in animal populations
Open the record for dataset details and reuse information.
Data from: Joint modelling of multi-scale animal movement data using hierarchical hidden Markov models
Open the record for dataset details and reuse information.
Data for "An application of upscaled optimal foraging theory using hidden Markov modelling: year-round behavioural variation in a large arctic herbivore"
<p>Data for the article “An application of upscaled optimal foraging theory using hidden Markov modelling: year-round behavioural variation in a large arctic herbivore”</p> <p>By LT Beumer, J Pohle, NMS Schmidt, M Chimienti, JP Desforges, LH Hansen, R Langrock, SH Pedersen, M Stelvig, FM van Beest</p> <p> </p> <p>The data set includes three files: A readme file describing the data files and two data files accompanying the above publication.</p> <p>Combined, the two data files represent the dataset collected by GPS collars fitted on 19 female muskoxen in northeast Greenland (28 muskox-years with 153-1062 observation days/animal) and associated extracted covariates, divided into a summer and winter season dataset as modelled in the article. Data here are given as included in the models (for a description of cleaning procedures, see article). All continuous, non-cyclical covariates were standardised to have zero mean and unit standard deviation to improve numerical stability of parameter estimation. This is indicated by “_scaled” in the column name.</p> <p>For further queries please contact nms@bios.au.dk</p>
Generalized hidden Markov models for phylogenetic comparative datasets
<ol> <li class="JamesManuscriptBody">Hidden Markov models (HMM) have emerged as an important tool for understanding the evolution of characters that take on discrete states. Their flexibility and biological sensibility make them appealing for many phylogenetic comparative applications.</li> <li class="JamesManuscriptBody">Previously available packages placed unnecessary limits on the number of observed and hidden states that can be considered when estimating transition rates and inferring ancestral states on a phylogeny.</li> <li class="JamesManuscriptBody">To address these issues, we expanded the capabilities of the R package corHMM to handle <i>n</i>-state and <i>n</i>-character problems and provide users with a streamlined set of functions to create custom HMMs for any biological question of arbitrary complexity.</li> <li class="JamesManuscriptBody">We show that increasing the number of observed states increases the accuracy of ancestral state reconstruction. We also explore the conditions for when an HMM is most effective, finding that an HMM is an appropriate model when the degree of rate heterogeneity is moderate to high.</li> <li class="JamesManuscriptBody">Finally, we demonstrate the importance of these generalizations by reconstructing the phyllotaxy of the ancestral angiosperm flower. Partially contradicting previous results, we find the most likely state to be a whorled perianth, whorled androecium, whorled gynoecium. The difference between our analysis and previous studies was that our modeling explicitly allowed for the correlated evolution of several flower characters.</li> </ol>
Supplementary material 1 from: Ferreira EM, Valerio F, Medinas D, Fernandes N, Craveiro J, Costa P, Silva JP, Carrapato C, Mira A, Santos SM (2022) Assessing behaviour states of a forest carnivore in a road-dominated landscape using Hidden Markov Models. In: Santos S, Grilo C, Shilling F, Bhardwaj M, Papp CR (Eds) Linear Infrastructure Networks with Ecological Solutions. Nature Conservation 47: 155-175. https://doi.org/10.3897/natureconservation.47.72781
Figures S1–S3
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.