Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20
datasets available to search
ShareScore release 0.9.0
Dataset results
20 results for “multilevel models”
Accompanying simulated data for "Go multivariate: a Monte Carlo study of a multilevel hidden Markov model with categorical data of varying complexity"
<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated for the manuscript: "Go multivariate: a Monte Carlo study of a multilevel hidden Markov model with categorical data of varying complexity". It comprehends: (1) model outputs (maximum a posteriori estimates) for each repetition (n=100) of each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for 4000 MCMC iterations) for two chains of each repetition (n=3) of each scenario (n=324). Please note that the empirical data used in the manuscript is not available as part of this repository. A subsample of the data used in the empirical example are openly available as an example data set in the R package <a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set is available on request from the authors.</p>
Accompanying empirical data for Kirchherr et al., 2023, "Bayesian multilevel hidden Markov models identify stable state dynamics in longitudinal recordings from macaque primary motor cortex"
<p>This repository contains data accompanying: Kirchherr et al., 2023, "Bayesian multilevel hidden Markov models identify stable state dynamics in longitudinal recordings from macaque primary motor cortex".</p> <p>Data collection methods:</p> <p>Two adult female rhesus macaques (Macaca mulatta) trained on a reaching, and grasping, and placing task served as the subjects. The animal handling as well as surgical and experimental procedures complied with European guideline (2010/63/UE) and authorized by the French Ministry for Higher Education and Research (project # 2016112713202878) in force on the care and use of laboratory animals, and were approved by the ethics committee CELYNE (comité d’éthique Lyonnais pour les neurosciences expérimentale, C2EA 42). After initial training, we performed a sterile surgery to implant six floating multielectrode arrays (FMA, Microprobes for Life Science, Gaithersburg, MD, USA) in the right (monkey 1) or left (monkey 2) cortical hemisphere. Each array was comprised of 32 platinum/iridium electrodes (impedance 0.5 MΩ at 1 kHz) with lengths ranging from 1 to 6 mm, and with an inter-electrode spacing of 400 μm. One electrode array was implanted in the primary motor cortex (M1), two were implanted in the ventral premotor cortex (F5), one in the dorsal premotor cortex (F2), and two in the prefrontal cortex (45a and 46/12r), as estimated according to a previous magnetic resonance imaging scan. For the purposes of this study, we analyzed data from the M1 array of each monkey.</p> <p>The wideband neural signal (bandpass filtered at 0.1 to 7500 kHz) was recorded at 30 kS/s, and amplified and digitized (16-bit; 0.192 μV resolution) with an Intan Tech-based (Intan Technologies, Los Angeles, CA, USA) open source acquisition system (Open Ephys; Siegle et al. 2017). This system uses a 256-channel Intan RHD2000 series acquisition board and 32-channel headstages (RHD2132). Spike detection was performed offline using Trisdesclous (Garcia & Pouzat,2015). The common reference was removed to reduce ambient noise. Spikes were then detected from each electrode using a threshold of 2 times the median absolute deviation (MAD), and analyzed as multi-unit activity (MUA) in 10 ms bins. All electrodes in which at least one well-isolated spike waveform was detected were selected for the following analyses. We thus used a sample of 21 electrodes out of 32 for monkey 1, and 25 out of 32 electrodes for monkey 2. Custom made detection panels were used to record the moments when the monkey’s hand released the handle, the hand contacted the target object, and when the object was placed in the groove. An Omniplex 16-channel recording system (Plexon, Dallas, TX, USA) was used to simultaneously record these behavioral events. Trials were discarded if the response time (time between the go signal and handle release) was less than 100 or greater than 1500 ms, the reach duration (time between handle release and object contact) was less than 100 or greater than 1000 ms, or the placing duration (time between object contact and placing the object in the groove) was less than 100 or greater than 1200 ms, leaving 19 - 68 trials per day for monkey 1 (M = 43.9, SD = 15.46, N = 439; left: M = 14.8, SD = 5.74; center: M = 14.4, SD = 5.15; right: M = 14.7, SD = 7.73), and 23 - 49 per day for monkey 2 (M = 38.3, SD = 9.87, N = 383; left: M = 14.2, SD = 3.91; center: M = 10.8, SD = 3.55; right: M = 13.3, SD = 3.37).</p> <p><br> Abstract:</p> <p>Neural populations, rather than single neurons, may be the fundamental unit of cortical computation. Analyzing chronically recorded neural population activity is challenging not only because of the high dimensionality of activity in many neurons, but also because of changes in the recorded signal that may or may not be due to neural plasticity. Hidden Markov models (HMMs) are a promising technique for analyzing such data in terms of discrete, latent states, but previous approaches have either not considered the statistical properties of neural spiking data, have not been adaptable to longitudinal data, or have not modeled condition specific differences. We present a multilevel Bayesian HMM which addresses these shortcomings by incorporating multivariate Poisson log-normal emission probability distributions, multilevel parameter estimation, and trial-specific condition covariates. We applied this framework to multi-unit neural spiking data recorded using chronically implanted multi-electrode arrays from macaque primary motor cortex during a cued reaching, grasping, and placing task. We show that the model identifies latent neural population states which are tightly linked to behavioral events, despite the model being trained without any information about event timing. We show that these events represent specific spatiotemporal patterns of neural population activity and that their relationship to behavior is consistent over days of recording. The utility and stability of this approach is demonstrated using a previously learned task, but this multilevel Bayesian HMM framework would be especially suited for future studies of long-term plasticity in neural populations.</p>
Data from: Functional traits and community composition: a comparison among community-weighted means, weighted correlations, and multilevel models
1. Of the several approaches that are used to analyze functional trait-environment relationships, the most popular is community-weighted mean regressions (CWMr) in which species trait values are averaged at the site level and then regressed against environmental variables. Other approaches include model-based methods and weighted correlations of different metrics of trait-environment associations, the best known of which is the fourth-corner correlation method. 2. We investigated these three general statistical approaches for trait-environment associations: CWMr, five weighted correlation metrics (Peres-Neto et al. 2017), and two multilevel models (MLM) using four different methods for computing p-values. We first compared the methods applied to a plant community dataset. To determine the validity of the statistical conclusions, we then performed a simulation study. 3. CWMr gave highly significant associations for both traits, while the other methods gave a mix of support. CWMr had inflated type I errors for some simulation scenarios, implying that the significant results for the data could be spurious. The weighted correlation methods had generally good type I error control but had low power. One of the multilevel models, that from Jamil et al. (2013), had both good type I error control and high power when an appropriate method was used to obtain p-values. In particular, if there was no correlation among species in their abundances among sites, a parametric bootstrap likelihood ratio test (LRT) gave the best power. When there was correlation among species in their abundances, a conditional parametric LRT had correct type I errors but had lower power. 4. There is no overall best method for identifying trait-environment associations. For the simple task of testing, one-by-one, associations between single environmental variables and single traits, the weighted correlations with permutation tests all had good type I error control, and their ease of implementation is an advantage. For the more complex task of multivariate analyses and model fitting, and when high statistical power is needed, we recommend MLM2 (Jamil et al. 2013); however, care must be taken to ensure against inflated type I errors. Because CWMr exhibited highly inflated type I error rates, it should always be avoided. 2. We investigated these three general statistical approaches for trait-environment associations: CWMr, five weighted correlation metrics (Peres-Neto et al. 2017), and two multilevel models (MLM) using five different methods for computing p-values. We first compared the methods applied to a plant community dataset. To determine the validity of the statistical conclusions, we then performed a simulation study. 3. CWMr gave highly significant associations for both traits, while the other methods gave a mix of support. CWMr had inflated type I errors for some simulation scenarios. The weighted correlation methods had generally good type I error control but had low power. One of the multilevel models, that from Jamil et al. (2013), had both good type I error control and high power when an appropriate method was used to obtain p-values. In particular, if there was no correlation among species in their abundances among sites, a parametric bootstrap likelihood ratio test (LRT) gave the best power. When there was correlation among species in their abundances, a conditional parametric LRT had correct type I errors but suffered from low power. 4. There is no overall best method for identifying trait-environment associations. For the simple task of testing, one-by-one, associations between single environmental variables and single traits, the weighted correlations with permutation tests all had good type I error control, and their ease of implementation is an advantage. For the more complex task of multivariate analyses and model fitting, and when high statistical power is needed, we recommend MLM2 (Jamil et al. 2013); however, care must be taken to ensure against inflated type I errors. Because CWMr exhibited highly inflated type I error rates, it should be avoided.
High-quality video files for Hermsen, R, "Emergent multilevel selection in a simple spatial model of the evolution of altruism" (2021)
<p>The supplementary movies published with the article<br> <br> R. Hermsen<em>, Emergent multilevel selection in a simple spatial model of the evolution of altruism</em><br> <br> have a relatively low resolution. Here, the same three movies are provided at a higher resolution.</p> <p>Note: In Version 1 of this deposit, Movie 1 was incorrect: it visualized a different simulation run than intended. This is corrected in Version 2.</p>
Estimated parameters for Bayesian Multilevel Models of KM and kcat values
<p>RData (.rds) files containing brmsfit model objects estimated with the brms R package from KM and kcat values reported in BRENDA and SABIO-RK.</p> <p>These models are used by the ENKIE python package to predict kinetic parameter values and uncertainties.</p>
Data from: Functional traits and community composition: a comparison among community-weighted means, weighted correlations, and multilevel models
Open the record for dataset details and reuse information.
Multilevel modeling of time-series cross-sectional data reveals the dynamic interaction between ecological threats and democratic development
<p>What is the relationship between environment and democracy? The framework of cultural evolution suggests that societal development is an adaptation to ecological threats. Pertinent theories assume that democracy emerges as societies adapt to ecological factors such as higher economic wealth, lower pathogen threats, less demanding climates, and fewer natural disasters. However, previous research confused within-country processes with between-country processes and erroneously interpreted between-country findings as if they generalize to within-country mechanisms. In this article, we analyze a time-series cross-sectional dataset to study the dynamic relationship between environment and democracy (1949-2016), accounting for previous misconceptions in levels of analysis. By separating within-country processes from between-country processes, we find that the relationship between environment and democracy not only differs by countries but also depends on the level of analysis. Economic wealth predicts increasing levels of democracy in between-country comparisons, but within-country comparisons show that democracy declines as countries become wealthier over time. This relationship is only prevalent among historically wealthy countries but not among historically poor countries, whose wealth also increased over time. By contrast, pathogen prevalence predicts lower levels of democracy in both between-country and within-country comparisons. Our longitudinal analyses identifying temporal precedence reveal that not only reductions in pathogen prevalence drive future democracy, but also democracy reduces future pathogen prevalence and increases future wealth. These nuanced results contrast with previous analyses using narrow, cross-sectional data. As a whole, our findings illuminate the dynamic process by which environment and democracy shape each other.</p>
Accompanying simulated data for "Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity"
<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated for the manuscript: "Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity". It comprehends: (1) model outputs (maximum a posteriori estimates) for each repetition (n=100) of each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for 4000 MCMC iterations) for two chains of each repetition (n=3) of each scenario (n=324). Please note that the empirical data used in the manuscript is not available as part of this repository. A subsample of the data used in the empirical example are openly available as an example data set in the R package <a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set is available on request from the authors.</p>
Source data and codes for the paper "Inviting atomic mechanics to macro-continua: A study on monocrystalline Si using a spatial multilevel coarsening model"
<p>Source data and codes for the paper "Inviting atomic mechanics to macro-continua: A study on monocrystalline Si using a spatial multilevel coarsening model"</p> <p>This file includes </p> <p>- Source data for Figs 1-5 and Supplementary Materials</p> <p>- LAMMPS codes and raw log files used to produce the results of this study</p> <p> </p>
Multilevel Modeling of Training Needs in Artificial Intelligence
<p>Nowadays, Artificial Intelligence (AI) is playing a rapidly increasing role in several fields of research and in almost all sectors of real life. However, few studies have assessed the effects of AI applications on training needs. This paper proposes an innovative multilevel modeling in order to investigate Awareness, Attitude and Trust towards AI and their reflections on learning needs. In particular, it is shown how a machine learning variable selection algorithm can support the definition of the optimal subset of all relevant covariates with respect to the outcome variable and improve the multilevel model performance for estimating the probability of educational needs. Thus, starting from a complex web survey to European citizens distributed in eight countries, the estimation of a multilevel binary model, defined on the basis of covariates selected through the Boruta random forest algorithm, is proposed. A discussion on the gender differences of the related estimated multilevel logit models is presented. A sensitivity analysis is also included in order to assess the prediction accuracy of the proposed multilevel logit modeling.</p> <p> </p> <p>This repository contains data generated for the manuscript: " A two-stage procedure for optimal modeling of the probability of training needs in artificial intelligence". It comprehends: (1) the dataset Data_Boruta_Random_Forest used to estimate the variables importance. (2) the dataset Data_Multilevel to perform the comparison among different multilevel binary models proposed in the paper.</p>
Multilevel modeling of time-series cross-sectional data reveals the dynamic interaction between ecological threats and democratic development
Open the record for dataset details and reuse information.
Data from: Competing metabolic strategies in a multilevel selection model
The evolutionary mechanisms of energy efficiency have been addressed. One important question is to understand how the optimized usage of energy can be selected in an evolutionary process, especially when the immediate advantage of gathering efficient individuals in an energetic context is not clear. We propose a model of two competing metabolic strategies differing in their resource usage, an efficient strain which converts resource into energy at high efficiency but displays a low rate of resource consumption, and an inefficient strain which consumes resource at a high rate but at low yield. We explore the dynamics in both well-mixed and structured populations. The selection for optimized energy usage is measured by the likelihood that an efficient strain can invade a population of inefficient strains. It is found that the parameter space at which the efficient strain can thrive in structured populations is always broader than observed in well-mixed populations.
Data from: Improving estimates of environmental change using multilevel regression models of Ellenberg indicator values
Ellenberg indicator values (EIVs) are a widely used metric in plant ecology comprising a semi-quantitative description of species' ecological requirements. Typically, point estimates of mean EIV scores are compared to infer differences in the environmental conditions structuring plant communities – particularly in resurvey studies with no historical environmental data available. However, the use of point estimates as a basis for inference does not take into account variance among species EIVs within sampled plots, and gives equal weighting to means calculated from sites with differing numbers of species. We present a set of multilevel models – fitted with and without group-level predictors – to improve precision and accuracy of site mean EIV scores, and to provide more reliable inference on changing environmental conditions over spatial and temporal gradients in re-visitation studies. We compare multilevel model performance to GLMM's fitted to point estimates of site mean EIVs. We also test the reliability of this method to improve inferences with incomplete species lists in some or all sample sites. Hierarchical modelling led to more accurate and precise estimates of site-level differences in mean EIV scores between time-periods, particularly for datasets with incomplete records of species occurrence. They also revealed directional environmental change within ecological habitat types, which estimates from GLMM's were inadequate to detect. Multilevel models also highlighted a prominent role of hydrological differences as a driver of community change in our case study, which traditional use of EIVs failed to reveal. We have demonstrated that multilevel modelling of EIVs allows for a nuanced estimation of environmental change underlying ecological communities from plant assemblage data, leading to a better understanding of temporal dynamics of ecosystems. Further, the ability of these methods to perform well with missing data should increase the total set of historical data which can be used to this end.
Highlighting the potential of multilevel statistical models for analysis of individual agroforestry systems (companion working R code)
<p>A companion working R code and crop yield dataset that illustrate key concepts presented in a scientific publication titled 'Highlighting the potential of multilevel statistical models for analysis of individual agroforestry systems' published in Agroforestry Systems (July 2023). For the shape file of tree strips, please refer to the publication's supplementary material (adjust the name accordingly). </p>
Healthy Children, Healthy Communities: Effectiveness of a Multilevel Rural Community Engagement Model for Improving Children's Dietary Intake in Family Child Care Homes
ClinicalTrials.gov study NCT07160530. IPD Sharing: YES. Countries: 1. Publications: 0.
Data from: Competing metabolic strategies in a multilevel selection model
Open the record for dataset details and reuse information.
Data from: Improving estimates of environmental change using multilevel regression models of Ellenberg indicator values
Open the record for dataset details and reuse information.
Multilevel Models of Therapeutic Response in the Lungs
ClinicalTrials.gov study NCT02947126. IPD Sharing: NO. Countries: 1. Publications: 0.
An Improved Model Predictive Control for the Virtual Synchronous Generator Based on Modular Multilevel Converters
Open the record for dataset details and reuse information.
test "Inviting atomic mechanics to macro-continua: A study of monocrystalline Si using spatial multilevel coarsening model"
<p>Law source data and LAMMPS codes of</p> <p>"Inviting atomic mechanics to macro-continua: A study of monocrystalline Si using spatial multilevel coarsening model"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.