Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: Tissue culture as a source of replicates in non-model plants: variation in cold response in Arabidopsis lyrata ssp. petraea
Whilst genotype-environment interaction is increasingly receiving attention by ecologists and evolutionary biologists, such studies need genetically homogeneous replicates-a challenging hurdle in outcrossing plants. This could potentially be overcome by using tissue culture techniques. However, plants regenerated from tissue culture may show aberrant phenotypes and "somaclonal" variation. Here we examined the somaclonal variation due to tissue culturing using the response to cold treatment of the photosynthetic efficiency (chlorophyll fluorescence measurements for Fv/Fm, Fv'/Fm' and ΦPSII, representing maximum efficiency of photosynthesis for dark- and light-adapted leaves, and the actual electron transport operating efficiency, respectively, which are reliable indicators of photoinhibition and damage to the photosynthetic electron transport system). We compared this to variation among half-sibling seedlings from three different families of Arabidopsis lyrata ssp. petraea. Somaclonal variation was limited and we could successfully detect within-family variation in change in chlorophyll fluorescence due to cold shock with the help of tissue-culture derived replicates. Icelandic and Norwegian families exhibited higher chlorophyll fluorescence, suggesting higher performance after cold shock, than a Swedish family. Although the main effect of tissue culture on Fv/Fm, Fv'/Fm' and ΦPSII, was small, there were significant interactions between tissue culture and family, suggesting that the effect of tissue culture is genotype-specific. Tissue-cultured plantlets were less affected by cold treatment than seedlings, but to a different extent in each family. These interactive effects, however, were comparable to, or much smaller than the single effect of family. These results suggest that tissue culture is a useful method for obtaining genetically homogenous replicates for studying genotype-environment interaction related to adaptively-relevant phenotypes, such as cold response, in non-model outcrossing plants.
Data from: Whole genome resequencing reveals extensive natural variation in the model green alga Chlamydomonas reinhardtii
We performed whole-genome resequencing of 12 field isolates and eight commonly studied laboratory strains of the model organism Chlamydomonas reinhardtii to characterize genomic diversity and provide a resource for studies of natural variation. Our data support previous observations that Chlamydomonas is among the most diverse eukaryotic species. Nucleotide diversity is ∼3% and is geographically structured in North America with some evidence of admixture among sampling locales. Examination of predicted loss-of-function mutations in field isolates indicates conservation of genes associated with core cellular functions, while genes in large gene families and poorly characterized genes show a greater incidence of major effect mutations. De novo assembly of unmapped reads recovered genes in the field isolates that are absent from the CC-503 assembly. The laboratory reference strains show a genomic pattern of polymorphism consistent with their origin as the recombinant progeny of a diploid zygospore. Large duplications or amplifications are a prominent feature of laboratory strains and appear to have originated under laboratory culture. Extensive natural variation offers a new source of genetic diversity for studies of Chlamydomonas, including naturally occurring alleles that may prove useful in studies of gene function and the dissection of quantitative genetic traits.
Data from: R2s for correlated data: phylogenetic models, LMMs, and GLMMs
Many researchers want to report an R2 to measure the variance explained by a model. When the model includes correlation among data, such as phylogenetic models and mixed models, defining an R2 faces two conceptual problems. (i) It is unclear how to measure the variance explained by predictor (independent) variables when the model contains covariances. (ii) Researchers may want the R2 to include the variance explained by the covariances by asking questions such as "How much of the data is explained by phylogeny?" Here, I investigate three R2s for phylogenetic and mixed models. R2resid is an extension of the ordinary least-squares R2 that weights residuals by variances and covariances estimated by the model; it is closely related to R2glmm presented by Nakagawa and Schielzeth (2013). R2pred is based on predicting each residual from the fitted model and computing the variance between observed and predicted values. R2lik is based on the likelihood of fitted models and therefore reflects the amount of information that the models contain. These three R2s are formulated as partial R2s, making it possible to compare the contributions of predictor variables and variance components (phylogenetic signal and random effects) to the fit of models. Because partial R2s compare a full model with a reduced model without components of the full model, they are distinct from marginal R2s that partition additive components of the variance. The properties of the R2s for phylogenetic models were assessed using simulations for continuous and binary response data (phylogenetic generalized least squares and phylogenetic logistic regression). Because the R2s are designed broadly for any model for correlated data, the R2s were also compared for LMMs and GLMMs. R2resid, R2pred, and R2lik all have similar performance in describing the variance explained by different components of models. However, R2pred gives the most direct answer to the question of how much variance in the data is explained by a model. R2resid is most appropriate for comparing models fit to different datasets, because it does not depend on sample sizes. And R2lik is most appropriate to assess the importance of different components within the same model applied to the same data, because it is most closely associated with statistical significance tests.
Data from: Multi-scale model of regional population decline in little brown bats due to white-nose syndrome
The introduced fungal pathogen Pseudogymnoascus destructans is causing decline of several species of bats in North America, with some even at risk of extinction or extirpation. The severity of the epidemic of white-nose syndrome caused by P. destructans has prompted investigation of the transmission and virulence of infection at multiple scales, but linking these scales is necessary to quantify the mechanisms of transmission and assess population-scale declines. We build a model connecting within-cave disease dynamics of little brown bats to regional scale dispersal, reproduction, and disease spread, including multiple plausible mechanisms of transmission. We parameterize the model using the approach of plausible parameter sets, by comparing stochastic simulation results to statistical probes from empirical data on within-cave prevalence and survival, as well as between-cave spread across a region. Our results are consistent with frequency-dependent transmission between bats, support an important role of environmental transmission, and show very little effect of dispersal among colonies on metapopulation survival. The model also offers a generalizable method to assess hypotheses about cave-to-cave transmission and to identify gaps in knowledge about key processes, and could be expanded to include additional mechanisms or bat species as research on this detrimental fungus progresses.
Data from: Combining landscape genomics and ecological modelling to investigate local adaptation of indigenous Ugandan cattle to East Coast fever
East Coast fever (ECF) is a fatal sickness affecting cattle populations of eastern, central, and southern Africa. The disease is transmitted by the tick Rhipicephalus appendiculatus, and caused by the protozoan Theileria parva parva, which invades host lymphocytes and promotes their clonal expansion. Importantly, indigenous cattle show tolerance to infection in ECF-endemically stable areas. Here, the putative genetic bases underlying ECF-tolerance were investigated using molecular data and epidemiological information from 823 indigenous cattle from Uganda. Vector distribution and host infection risk were estimated over the study area and subsequently tested as triggers of local adaptation by means of landscape genomics analysis. We identified 41 and seven candidate adaptive loci for tick resistance and infection tolerance, respectively. Among the genes associated with the candidate adaptive loci are PRKG1 and SLA2. PRKG1 was already described as associated with tick resistance in indigenous South African cattle, due to its role into inflammatory response. SLA2 is part of the regulatory pathways involved into lymphocytes' proliferation. Additionally, local ancestry analysis suggested the zebuine origin of the genomic region candidate for tick resistance.
Data from: How do foragers decide when to leave a patch? A test of alternative models under natural and experimental conditions
1. A forager's optimal patch-departure time can be predicted by the prescient marginal value theorem (pMVT), which assumes they have perfect knowledge of the environment, or by approaches such as Bayesian-updating and learning rules, which avoid this assumption by allowing foragers to use recent experiences to inform their decisions. 2. In understanding and predicting broader scale ecological patterns, individual-level mechanisms, such as patch-departure decisions, need to be fully elucidated. Unfortunately, there are few empirical studies that compare the performance of patch-departure models that assume perfect knowledge with those that do not, resulting in a limited understanding of how foragers decide when to leave a patch. 3. We tested the patch-departure rules predicted by fixed-rule, pMVT, Bayesian-updating and learning models against one another, using patch residency times recorded from 54 chacma baboons (Papio ursinus) across two groups in natural (n = 6,594 patch visits) and field-experimental (n = 8,569) conditions. 4. We found greater support in the experiment for the model based on Bayesian-updating rules, but greater support for the model based on the pMVT in natural foraging conditions. This suggests that foragers may place more importance on recent experiences in predictable environments, like our experiment, where these experiences provide more reliable information about future opportunities.5. Furthermore, the effect of a single recent foraging experience on patch residency times was uniformly weak across both conditions. This suggests that foragers' perception of their environment may incorporate many previous experiences, thus approximating the perfect knowledge assumed by the pMVT. Foragers may, therefore, optimise their patch-departure decisions in line with the pMVT through the adoption of rules similar to those predicted by Bayesian-updating.
Data from: Trial marriage model – female mate choice under male interference
<ol> <li>In sexually reproducing animals, the process of mate choice by females is often mixed with the process of male-male competition. Current models of female male choice focus mainly on how females identify the higher quality of males, but neglect the effect of male-male competition on the mate choice of females. Therefore, it remains controversial what is the relative importance of two processes in forming a social bond.</li> <li>We propose a new "trial marriage" model for females' mate choice. The model assumes that females unconditionally accept any male they first encounter as their mating partner, and then conditionally switch mates to a new male of higher quality than their current partner when male-male competition occurs. This model was tested in the green weevil, <i><span>Hypomeces squamosus</span></i><span>, by exploring how </span>females switched mates when males' mating interference was experimentally induced.</li> <li>The likelihood that females switched mates, as well as their conditional acceptance criteria of a new mate, was both raised with the intensity of males' mating interference that was manipulated in an enhanced encounter rate experiment, and in male introduction or stepwise removal experiments. These experimental findings confirm that a "trial marriage" strategy occurs during females' mate choice.</li> <li>Compared with other strategies, it is more beneficial for females to choose a better mate without paying the costs of identifying males as suggested by the "trial marriage" strategy. More importantly, by using the current partner quality as the conditional acceptance threshold of new mates, females can choose better males in future encounters with potential mates. In the green weevils, males' preference for larger females and the higher possibility of the largest male winning an interference are mixed together when males' mating interference reaches a higher intensity. Therefore, the consequence of a male interference will determine which male could be chosen by a female. Under this condition, conditional acceptance of the winner becomes the most beneficial strategy of females in choosing their mates. We thus suggest that the "trial marriage" strategy would be more efficient when males' mating interference becomes the determinant factor of females' mate choice.</li> </ol>
Data from: Developing allometric models to predict the individual aboveground biomass of shrubs worldwide
Aim Existing global models to predict standing woody biomass are based on trees characterized by a single principal stem, well-developed in height. However, their use in open woodlands and shrublands, characterized by multistemmed species with substantial crown development, generates a high level of uncertainty in biomass estimates. This limitation led us to i) develop global predictive models of shrub individual aboveground biomass based on simple allometric variables; ii) to compare the fit of these models with existing global biomass models; and iii) to assess whether models fit change when bioclimatic variables are considered. Location Global. Time period Present. Major taxa studied 118 species. Methods We compile a database of 3243 individuals across 49 sites distributed worldwide. Including basal diameter, height and crown diameter as predictor variables, we built potential models and compared their fit using generalized least squares. We used mixed effects models to determine if bioclimatic variables improved the accuracy of biomass models. Results Although the most important variable in terms of predictive capacity was stem basal diameter, crown diameter significantly improved the models fit, followed by height. Four models were finally chosen, with the best model combining all these variables in the same equation (R2 = 0.930, RMSE = 0.476). Selected models performed as well as established global biomass models. Including the individual bioform significantly improved the models fit. Main conclusions Basal diameter, crown diameter and height measures could be combined to provide robust AGB estimates of individual shrub species. Our study supplements well-established models developed for trees, allowing more accurate biomass estimation of multistemmed woody individuals. We further provide tools for a methodological standardization of individual biomass quantification in these species. We expect these results contribute to improve the quality of biomass estimates across ecosystems, but also to generate methodological consensus on field biomass assessments in shrubs.
Data from: Identifying drivers of breeding success in a long-distance migrant using structural equation modelling
In migrant animals, conditions encountered at various times and places throughout their annual cycle may affect breeding success. Yet, most studies so far have only investigated the effect of specific parts of the annual cycle, despite the importance to understand how different stages can interact and how these stages compare to intrinsic quality to properly modulate breeding success. Using a structural equation modelling approach, we investigated drivers of breeding success (migration cycle, individual quality, breeding conditions) in hoopoes (Upupa epops), a long-distant migrant. Our causal framework explained 75% of the variation in breeding success. The effect of the migration schedule was negligible, whereas the previous breeding attempt strongly influenced current breeding success. We suggest that the interplay of individual quality and environmental conditions during both previous and current breeding season may be more important drivers of breeding success than migration schedules, even in a long-distance migrant. We conclude that structural equation modeling is a promising tool to investigate causal relationships. Applied to hoopoes, we demonstrated that current breeding success is strongly linked to previous breeding success. Complementary analysis integrating weather and climate conditions during migration and the breeding season may provide a deeper and wider overview of the annual cycle of hoopoes and additional insights into the existence of carry-over effects in breeding success.
Data from: Ecological niche modelling for conservation planning of an endemic snail in the verge of becoming a pest in cardamom plantations in the Western Ghats biodiversity hotspot
Conservation managers and policy makers are often confronted with a challenging dilemma of devising suitable strategies to maintain agricultural productivity while conserving endemic species that at the early stages of becoming pests of agricultural crops. Identification of environmental factors conducive to species range expansion for forecasting species distribution patterns will play a central role in devising management strategies to minimize the conflict between the agricultural productivity and biodiversity conservation. Here, we present results of a study that predicts the distribution of Indrella ampulla, a snail endemic to the Western Ghats biodiversity hotspot, which is becoming a pest in cardamom (Ellettaria cardamomum) plantations. We determined the distribution patterns and niche overlap between I. ampulla and Ellettaria cardamomum using maximum entropy (MaxEnt) niche modeling techniques under current and future (2020–2080) climatic scenarios. The results showed that climatic (precipitation of coldest quarter and isothermality) and soil (cation exchange capacity of soil [CEC]) parameters are major factors that determine the distribution of I. ampulla in Western Ghats. The model predicted cardamom cultivation areas in southern Western Ghats are highly sensitive to invasion of I. ampulla under both present and future climatic conditions. While the land area in the central Western Ghats is predicted to become unsuitable for I. ampulla and Ellettaria cardamomum in future, we found 71% of the Western Ghats land area is suitable for Ellettaria cardamomum cultivation and 45% suitable for I. ampulla, with an overlap of 35% between two species. The resulting distribution maps are invaluable for policy makers and conservation managers to design and implement management strategies minimizing the conflicts to sustain agricultural productivity while maintaining biodiversity in the region.
Data from: A spatially explicit hierarchical model to characterize population viability
Many of the processes that govern the viability of animal populations vary spatially, yet population viability analyses (PVAs) that account explicitly for spatial variation are rare. We develop a PVA model that incorporates autocorrelation into the analysis of local demographic information to produce spatially explicit estimates of demography and viability at relatively fine spatial scales across a large spatial extent. We use a hierarchical, spatial autoregressive model for capture-recapture data from multiple locations to obtain spatially explicit estimates of adult survival (Φad), juvenile survival (Φjuv), and juvenile-to-adult transition rates (ψ), and a spatial autoregressive model for recruitment data from multiple locations to obtain spatially explicit estimates of recruitment (R). We combine local estimates of demographic rates in stage-structured population models to estimate the rate of population change (λ), then use estimates of λ (and its uncertainty) to forecast changes in local abundance and produce spatially explicit estimates of viability (probability of extirpation, Pex). We apply the model to demographic data for the Sonoran desert tortoise (Gopherus morafkai) collected across its geographic range in Arizona. There was modest spatial variation in λ (0.94–1.03), which reflected spatial variation in Φad (0.85–0.95), Φjuv (0.70–0.89), and ψ (0.07–0.13). Recruitment data were too sparse for spatially explicit estimates, therefore we used a range-wide estimate (R = 0.32 one-year old females per female per year). Spatial patterns in demographic rates were complex, but Φad, Φjuv, and λ tended to be lower and ψ higher in the northwestern portion of the range. Spatial patterns in Pex varied with local abundance. For local abundances > 500, Pex was near zero (Pex approached one in the northwestern portion of the range and remained low elsewhere. When local abundances were Pex > 0.25). This approach to PVA offers the potential to reveal spatial patterns in demography and viability that can inform conservation and management at multiple spatial scales, provide insight into scale-related investigations in population ecology, and improve basic ecological knowledge of landscape-level phenomena.
Data from: Cladogenetic and anagenetic models of chromosome number evolution: a Bayesian model averaging approach
Chromosome number is a key feature of the higher-order organization of the genome, and changes in chromosome number play a fundamental role in evolution. Dysploid gains and losses in chromosome number, as well as polyploidization events, may drive reproductive isolation and lineage diversification. The recent development of probabilistic models of chromosome number evolution in the groundbreaking work by Mayrose et al. (2010, ChromEvol) have enabled the inference of ancestral chromosome numbers over molecular phylogenies and generated new interest in studying the role of chromosome changes in evolution. However, the ChromEvol approach assumes all changes occur anagenetically (along branches), and does not model events that are specifically cladogenetic. Cladogenetic changes may be expected if chromosome changes result in reproductive isolation. Here we present a new class of models of chromosome number evolution (called ChromoSSE) that incorporate both anagenetic and cladogenetic change. The ChromoSSE models allow us to determine the mode of chromosome number evolution; is chromosome evolution occurring primarily within lineages, primarily at lineage splitting, or in clade-specific combinations of both? Furthermore, we can estimate the location and timing of possible chromosome speciation events over the phylogeny. We implemented ChromoSSE in a Bayesian statistical framework, specifically in the software RevBayes, to accommodate uncertainty in parameter estimates while leveraging the full power of likelihood based methods. We tested ChromoSSE's accuracy with simulations and re-examined chromosomal evolution in Aristolochia, Carex section Spirostachyae, Helianthus, Mimulus sensu lato (s.l.), and Primula section Aleuritia, finding evidence for clade-specific combinations of anagenetic and cladogenetic dysploid and polyploid modes of chromosome evolution.
Simulation data for WRF-GC (v2.0): online two-way coupling of WRF (v3.9.1.1) and GEOS-Chem (v12.7.2) for modeling regional atmospheric chemistry–meteorology interactions
<p>This repository provides the test simulation data for "WRF-GC (v2.0): online two-way coupling of WRF (v3.9.1.1) and GEOS-Chem (v12.7.2) for modeling regional atmospheric chemistry–meteorology interactions" published in Geoscientific Model Development. The configurations for sensitivity experiments are described in this paper. Please contact the corresponding author Tzung-May Fu (fuzm@sustech.edu.cn) for more details.</p>
Data from: Demographic modelling reveals a history of divergence with gene flow for a glacially tied stonefly in a changing post-Pleistocene landscape
Aim: Climate warming is causing extensive loss of glaciers in mountainous regions, yet our understanding of how glacial recession influences evolutionary processes and genetic diversity is limited. Linking genetic structure with the influences shaping it can improve understanding of how species respond to environmental change. Here, we used genome-scale data and demographic modelling to resolve the evolutionary history of Lednia tumana, a rare, aquatic insect endemic to alpine streams. We also employed a range of widely used data filtering approaches to quantify how they influenced population structure results. Location: Alpine streams in the Rocky Mountains of Glacier National Park, Montana, USA. Taxon: Lednia tumana, a stonefly (Order Plecoptera) in the family Nemouridae. Methods: We generated single nucleotide polymorphism data through restriction-site associated DNA sequencing to assess contemporary patterns of genetic structure for 11 L. tumana populations. Using identified clusters, we assessed demographic history through model selection and parameter estimation in a coalescent framework. During population structure analyses, we filtered our data to assess the influence of singletons, missing data and total number of markers on results. Results: Contemporary patterns of population structure indicate that L. tumana exhibits a pattern of isolation-by-distance among populations within three genetic clusters that align with geography. Mean pairwise genetic differentiation (FST) among populations was 0.033. Coalescent-based demographic modelling supported divergence with gene flow among genetic clusters since the end of the Pleistocene (~13-17 kya), likely reflecting the south-to-north recession of ice sheets that accumulated during the Wisconsin glaciation. Main conclusions: We identified a link between glacial retreat, evolutionary history and patterns of genetic diversity for a range-restricted stonefly imperiled by climate change. This finding included a history of divergence with gene flow, an unexpected conclusion for a mountaintop species. Beyond L. tumana, this study demonstrates the complexity of assessing genetic structure for weakly differentiated species, shows the degree to which rare alleles and missing data may influence results, and highlights the usefulness of genome-scale data to extend population genetic inquiry in non-model species.
Data from: A refined modelling approach to assess the influence of sampling on palaeobiodiversity curves: new support for declining Cretaceous dinosaur richness
Modelling has been underdeveloped with respect to constructing palaeobiodiversity curves, but it offers an additional tool for removing sampling from their estimation. Here an alternative to subsampling approaches, which often require large sample sizes, is explored by the extension and refinement of a pre-existing modelling technique that uses a geological proxy for sampling. Application of the model to the three main clades of dinosaurs suggests that much of their diversity fluctuations cannot be explained by sampling alone. Furthermore, there is new support for a long-term decline in their diversity leading up to the K/Pg extinction event. At present use of this method with data that includes either lagerstätten or 'Pull of the Recent' biases is inappropriate, although partial solutions are offered.
Data from: Detection error influences both temporal seroprevalence predictions and risk factors associations in wildlife disease models
Understanding the prevalence of pathogens in invasive species is essential to guide efforts to prevent transmission to agricultural animals, wildlife, and humans. Pathogen prevalence can be difficult to estimate for wild species due to imperfect sampling and testing (pathogens may not be detected in infected individuals and erroneously detected in individuals that are not infected). The invasive wild pig (Sus scrofa, also referred to as wild boar and feral swine) is one of the most widespread hosts of domestic animal and human pathogens in North America. We developed hierarchical Bayesian models that account for imperfect detection to estimate the seroprevalence of five pathogens (porcine reproductive and respiratory syndrome virus, pseudorabies virus, Influenza A virus in swine, Hepatitis E virus, and Brucella spp.) in wild pigs in the United States using a dataset of over 50,000 samples across nine years. To assess the effect of incorporating detection error in models, we also evaluated models that ignored detection error. Both sets of models included effects of demographic parameters on seroprevalence. We compared our predictions of seroprevalence to 40 published studies, only one of which accounted for imperfect detection. We found a range of seroprevalence among the pathogens with a high seroprevalence of pseudorabies virus, indicating significant risk to livestock and wildlife. Demographics had mostly weak effects, indicating that other variables may have greater effects in predicting seroprevalence. Models that ignored detection error led to different predictions of seroprevalence as well as different inferences on the effects of demographic parameters. Our results highlight the importance of incorporating detection error in models of seroprevalence and demonstrate that ignoring such error may lead to erroneous conclusions about the risk associated with pathogen transmission. When using opportunistic sampling data to model seroprevalence and evaluate risk factors, detection error should be included.
Data from: Assessing the expected response to genomic selection of individuals and families in Eucalyptus breeding with an additive-dominant model
We report a genomic selection (GS) study of growth and wood quality traits in an outbred F2 hybrid Eucalyptus population (n=768) using high-density single-nucleotide polymorphism (SNP) genotyping. Going beyond previous reports in forest trees, models were developed for different selection targets, namely, families, individuals within families and individuals across the entire population using a genomic model including dominance. To provide a more breeder-intelligible assessment of the performance of GS we calculated the expected response as the percentage gain over the population average expected genetic value (EGV) for different proportions of genomically selected individuals, using a rigorous cross-validation (CV) scheme that removed relatedness between training and validation sets. Predictive abilities (PAs) were 0.40–0.57 for individual selection and 0.56–0.75 for family selection. PAs under an additive+dominance model improved predictions by 5 to 14% for growth depending on the selection target, but no improvement was seen for wood traits. The good performance of GS with no relatedness in CV suggested that our average SNP density (~25 kb) captured some short-range linkage disequilibrium. Truncation GS successfully selected individuals with an average EGV significantly higher than the population average. Response to GS on a per year basis was ~100% more efficient than by phenotypic selection and more so with higher selection intensities. These results contribute further experimental data supporting the positive prospects of GS in forest trees. Because generation times are long, traits are complex and costs of DNA genotyping are plummeting, genomic prediction has good perspectives of adoption in tree breeding practice.
Data from: Novel application of explicit dynamics occupancy models to ongoing aquatic invasions
1. Identification of suitable habitat where invasive species can establish is an important step towards controlling their spread. Accurate identification is difficult for new or slow invaders because unoccupied habitats may be suitable given enough time for dispersal and occupied habitats may prove to be unsuitable for establishment. 2. To identify suitable habitat of a recent invader, I used an explicit dynamics occupancy modeling framework to evaluate habitat covariates related to successful and failed establishments of American bullfrogs (Lithobates catesbeianus) within the Yellowstone River floodplain of Montana, USA from 2012-2016. 3. During this 5-year period, bullfrogs failed to establish at most sites they colonized. Bullfrog establishment was most likely to occur and least likely to fail at sites closest to human-modified ponds and lakes and those with emergent vegetation. These habitat covariates were generally associated with the presence of permanent water. 4. Suitable habitat for bullfrog establishment is abundant in the Yellowstone River floodplain, though many sites with suitable habitat remain uncolonized. Thus, the maximum distribution of bullfrogs is much greater than their current distribution. 5. Synthesis and applications. Focused control efforts on habitats with or proximate to permanent waters are most likely to reduce the potential for bullfrog establishment and spread. The novel application of explicit dynamics occupancy models is a useful and widely applicable tool for guiding management efforts towards those habitats where new or slow invaders are most likely to establish and persist.07-Aug-2017
Data from: Analysis of animal accelerometer data using hidden Markov models
Use of accelerometers is now widespread within animal biologging as they provide a means of measuring an animal's activity in a meaningful and quantitative way where direct observation is not possible. In sequential acceleration data, there is a natural dependence between observations of behaviour, a fact that has been largely ignored in most analyses. Analyses of acceleration data where serial dependence has been explicitly modelled have largely relied on hidden Markov models (HMMs). Depending on the aim of an analysis, an HMM can be used for state prediction or to make inferences about drivers of behaviour. For state prediction, a supervised learning approach can be applied. That is, an HMM is trained to classify unlabelled acceleration data into a finite set of pre-specified categories. An unsupervised learning approach can be used to infer new aspects of animal behaviour when biologically meaningful response variables are used, with the caveat that the states may not map to specific behaviours. We provide the details necessary to implement and assess an HMM in both the supervised and unsupervised learning context and discuss the data requirements of each case. We outline two applications to marine and aerial systems (shark and eagle) taking the unsupervised learning approach, which is more readily applicable to animal activity measured in the field. HMMs were used to infer the effects of temporal, atmospheric and tidal inputs on animal behaviour. Animal accelerometer data allow ecologists to identify important correlates and drivers of animal activity (and hence behaviour). The HMM framework is well suited to deal with the main features commonly observed in accelerometer data and can easily be extended to suit a wide range of types of animal activity data. The ability to combine direct observations of animal activity with statistical models, which account for the features of accelerometer data, offers a new way to quantify animal behaviour and energetic expenditure and to deepen our insights into individual behaviour as a constituent of populations and ecosystems.
Data from: How does spatial resolution affect model performance? A case for ensemble approaches for marine benthic mesophotic communities
Aim: To investigate how changing grid size can alter model predictions of the distribution of mesophotic taxa and how it affects different modelling methods. Location: Ningaloo Marine Park, Western Australia. Taxon: Benthic mesophotic taxa: corals, macroalgae, and sponges. Methods: We determined the distributions of the major benthic taxonomic groups: corals, macroalgae, and sponges, using a number of modelling techniques and an ensemble using the 'sdm' R package. A range of grid sizes were used (10 m, 50 m, 100 m, and 250 m) to identify how model predictions were altered. Models were evaluated using the area under the curve of a receiver operator characteristic plot (AUC) and the true skill statistic (TSS) using a spatially independent dataset. Results: Grid size had a large effect on model performance across the taxonomic groups. Model outputs were compared to null surfaces and 88.8% of models performed significantly better than null. Distribution of corals was best predicted using the finest grid size (10 m) regardless of modelling method, although a model ensemble produced the best results (AUC = 0.80, TSS = 0.52). Macroalgae and sponges were better predicted at coaster grids sizes (250 m). Again, ensembles performed well for both macroalgae (AUC = 0.83, TSS = 0.63) and sponges (AUC = 0.88, TSS = 0.66). Model ensembles maintained high accuracy across grid sizes and were consistently the best, or second-best, performing method. Main Conclusions: This study has shown how grid size should be considered when producing distribution models. Identifying the most relevant grid size and being aware of the influence it may have will provide more accurate predictions of the distributions of taxa. Ensemble methods maintained good performance across scenarios and thus provide a useful tool for conservation and management especially where single modelling methods showed high levels of variability.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.