Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Data from: A hierarchical distance sampling model to estimate abundance and covariate associations of species and communities
Distance sampling is a common survey method in wildlife studies, because it allows accounting for imperfect detection. The framework has been extended to hierarchical distance sampling (HDS), which accommodates the modelling of abundance as a function of covariates, but rare and elusive species may not yield enough observations to fit such a model. We integrate HDS into a community modelling framework that accommodates multi-species spatially replicated distance sampling data. The model allows species-specific parameters, but these come from a common underlying distribution. This form of information sharing enables estimation of parameters for species with sparse data sets that would otherwise be discarded from analysis. We evaluate the performance of the model under varying community sizes with different species-specific abundances through a simulation study. We further fit the model to a seabird data set obtained from shipboard distance sampling surveys off the East Coast of the USA. Comparing communities comprised of 5, 15 or 30 species, bias of all community-level parameters and some species-level parameters decreased with increasing community size, while precision increased. Most species-level parameters were less biased for more abundant species. For larger communities, the community model increased precision in abundance estimates of rarely observed species when compared to single-species models. For the seabird application, we found a strong negative association of community and species abundance with distance to shore. Water temperature and prey density had weak effects on seabird abundance. Patterns in overall abundance were consistent with known seabird ecology. The community distance sampling model can be expanded to account for imperfect availability, imperfect species identification or other missing individual covariates. The model allowed us to make inference about ecology of species communities, including rarely observed species, which is particularly important in conservation and management. The approach holds great potential to improve inference on species communities that can be surveyed with distance sampling.
Data from: The sampling and estimation of marine paleodiversity patterns: implications of a Pliocene model
Data that accurately capture the spatial structure of biodiversity are required for many paleobiological questions, from assessments of changing provinciality and the role of geographic ranges in extinction and originations, to estimates of global taxonomic or morphological diversity through time. Studies of temporal changes in diversity and global biogeographic patterns have attempted to overcome fossil sampling biases through sampling standardization protocols, but such approaches must ultimately be limited by available literature and museum collections. One approach to evaluating such limits is to compare results from the fossil record with models of past diversity patterns informed by modern relationships between diversity and climatic factors. Here we use present-day patterns for marine bivalves, combined with data on the geologic ages and distributions of extant taxa, to develop a model for Pliocene diversity patterns, which is then compared with diversity patterns retrieved from the literature as compiled by the Paleobiology Database (PaleoDB). The published Pliocene bivalve data (PaleoDB) lack the first-order spatial structure required to generate the modern biogeography within the time available (<3 Myr). Instead, the published data (raw and standardized) show global diversity maxima in the Tropical West Atlantic, followed closely by a peak in the cool-temperate East Atlantic. Either today's tropical West Pacific diversity peak, double that of any other tropical region, is a purely Pleistocene phenomenon—highly unlikely given the geologic ages of extant genera and the topology of molecular phylogenies—or the paleontological literature is such a distorted sample of tropical Pliocene diversity that current sampling standardization methods cannot compensate for existing biases. A rigorous understanding of large-scale spatial and temporal diversity patterns will require new approaches that can compensate for such strong bias, presumably by drawing more fully on our understanding of the factors that underlie the deployment of diversity today.
Data from: Evaluating noninvasive genetic sampling techniques to estimate large carnivore abundance
Monitoring large carnivores is difficult because of intrinsically low densities and can be dangerous if physical capture is required. Noninvasive genetic sampling (NGS) is a safe and cost-effective alternative to physical capture. We evaluated the utility of two NGS methods (scat detection dogs and hair sampling) to obtain genetic samples for abundance estimation of coyotes, black bears and Canada lynx in three areas of Newfoundland, Canada. We calculated abundance estimates using program capwire, compared sampling costs, and the cost/sample for each method relative to species and study site, and performed simulations to determine the sampling intensity necessary to achieve abundance estimates with coefficients of variation (CV) of <10%. Scat sampling was effective for both coyotes and bears and hair snags effectively sampled bears in two of three study sites. Rub pads were ineffective in sampling coyotes and lynx. The precision of abundance estimates was dependent upon the number of captures/individual. Our simulations suggested that ~3.4 captures/individual will result in a < 10% CV for abundance estimates when populations are small (23–39), but fewer captures/individual may be sufficient for larger populations. We found scat sampling was more cost-effective for sampling multiple species, but suggest that hair sampling may be less expensive at study sites with limited road access for bears. Given the dependence of sampling scheme on species and study site, the optimal sampling scheme is likely to be study-specific warranting pilot studies in most circumstances.
Data from: DNA barcoding meets molecular scatology: short mtDNA sequences for standardized species assignment of carnivore noninvasive samples
Although species assignment of scats is important to study carnivoran biology, there is still no standardized assay for the identification of carnivores worldwide, which would allow large-scale routine assessments and reliable cross-comparison of results. Here we evaluate the potential of two short mtDNA fragments (ATP6 [126 bp] and COI [187 bp]) to serve as standard markers for the Carnivora. Samples of 66 species were sequenced for one or both of these segments. Alignments were complemented with archival sequences, and analyzed with three approaches (tree-based, distance-based and character-based). Intraspecific genetic distances were generally lower than between-species distances, resulting in diagnosable clusters for 86% (ATP6) and and 85% (COI) of the species. Notable exceptions were recently diverged species, most of which could still be identified using diagnostic characters, uniqueness of haplotypes, or by reducing the geographic scope of the comparison. In silico comparative analyses were also performed with a 110-bp cytochrome b (cytb) segment, whose identification success was lower (70%), possibly due to the smaller number of informative sites and/or the influence of misidentified sequences obtained from GenBank. Finally, we performed case-studies with faecal samples, which supported the suitability of our two focal markers for poor-quality DNA, and allowed an assessment of prey-DNA co-amplification. No evidence of prey DNA contamination was found for ATP6, while some cases were observed for COI and subsequently eliminated by the design of more specific primers. Overall, our results indicate that these segments hold good potential as standard markers for accurate species-level identification in the Carnivora.
Data from: Soil sampling and isolation of extracellular DNA from large amount of starting material suitable for metabarcoding studies
DNA metabarcoding corresponds to the DNA-based identification of multiple species from a single complex and degraded environmental sample. We developed new sampling and extraction protocols suitable for DNA metabarcoding analyses, targeting soil extracellular DNA. The proposed sampling protocol has been designed to reduce as much as possible the influence of the local heterogeneity by processing large amount of soil, resulting from the mixing of many different cores. The DNA extraction is based on the use of saturated phosphate buffer. The sampling and extraction protocols were validated first by analyzing plant DNA from a set of 12 plots corresponding to four plant communities in alpine meadows, and second by conducting pilot experiments on fungi and earthworms. The results of the validation experiments clearly demonstrated that sound biological information can be retrieved when following these sampling and extraction procedures. Such a protocol can be implemented at any time of the year without any preliminary knowledge of specific types of organisms during the sampling. It offers the opportunity to analyze all groups of organisms using a single sampling/extraction procedure and opens the possibility to fully standardize biodiversity surveys.
Data from: Model selection with overdispersed distance sampling data
1. Distance sampling (DS) is a widely-used framework for estimating animal abundance. DS models assume that observations of distances to animals are independent. Non-independent observations introduce overdispersion, causing model selection criteria such as AIC or AICc to favour overly complex models, with adverse effects on accuracy and precision. 2. We describe, and evaluate via simulation and with real data, estimators of an overdispersion factor (c ̂), and associated adjusted model selection criteria (QAIC) for use with overdispersed DS data. In other contexts, a single value of c ̂ is calculated from the "global" model, i.e., the most highly-parameterized model in the candidate set, and used to calculate QAIC for all models in the set; the resulting QAIC values, and associated ΔQAIC values and QAIC weights, are comparable across the entire set. Candidate models of the DS detection function include models with different general forms (e.g., half-normal, hazard rate, uniform), so it may not be possible to identify a single global model. We therefore propose a two-step model selection procedure by which QAIC is used to select among models with the same general form, and then a goodness-of-fit statistic is used to select among models with different forms. A drawback of this approach is that QAIC values are not comparable across all models in the candidate set. 3. Relative to AIC, QAIC and the two-step model selection procedure avoided overfitting and improved the accuracy and precision of densities estimated from simulated data. When applied to six real data sets, adjusted criteria and procedures selected either the same model as AIC or a model that yielded a more accurate density estimate in 5 cases, and a model that yielded a less accurate estimate in 1 case. 4. Many DS surveys yield overdispersed data, including cue counting surveys of songbirds and cetaceans, surveys of social species including primates, and camera-trapping surveys. Methods that adjust for overdispersion during the model selection stage of DS analyses therefore address a conspicuous gap in the DS analytical framework as applied to species of conservation concern.
Data from: How many dinosaur species were there? Fossil bias and true richness estimated using a Poisson sampling model
The fossil record is a rich source of information about biological diversity in the past. However, the fossil record is not only incomplete but has also inherent biases due to geological, physical, chemical and biological factors. Our knowledge of past life is also biased because of differences in academic and amateur interests and sampling efforts. As a result, not all individuals or species that lived in the past are equally likely to be discovered at any point in time or space. To reconstruct temporal dynamics of diversity using the fossil record, biased sampling must be explicitly taken into account. Here, we introduce an approach that uses the variation in the number of times each species is observed in the fossil record to estimate both sampling bias and true richness. We term our technique TRiPS (True Richness estimated using a Poisson Sampling model) and explore its robustness to violation of its assumptions via simulations. We then venture to estimate sampling bias and absolute species richness of dinosaurs in the geological stages of the Mesozoic. Using TRiPS, we estimate that 1936 (1543–2468) species of dinosaurs roamed the Earth during the Mesozoic. We also present improved estimates of species richness trajectories of the three major dinosaur clades: the sauropodomorphs, ornithischians and theropods, casting doubt on the Jurassic–Cretaceous extinction event and demonstrating that all dinosaur groups are subject to considerable sampling bias throughout the Mesozoic.
Data from: Evolutionary genomics of gypsy moth populations sampled along a latitudinal gradient
The European gypsy moth (Lymantria dispar L.) was first introduced to Massachusetts in 1869 and within 150 years has spread throughout eastern North America. This large-scale invasion across a heterogeneous landscape allows examination of the genetic signatures of adaptation potentially associated with rapid geographic spread. We tested the hypothesis that spatially divergent natural selection has driven observed changes in three developmental traits that were measured in a common garden for 165 adult moths sampled from six populations across a latitudinal gradient covering the entirety of the range. We generated genotype data for 91,468 single nucleotide polymorphisms (SNPs) based on double digest restriction-site associated DNA sequencing (ddRADseq) and used these data to discover genome-wide associations for each trait, as well as to test for signatures of selection on the discovered architectures. Genetic structure across the introduced range of gypsy moth was small in magnitude (FST = 0.069), with signatures of bottlenecks and spatial expansion apparent in the rare portion of the allele frequency spectrum. Results from applications of Bayesian sparse linear mixed models were consistent with the presumed polygenic architectures of each trait. Further analyses were indicative of spatially divergent natural selection acting on larval development time and pupal mass, with the linkage disequilibrium like component of this test acting as the main driver of observed patterns. The populations most important for these signals were two range-edge populations established less than 30 generations ago. We discuss the importance of rapid polygenic adaptation to the ability of non-native species to invade novel environments.
Data from: Evidence of neutral and adaptive genetic divergence between European trout populations sampled along altitudinal gradients
Species with a wide geographical distribution are often composed of distinct subgroups which may be adapted to their local environment. European trout (Salmo trutta species complex) provide an example of such a complex consisting of several genetically and ecologically distinct forms. However, trout populations are strongly influenced by human activities, and it is unclear to what extent neutral and adaptive genetic differences have persisted. We sampled 30 Swiss trout populations from heterogeneous environments along replicated altitudinal gradients in three major European drainages. More than 850 individuals were genotyped at 18 microsatellite loci which included loci diagnostic for evolutionary lineages and candidate markers associated with temperature tolerance, reproductive timing and immune defence. We find that the phylogeographic structure of Swiss trout populations has not been completely erased by stocking. Distinct genetic clusters corresponding to the different drainages could be identified, although nonindigenous alleles were clearly present, especially in the two Mediterranean drainages. We also still detected neutral genetic differentiation within rivers which was often associated with the geographical distance between populations. Five loci showed evidence of divergent selection between populations with several drainage-specific patterns. Lineage-diagnostic markers, a marker linked to a quantitative trait locus for upper temperature tolerance in other salmonids and a marker linked to the major histocompatibility class I gene were implicated in local adaptation and some patterns were associated with altitude. In contrast, tentative evidence suggests a signal of balancing selection at a second immune relevant gene (TAP2). Our results confirm the persistence of both neutral and potentially adaptive genetic differences between trout populations in the face of massive human-mediated dispersal.
Data from: Inference of genetic architecture from chromosome partitioning analyses is sensitive to genome variation, sample size, heritability and effect size distribution
Genomewide association studies have contributed immensely to our understanding of the genetic basis of complex traits. One major conclusion arising from these studies is that most traits are controlled by many loci of small effect, confirming the infinitesimal model of quantitative genetics. A popular approach to test for polygenic architecture involves so‐called "chromosome partitioning" where phenotypic variance explained by each chromosome is regressed on the size of the chromosome. First developed for humans, this has now been repeatedly used in other species, but there has been no evaluation of the suitability of this method in species that can differ in their genome characteristics such as number and size of chromosomes. Nor has the influence of sample size, heritability of the trait, effect size distribution of loci controlling the trait or the physical distribution of the causal loci in the genome been examined. Using simulated data, we show that these characteristics have major influence on the inferences of the genetic architecture of traits we can infer using chromosome partitioning analyses. In particular, small variation in chromosome size, small sample size, low heritability, a skewed effect size distribution and clustering of loci can lead to a loss of power and consequently altered inference from chromosome partitioning analyses. Future studies employing this approach need to consider and derive an appropriate null model for their study system, taking these parameters into consideration. Our simulation results can provide some guidelines on these matters, but further studies examining a broader parameter space are needed.
Data from: Mimicry among unequally defended prey should be mutualistic when predators sample optimally
Understanding the conditions under which moderately defended prey evolve to resemble better-defended prey and whether this mimicry is parasitic (quasi-Batesian) or mutualistic (Müllerian) is central to our understanding of warning signals. Models of predator learning generally predict quasi-Batesian relationships. However, predators' attack decisions are based not only on learning alone but also on the potential future rewards. We identify the optimal sampling strategy of predators capable of classifying prey into different profitability categories and contrast the implications of these rules for mimicry evolution with a classical Pavlovian model based on conditioning. In both cases, the presence of moderately unprofitable mimics causes an increase in overall consumption. However, in the case of the optimal sampling strategy, this increase in consumption is typically outweighed by the increase in overall density of prey sharing the model appearance (a dilution effect), causing a decrease in mortality. It suggests that if predators forage efficiently to maximize their long-term payoff, genuine quasi-Batesian mimicry should be rare, which may explain the scarcity of evidence for it in nature. Nevertheless, we show that when moderately defended mimics are profitable to attack by hungry predators, then they can be parasitic on their models, just as classical Batesian mimics are.
Data from: Aβ42 fibril formation from predominantly oligomeric samples suggests a link between oligomer heterogeneity and fibril polymorphism
Aβ oligomers play a central role in the pathogenesis of Alzheimer's disease. Oligomers of different sizes, morphology, and structures have been reported in both in vivo and in vitro studies, but there is a general lack of understanding about where to place these oligomers in the overall process of Aβ aggregation and fibrillization. Here we show that Aβ42 spontaneously forms oligomers with a wide range of sizes in the same sample. These Aβ42 samples contain predominantly oligomers, and they quickly form fibrils upon incubation at 37°C. When fractionated using ultrafiltration filters, the samples enriched with smaller oligomers form fibrils at a faster rate than the samples enriched with larger oligomers, with both a shorter lag time and faster fibril growth rate. This observation is independent of Aβ42 batches and HFIP treatment. Furthermore, the fibrils formed by the samples enriched with larger oligomers are more readily solubilized by EGCG, a main catechin component of green tea. These results suggest that the fibrils formed by larger oligomers may adopt a different structure from fibrils formed by smaller oligomers, pointing to a link between oligomer heterogeneity and fibril polymorphism.
Data from: Scalable, semi-automated fluorescence reduction neutralization assay for qualitative assessment of Ebola virus-neutralizing antibodies in human clinical samples
Antibody titers against a viral pathogen are typically measured using an antigen binding assay, such as an enzyme-linked immunosorbent assay (ELISA), which only measures the ability of antibodies to identify a viral antigen of interest. Neutralization assays measure the presence of virus-neutralizing antibodies in a sample. Traditional neutralization assays, such as the plaque reduction neutralization test (PRNT), are often difficult to use on a large scale due to being both labor and resource intensive. Here we describe an Ebola virus fluorescence reduction neutralization assay (FRNA), which tests for neutralizing antibodies, that requires only a small volume of sample in a 96-well format and is easy to automate. The readout of the FRNA is the percentage of Ebola virus-infected cells measured with an optical reader or overall chemiluminescence that can be generated by multiple reading platforms and the readout is compatible with lytic and non-lytic viruses. Using blinded human clinical samples (EVD survivors or contacts) obtained in Liberia during the 2013–2016 Ebola virus disease outbreak, we demonstrate that FRNA-measured antibody titers are highly correlated with those measured by the Filovirus Animal Non-clinical Group (FANG) ELISA - the current standard for anti-EBOV antibody measurement with the important distinction of providing information on the neutralizing capabilities of the antibodies.
Data from: High-throughput adaptive sampling for whole-slide histopathology image analysis (HASHI) via convolutional neural networks: application to invasive breast cancer detection
Precise detection of invasive cancer on whole-slide images (WSI) is a critical first step in digital pathology tasks of diagnosis and grading. Convolutional neural network (CNN) is the most popular representation learning method for computer vision tasks, which have been successfully applied in digital pathology, including tumor and mitosis detection. However, CNNs are typically only tenable with relatively small image sizes (200x200 pixels). Only recently, Fully convolutional networks (FCN) are able to deal with larger image sizes (500x500 pixels) for semantic segmentation. Hence, the direct application of CNNs to WSI is not computationally feasible because for a WSI, a CNN would require billions or trillions of parameters. To alleviate this issue, this paper presents a novel method, High-throughput Adaptive Sampling for whole-slide Histopathology Image analysis (HASHI), which involves: i) a new efficient adaptive sampling method based on probability gradient and quasi-Monte Carlo sampling, and, ii) a powerful representation learning classifier based on CNNs. We applied HASHI to automated detection of invasive breast cancer on WSI. HASHI was trained and validated using three different data cohorts involving near 500 cases and then independently tested on 195 studies from The Cancer Genome Atlas. The results show that (1) the adaptive sampling method is an effective strategy to deal with WSI without compromising prediction accuracy by obtaining comparative results of a dense sampling (~6 million of samples in 24 hours) with far fewer samples (~2,000 samples in 1 minute), and (2) on an independent test dataset, HASHI is effective and robust to data from multiple sites, scanners, and platforms, achieving an average Dice coefficient of 76%.
Data from: The relative power of genome scans to detect local adaptation depends on sampling design and statistical method
Although genome scans have become a popular approach towards understanding the genetic basis of local adaptation, the field still does not have a firm grasp on how sampling design and demographic history affect the performance of genome scans on complex landscapes. To explore these issues, we compared 20 different sampling designs in equilibrium (i.e. island model and isolation by distance) and nonequilibrium (i.e. range expansion from one or two refugia) demographic histories in spatially heterogeneous environments. We simulated spatially complex landscapes, which allowed us to exploit local maxima and minima in the environment in 'pair' and 'transect' sampling strategies. We compared FST outlier and genetic–environment association (GEA) methods for each of two approaches that control for population structure: with a covariance matrix or with latent factors. We show that while the relative power of two methods in the same category (FST or GEA) depended largely on the number of individuals sampled, overall GEA tests had higher power in the island model and FST had higher power under isolation by distance. In the refugia models, however, these methods varied in their power to detect local adaptation at weakly selected loci. At weakly selected loci, paired sampling designs had equal or higher power than transect or random designs to detect local adaptation. Our results can inform sampling designs for studies of local adaptation and have important implications for the interpretation of genome scans based on landscape data.
Data from: Model selection and parameter inference in phylogenetics using nested sampling
Bayesian inference methods rely on numerical algorithms for both model selection and parameter inference. In general, these algorithms require a high computational effort to yield reliable estimates. One of the major challenges in phylogenetics is the estimation of the marginal likelihood. This quantity is commonly used for comparing different evolutionary models, but its calculation, even for simple models, incurs high computational cost. Another interesting challenge relates to the estimation of the posterior distribution. Often, long Markov chains are required to get sufficient samples to carry out parameter inference, especially for tree distributions. In general, these problems are addressed separately by using different procedures. Nested sampling (NS) is a Bayesian computation algorithm which provides the means to estimate marginal likelihoods together with their uncertainties, and to sample from the posterior distribution at no extra cost. The methods currently used in phylogenetics for marginal likelihood estimation lack in practicality due to their dependence on many tuning parameters and their inability of most implementations to provide a direct way to calculate the uncertainties associated with the estimates, unlike NS. In this paper, we introduce NS to phylogenetics. Its performance is analysed under different scenarios and compared to established methods. We conclude that NS is a competitive and attractive algorithm for phylogenetic inference. An implementation is available as a package for BEAST 2 under the LGPL licence, accessible at https://github.com/BEAST2-Dev/nested-sampling.
Data from: The efficacy of consensus tree methods for summarising phylogenetic relationships from a posterior sample of trees estimated from morphological data
Consensus trees are required to summarise trees obtained through MCMC sampling of a posterior distribution, providing an overview of the distribution of estimated parameters such as topology, branch lengths and divergence times. Numerous consensus tree construction methods are available, each presenting a different interpretation of the tree sample. The rise of morphological clock and sampled-ancestor methods of divergence time estimation, in which times and topology are co-estimated, has increased the popularity of the maximum clade credibility (MCC) consensus tree method. The MCC method assumes that the sampled, fully resolved topology with the highest clade credibility contains an adequate summary of the most probable clades, with parameter estimates from compatible sampled trees used to obtain the marginal distributions of parameters such as clade ages and branch lengths. Using both simulated and empirical data, we demonstrate that MCC trees, and trees constructed using the similar maximum a posteriori (MAP) method, often include poorly supported and incorrect clades when summarising diffuse posterior samples of trees. We demonstrate that the paucity of information in morphological datasets contributes to the inability of MCC and MAP trees to present an accurate summary of the posterior distribution. Conversely, majority-rule consensus (MRC) trees report a lower proportion of incorrect nodes when summarising the same posterior samples of trees. Thus, we advocate the use of MRC trees, in place of MCC or MAP trees, in attempts to summarise the results of Bayesian phylogenetic analyses of morphological data.
Data from: The importance of microhabitat for biodiversity sampling
Responses to microhabitat are often neglected when ecologists sample animal indicator groups. Microhabitats may be particularly influential in non-passive biodiversity sampling methods, such as baited traps or light traps, and for certain taxonomic groups which respond to fine scale environmental variation, such as insects. Here we test the effects of microhabitat on measures of species diversity, guild structure and biomass of dung beetles, a widely used ecological indicator taxon. We demonstrate that choice of trap placement influences dung beetle functional guild structure and species diversity. We found that locally measured environmental variables were unable to fully explain trap-based differences in species diversity metrics or microhabitat specialism of functional guilds. To compare the effects of habitat degradation on biodiversity across multiple sites, sampling protocols must be standardized and scale-relevant. Our work highlights the importance of considering microhabitat scale responses of indicator taxa and designing robust sampling protocols which account for variation in microhabitats during trap placement. We suggest that this can be achieved either through standardization of microhabitat or through better efforts to record relevant environmental variables that can be incorporated into analyses to account for microhabitat effects. This is especially important when rapidly assessing the consequences of human activity on biodiversity loss and associated ecosystem function and services.
Data from: The impact of anchored phylogenomics and taxon sampling on phylogenetic inference in narrow-mouthed frogs (Anura, Microhylidae)
Despite considerable progress in unravelling the phylogenetic relationships of microhylid frogs, relationships among subfamilies remain largely unstable and many genera are not demonstrably monophyletic. Here, we used five alternative combinations of DNA sequence data (ranging from seven loci for 48 taxa to up to 73 loci for as many as 142 taxa) generated using the anchored phylogenomics sequencing method (66 loci, derived from conserved genome regions, for 48 taxa) and Sanger sequencing (seven loci for up to 142 taxa) to tackle this problem. We assess the effects of character sampling, taxon sampling, analytical methods and assumptions in phylogenetic inference of microhylid frogs. The phylogeny of microhylids shows high susceptibility to different analytical methods and datasets used for the analyses. Clades inferred from maximum-likelihood are generally more stable across datasets than those inferred from parsimony. Parsimony trees inferred within a tree-alignment framework are generally better resolved and better supported than those inferred within a similarity-alignment framework, even under the same cost matrix (equally weighted) and same treatment of gaps (as a fifth nucleotide state). We discuss potential causes for these differences in resolution and clade stability among discovery operations. We also highlight the problem that commonly used algorithms for model-based analyses do not explicitly model insertion and deletion events (i.e. gaps are treated as missing data). Our results corroborate the monophyly of Microhylidae and most currently recognized subfamilies but fail to provide support for relationships among subfamilies. Several taxonomic updates are provided, including naming of two new subfamilies, both monotypic.
Data from: Sampling strategies for delimiting species: genes, individuals, and populations in the Liolaemus elongatus-kriegi complex (Squamata: Liolaemidae) in Andean-Patagonian South America
Recovery of evolutionary history and delimiting species boundaries in widely distributed, poorly-known groups requires extensive geographic sampling, but this is difficult to design a priori because evolutionary diversity is often "hidden" by an inadequate taxonomy. Large data sets are needed, and these provide unique challenges for analysis when they span intra and inter-specific levels of divergence. Protocols have been designed to combine methods of analysis for DNA sequences that exhibit both very shallow and relatively deeper divergences (Crandall and Fitzpatrick, 1996). In this study we combine several tree-based phylogeny reconstruction methods with nested clade analysis, to extract maximum historical signal at various levels, in the poorly-known Liolaemus elongatus-kriegi complex in temperate South America. We implement the basic protocol of Wiens and Penkrot (2002) to test for species boundaries, and propose modifications to accommodate large data sets and gene regions with heterogeneous substitution rates. Combining haplotype trees with nested-clade analyses allowed testing of species boundaries on the basis of a priori defined criteria, and this approach suggests that the number of putative species could be doubled. We discuss these findings in the context of the advantages and limitations of a combined approach for retrieval of maximum historical information in large data sets, in the context of the yet formidable unresolved issues of sampling strategies.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.