Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
750
datasets available to search
ShareScore release 0.9.0
Dataset results
750 results for “heterogeneous data”
Data from: The effects of selective history and environmental heterogeneity on inbreeding depression in experimental populations of Drosophila melanogaster
Inbreeding depression varies considerably among populations, but only some aspects of this variation have been thoroughly studied. Because inbreeding depression requires genetic variation, factors that influence the amount of standing variation can affect the magnitude of inbreeding depression. Environmental heterogeneity has long been considered an important contributor to the maintenance of genetic variation, but its effects on inbreeding depression have been largely ignored by empiricists. Here we compare inbreeding depression, measured in two environments, for 20 experimental populations of Drosophila melanogaster that have been maintained under four different selection regimes, including two types of environmentally homogeneous selection and two types environmentally heterogeneous selection. In line with theory, we find considerably higher inbreeding depression in populations from heterogeneous selection regimes. We also use our data set to test whether inbreeding depression is correlated with either stress or the phenotypic coefficient of variation (CV), as suggested by some recent studies. Though both of these factors are significant predictors of inbreeding depression in our study, there is an effect of assay environment on inbreeding depression that cannot be explained by either stress or CV.
Data from: Modeling character change heterogeneity in phylogenetic analyses of morphology through the use of priors
The Mk model was developed for estimating phylogenetic trees from discrete morphological data, whether for living or fossil taxa. Like any model, the Mk model makes a number of assumptions. One assumption is that transitions between character states are symmetric (i.e., the probability of changing from 0 to 1 is the same as 1 to 0). However, some characters in a data matrix may not satisfy this assumption. Here, we test methods for relaxing this assumption in a Bayesian context. Using empirical datasets, we perform model fitting to illustrate cases in which modeling asymmetric transition rates among characters is preferable to the standard Mk model. We use simulated datasets to demonstrate that choosing the best-fit model of transition-state symmetry can improve model fit and phylogenetic estimation.
Data from: Invariant versus classical quartet inference when evolution is heterogeneous across sites and lineages
One reason why classical phylogenetic reconstruction methods fail to correctly infer the underlying topology is because they assume oversimplified models. In this paper we propose a quartet reconstruction method consistent with the most general Markov model of nucleotide substitution, which can also deal with data coming from mixtures on the same topology. Our proposed method uses phylogenetic invariants and provides a system of weights that can be used as input for quartet-based methods. We study its performance on real data and on a wide range of simulated 4-taxon data (both time-homogeneous and nonhomogeneous, with or without among-site rate heterogeneity, and with different branch length settings). We compare it to the classical methods of neighbor-joining (with paralinear distance), maximum likelihood (with different underlying models), and maximum parsimony. Our results show that this method is accurate and robust, has a similar performance to ML when data satisfies the assumptions of both methods, and outperforms the other methods when these are based on inappropriate substitution models. If alignments are long enough, then it also outperforms other methods when some of its assumptions are violated.
Data from: Demographic variability and heterogeneity among individuals within and among clonal bacteria strains
Identifying what drives individual heterogeneity has been of long interest to ecologists, evolutionary biologists and biodemographers, because only such identification provides deeper understanding of ecological and evolutionary population dynamics. In natural populations one is challenged to accurately decompose the drivers of heterogeneity among individuals as genetically fixed or selectively neutral. Rather than working on wild populations we present here data from a simple bacterial system in the lab, Escherichia coli. Our system, based on cutting-edge microfluidic techniques, provides high control over the genotype and the environment. It therefore allows to unambiguously decompose and quantify fixed genetic variability and dynamic stochastic variability among individuals. We show that within clonal individual variability (dynamic heterogeneity) in lifespan and lifetime reproduction is dominating at about 82–88%, over the 12–18% genetically (adaptive fixed) driven differences. The genetic differences among the clonal strains still lead to substantial variability in population growth rates (fitness), but, as well understood based on foundational work in population genetics, the within strain neutral variability slows adaptive change, by enhancing genetic drift, and lowering overall population growth. We also revealed a surprising diversity in senescence patterns among the clonal strains, which indicates diverse underlying cell-intrinsic processes that shape these demographic patterns. Such diversity is surprising since all cells belong to the same bacteria species, E. coli, and still exhibit patterns such as classical senescence, non-senescence, or negative senescence. We end by discussing whether similar levels of non-genetic variability might be detected in other systems and close by stating the open questions how such heterogeneity is maintained, how it has evolved, and whether it is adaptive.
Data from: Environmental heterogeneity does not affect levels of phenotypic plasticity in natural populations of three Drosophila species
Adaptation of natural populations to variable environmental conditions may occur by changes in trait means and/or in the levels of plasticity. Theory predicts that environmental heterogeneity favors plasticity of adaptive traits. Here we investigated the performance in several traits of three sympatric Drosophila species freshly collected in two environments that differ in the heterogeneity of environmental conditions. Differences in trait means within species were found in several traits, indicating that populations differed in their evolutionary response to the environmental conditions of their origin. Different species showed distinct adaptation with a very different role of plasticity across species for coping with environmental changes. However, geographically distinct populations of the same species generally displayed the same levels of plasticity as induced by fluctuating thermal regimes. This indicates a weak and trait-specific effect of environmental heterogeneity on plasticity. Furthermore, similar levels of plasticity were found in a laboratory-adapted population of Drosophila melanogaster with a common geographic origin but adapted to the laboratory conditions for more than 100 generations. Thus, this study does not confirm theoretical predictions on the degree of adaptive plasticity among populations in relation to environmental heterogeneity but shows a very distinct role of species-specific plasticity.
Data from: The influence of geographic heterogeneity in predation pressure on mating signal divergence in an Amazonian frog species complex
Sexual section plays an important role in mating signal divergence, but geographic variation in ecological factors can also contribute to divergent signal evolution. We tested the hypothesis that geographic heterogeneity in predation causes divergent selection on advertisement call complexity within the Engystomops petersi frog species complex. We conducted predator phonotaxis experiments at two sites where female choice is consistent with call trait divergence. Engystomops at one site produces complex calls, while the closely related species at the other site produces simple calls. Bats approached complex calls more than simple calls at both sites, suggesting selection against complex calls. Moreover, bat predation pressure was greater at the site with simple calls, suggesting stronger selection against complex calls and potentially precluding evolution of complex calls at this site. Our results show that geographic variation in predation may play an important role in the evolution and maintenance of mating signal divergence.
Data from: Optimal management strategy of insecticide resistance under various insect life histories: heterogeneous timing of selection and inter-patch dispersal
Although theoretical studies have shown that the mixture strategy, which uses multiple toxins simultaneously, can effectively delay the evolution of insecticide resistance, whether it is the optimal management under different insect life histories and insecticide types remains unknown. To test the robustness of the management strategy over the life histories, we developed a series of simulation models which cover almost all the diploid insect types and have the same basic structure describing pest population dynamics and resistance evolution with discrete time-steps. For each of two insecticidal toxins, one-locus two-allele autosomal inheritance of resistance was assumed. The simulations demonstrated the optimality of the mixture strategy either when insecticide efficacy was incomplete or when some part of the population disperses between patches before mating. The rotation strategy, which uses one insecticide on one pest generation and a different one on the next, did not differ from sequential usage in the time to resistance, except when dominance was low. It was the optimal strategy when insecticide efficacy was high and pre-mating selection and dispersal occur.
Data from: The relative importance of modeling site pattern heterogeneity versus partition-wise heterotachy in phylogenomic inference
Large taxa-rich genome-scale data sets are often necessary for resolving ancient phylogenetic relationships. But accurate phylogenetic inference requires that they are analyzed with realistic models that account for the heterogeneity in substitution patterns amongst the sites, genes and lineages. Two kinds of adjustments are frequently used: models that account for heterogeneity in amino acid frequencies at sites in proteins, and partitioned models that accommodate the heterogeneity in rates (branch lengths) among different proteins in different lineages (protein-wise heterotachy). Although partitioned and site-heterogeneous models are both widely used in isolation, their relative importance to the inference of correct phylogenies has not been carefully evaluated. We conducted several empirical analyses and a large set of simulations to compare the relative performances of partitioned models, site-heterogeneous models and combined partitioned site heterogeneous models. In general, site-homogeneous models (partitioned or not) performed worse than site heterogeneous, except in simulations with extreme protein-wise heterotachy. Furthermore, simulations using empirically-derived realistic parameter settings showed a marked long-branch attraction (LBA) problem for analyses employing protein-wise partitioning even when the generating model included partitioning. This LBA problem results from a small sample bias compounded over many single protein alignments. In some cases, this problem was ameliorated by clustering similarly-evolving proteins together into larger partitions using the PartitionFinder method. Similar results were obtained under simulations with larger numbers of taxa or heterogeneity in simulating topologies over genes. For an empirical Microsporidia test data set, all but one tested site-heterogeneous models (with or without partitioning) obtain the correct Microsporidia+Fungi grouping, whereas site-homogenous models (with or without partitioning) did not. The single exception was the fully partitioned site-heterogeneous analysis that succumbed to the compounded small sample LBA bias. In general unless protein-wise heterotachy effects are extreme, it is more important to model site-heterogeneity than protein-wise heterotachy in phylogenomic analyses. Complete protein-wise partitioning should be avoided as it can lead to a serious LBA bias. In cases of extreme protein-wise heterotachy, approaches that cluster similarly-evolving proteins together and coupled with site-heterogeneous models work well for phylogenetic estimation.
Data from: Lifetime fitness in wild baboons: tradeoffs and individual heterogeneity in quality
Understanding the evolution of life histories requires information on how life histories vary among individuals, and how such variation predicts individual fitness. Using complete life histories for females in a well-studied population of wild baboons, we tested two non-exclusive hypotheses about the relationships among survival, reproduction, and fitness: the quality hypothesis, which predicts positive correlations between life history traits, mediated by variation in resource acquisition, and the tradeoff hypothesis, which predicts negative correlations between life history traits, mediated by tradeoffs in resource allocation. In support of the quality hypothesis, we found that females with higher rates of offspring survival were themselves better at surviving. Further, after statistically controlling for variation in female quality, we found evidence for two types of tradeoffs: females who produced surviving offspring at a slower rate had longer lifespans than those who produced surviving offspring at a faster rate, and females who produced surviving offspring at a slower rate had a higher overall proportion of offspring survive infancy than females who produced surviving offspring at a faster rate. Importantly, these tradeoffs were evident even when accounting for: (i) the influence of offspring survival on maternal birth rate, (ii) the dependence of offspring survival on maternal survival, and (iii) potential age-related changes in birth rate and/or offspring survival. Our results shed light on why tradeoffs are evident in some populations, while variation in individual quality masks tradeoffs in others.
Data from: Lizards paid a greater opportunity cost to thermoregulate in a less heterogeneous environment
The theory of thermoregulation has developed slowly, hampering efforts to predict how individuals can buffer climate change through behaviour. Mixed results of field and laboratory experiments underscore the need to test hypotheses about thermoregulation explicitly, while measuring costs and benefits in different thermal landscapes. We simulated body temperature and energy expenditure of a virtual lizard that either thermoregulates optimally or thermoconforms in a landscape of either low or high quality (one or four basking sites, respectively). We then compare the predicted values in each landscape with the observed values for real lizards in experimental arenas. Lizards thermoregulated more accurately in the high-quality landscape than they did on the low-quality landscape, albeit only slightly so, but spent similar amounts of energy in these landscapes. Basking, rather than shuttling between heat sources, accounted for the majority of the energy consumed in both landscapes. These results did not support the predictions of our model. In the low-quality landscape, real lizards thermoregulated intensely despite the potential to save energy by thermoconforming. In the high-quality landscape, lizards moved more than expected, suggesting that lizards explored their surroundings despite being able to thermoregulate without doing so. Our results suggest that non-energetic benefits drive thermoregulatory behaviour in costly environments, despite the missed opportunities arising from thermoregulation. We propose that energetic costs associated with thermoregulatory movement will become substantial in homogeneous environments such as flat plains and dense forests. The theory of thermoregulation should incorporate these aspects if biologists wish to predict responses of ectotherms to changing climates and habitats.
Data from: Host use dynamics in a heterogeneous fitness landscape generates oscillations in host range and diversification
Colonization of novel hosts is thought to play an important role in parasite diversification, yet little consensus has been achieved about the macroevolutionary consequences of changes in host use. Here we offer a mechanistic basis for the origins of parasite diversity by simulating lineages evolved in silico. We describe an individual-based model in which (i) parasites undergo sexual reproduction limited by genetic proximity, (ii) hosts are uniformly distributed along a one-dimensional resource gradient, and (iii) host use is determined by the interaction between the phenotype of the parasite and a heterogeneous fitness landscape. We found two main effects of host use on the evolution of a parasite lineage. First, the colonization of a novel host allowed parasites to explore new areas of the resource space, increasing phenotypic and genotypic variation. Second, hosts produced heterogeneity in the parasite fitness landscape, which led to reproductive isolation and therefore, speciation. As a validation of the model, we analyzed empirical data from Nymphalidae butterflies and their host plants. We then assessed the number of hosts used by parasite lineages and the diversity of resources they encompass. In both simulated and empirical systems, host diversity emerged as the main predictor of parasite species richness.
Data from: Heterogeneous models place the root of the placental mammal phylogeny
Heterogeneity among life traits in mammals has resulted in considerable phylogenetic conflict, particularly concerning the position of the placental root. Layered upon this are gene- and lineage-specific variation in amino acid substitution rates and compositional biases. Life trait variations that may impact upon mutational rates are longevity, metabolic rate, body size and germ line generation time. Over the past 12 years, three main conflicting hypotheses have emerged for the placement of the placental root. These hypotheses place: the Atlantogenata (common ancestor of Xenarthra plus Afrotheria), the Afrotheria, or the Xenarthra as the sister group to all other placental mammals. Model adequacy is critical for accurate tree reconstruction and by failing to account for these compositional and character exchange heterogeneities across the tree and dataset, previous studies have not provided a strongly supported hypothesis for the placental root. For the first time, models that accommodate both tree and dataset heterogeneity have been applied to mammal data. Here we show the impact of accurate model assignment and the importance of datasets in accommodating model parameters while maintaining the power to reject competing hypotheses. Through these sophisticated methods, we demonstrate the importance of model adequacy, dataset power and provide strong support for the Atlantogenata over other competing hypotheses for the position of the placental root.
Data from: Gametic selection, developmental trajectories, and extrinsic heterogeneity in Haldane's rule
Deciphering the genetic and developmental causes of the disproportionate rarity, inviability and sterility of hybrid males, Haldane's rule, is important for understanding the evolution of reproductive isolation between species. Moreover, extrinsic and pre-zygotic factors can contribute to the magnitude of intrinsic isolation experienced between species with partial reproductive compatibility. Here we use the nematodes Caenorhabditis briggsae and C. nigoni to quantify the sensitivity of hybrid male viability to extrinsic temperature and developmental timing, and test for a role of mito-nuclear incompatibility as a genetic cause. We demonstrate that hybrid male inviability manifests almost entirely as embryonic, not larval, arrest and is maximal at the lowest rearing temperatures, indicating an intrinsic-by-extrinsic interaction to hybrid inviability. Crosses using mitochondrial substitution strains that have reciprocally introgressed mitochondrial and nuclear genomes show that mito-nuclear incompatibility is not a dominant contributor to post-zygotic isolation and does not drive Haldane's rule in this system. Crosses also reveal that competitive superiority of X-bearing sperm provides a novel means by which post-mating pre-zygotic factors exacerbate the rarity of hybrid males. These findings highlight the important roles of gametic, developmental, and extrinsic factors in modulating the manifestation of Haldane's rule.
Data from: Environmental heterogeneity generates opposite gene-by-environment interactions for two fitness-related traits within a population
Theory predicts that environmental heterogeneity offers a potential solution to the maintenance of genetic variation within populations, but empirical evidence remains sparse. The livebearing fish Xiphophorus variatus exhibits polymorphism at a single locus, with different alleles resulting in up to five distinct melanistic "tailspot" patterns within populations. We investigated the effects of heterogeneity in two ubiquitous environmental variables (temperature and food availability) on two fitness-related traits (upper thermal limits and body condition) in two different tailspot types (wildtype and upper cut crescent). We found gene-by-environment (GxE) interactions between tailspot type and food level affecting upper thermal limits (UTL), as well as between tailspot type and thermal environment affecting body condition. Exploring mechanistic bases underlying these GxE patterns, we found no differences between tailspot types in hsp70 gene expression despite significant overall increases in expression under both thermal and food stress. Similarly, there was no difference in routine metabolic rates between the tailspot types. The reversal of relative performance of the two tailspot types under different environmental conditions revealed a mechanism by which environmental heterogeneity can balance polymorphism within populations through selection on different fitness-related traits.
Data for: Density dependence and spatial heterogeneity limit the population growth rate of invasive pines at the landscape scale
<p class="MsoBodyText"><span><span><span><span><span><span><span><span><span><span><span>Determining population growth across large scales is difficult because it is often impractical to collect data at large scales and over long timespans. Instead, the growth of a population is often only measured at a small, plot-level scale and then extrapolated to derive a mean field estimate. However, this approach is prone to error since it simplifies spatial processes such as the neighbourhood effects of density and dispersal. We present a novel approach that estimates how spatial processes derived from the effects of density and dispersal affect population growth between plot scales and landscape scales. The method is based on a scale transition theory and calculates a transition term to measure the spatial scaling of population growth, which we extend to unstable, expanding populations in order to assess whether landscape-scale population dynamics are different from those estimated at smaller spatial scales. We illustrate this approach using aerial imagery of eight locations in New Zealand experiencing non-native pine invasions. Analyses examined the dynamics at a plot scale (1 hectare) and compared this to estimates across entire landscapes (between 24 and 1600 hectares), in several cases for more than one time period. We used a Bayesian spatial random effects model to examine population growth and to account for neighbourhood effects and dispersal between plots in a rapidly changing system. </span></span></span></span></span></span></span></span></span></span></span></p> <p class="MsoBodyText"><span><span><span><span><span><span><span><span><span><span><span>We found that the estimates of the scale transition term were typically 10-25% of the mean field estimates, which led to mean field estimates of population growth extrapolated from plots being considerably higher than landscape estimates. The approach we have developed will not only have applications for predicting the populations' growth of invasive species, but also for studies examining the scaling of landscape-scale phenomena.</span></span></span></span></span></span></span></span></span></span></span></p>
Data for manuscript "Integrated microwave photonic notch filter using a heterogeneously integrated Brillouin and active-silicon photonic circuit"
<p>Experimental data and scripts to plot results for manuscript "Integrated microwave photonic notch filter using a heterogeneously integrated Brillouin and active-silicon photonic circuit"</p>
Developmental Basis of SHH Medulloblastoma Heterogeneity (mIHC Data)
Open the record for dataset details and reuse information.
Data from: Assessing urban-scale spatiotemporal heterogeneous metro station coverage using multi-source mobility data
Open the record for dataset details and reuse information.
Data from: Mitochondrial phylogenomics of early land plants: mitigating the effects of saturation, compositional heterogeneity, and codon-usage bias
Phylogenetic analyses using concatenation of genomic-scale data have been seen as the panacea to resolving the incongruences among inferences from few or single genes. However, phylogenomics may also suffer from systematic errors, due to the, perhaps cumulative, effects of saturation, among-taxa compositional (GC content) heterogeneity, or codon-usage bias plaguing the individual nucleotide loci that are concatenated. Here we provide an example of how these factors affect the inferences of the phylogeny of early land plants based on mitochondrial genomic data. Mitochondrial sequences evolve slowly in plants and hence are thought to be suitable for resolving deep relationships. We newly assembled mitochondrial genomes from 20 bryophytes, complemented these with 40 other streptophytes (land plants plus algal outgroups), compiling a data matrix of 60 taxa and 41 mitochondrial genes. Homogeneous analyses of the concatenated nucleotide data resolve mosses as sister-group to the remaining land plants. However, the corresponding translated amino acid data support the liverwort lineage in this position. Both results receive weak to moderate support in maximum likelihood analyses, but strong support in Bayesian inferences. Tests of alternative hypotheses using either nucleotide or amino-acid data provide implicit support for the respective optimal topologies. By analyzing the nucleotide data, we found that the 3rd codon positions are more saturated than the 1st and 2nd codon positions, and excluding these from the analyses leads to a topology congruent with that obtained using amino-acid data. Further, we determined that land plant lineages differ in their nucleotide composition, and in their usage of synonymous codon variants. Composition heterogeneous Bayesian analyses employing a non-stationary model that accounts for variation in among-lineage composition, and inferences from degenerated nucleotide data that avoids the effects of synonymous mutations that underlie codon-usage bias, again recovered liverworts being sister to the remaining land plants. These analyses indicate that the discrepancy between the nucleotide-based and the amino acid-based trees is caused by the lineage specific, parallel compositional bias, or synonymous mutations driving codon-usage bias, as well as saturation in the 3rd codon positions. While genomic data may generate highly supported phylogenetic trees, these inferences may be artifacts. We suggest that phylogenomic analyses should assess the possible impact of potential biases through comparisons of protein coding gene data and their amino-acids translations, by analyzing data modeling compositional bias, and by excluding nucleotide noisy signals due to saturation or codon-usage bias. We caution against relying on any one presentation of the data (nucleotide or amino acid) or any one type of analysis even when analyzing large-scale data sets, no matter how well-supported, without fully exploring the effects of substitution models.
Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species
The ribosomal RNA encapsulates a wealth of evolutionary information, including genetic variation that can be used to discriminate between organisms at a wide range of taxonomic levels. For example, the prokaryotic 16S rDNA sequence is very widely used both in phylogenetic studies and as a marker in metagenomic surveys and the ITS region, frequently used in plant phylogenetics, is now recognised as a fungal DNA barcode. However, this widespread use does not escape criticism, principally due to issues such as difficulties in classification of paralogous versus orthologous rDNA units and intragenomic variation, both of which may be significant barriers to accurate phylogenetic inference. We recently analysed datasets from the Saccharomyces Genome Resequencing Project, characterising rDNA sequence variation within multiple strains of the baker's yeast <i>Saccharomyces cerevisiae</i> and its nearest wild relative <i>Saccharomyces paradoxus</i> in unprecedented detail. Notably, both species possess single locus rDNA systems. Here, we use these new variation datasets to assess whether a more detailed characterisation of the rDNA locus can alleviate the second of these phylogenetic issues, sequence heterogeneity, while controlling for the first. We demonstrate that a strong phylogenetic signal exists within both datasets and illustrate how they can be used, with existing methodology, to estimate intra-species phylogenies of yeast strains consistent with those derived from whole-genome approaches. We also describe the use of partial Single Nucleotide Polymorphisms, a type of sequence variation found only in repetitive genomic regions, in identifying key evolutionary features such as genome hybridisation events and show their consistency with whole-genome Structure analyses. We conclude that our approach can transform rDNA sequence heterogeneity from a problem to a useful source of evolutionary information, enabling the estimation of highly accurate phylogenies of closely related organisms, and discuss how it could be extended to future studies of multi-locus rDNA systems.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.