Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
266
datasets available to search
ShareScore release 0.9.0
Dataset results
266 results for “data partitioning”
Data from: Complementarity in spatial subsidies of carbon associated with resource partitioning along multiple niche axes
Open the record for dataset details and reuse information.
Data Set "Systematic partitioning of proteins for quantum-chemical fragmentation methods using graph algorithms"
<p>Data set accompanying the publication "Systematic partitioning of proteins for quantum-chemical fragmentation methods using graph algorithms"</p> <p>The data set contains:</p> <p>- Input script for PyADF (v0.97) for calculating (a) all two body terms to use as graph weights and (b) fragmentation error for all k and nmax (aspf)</p> <p>- PDB files of proteins and the "regions of interest" (RoI) used in this work.</p> <p>- Raw data: protein graph representations, resulting partitions, data underlying all figures shown in our article.</p> <p>- Jupiter notebook to create all figures shown in the article and in the supporting information from data in the results folder.</p> <p>- Images of protein structures and graph representations of ubiquitin.</p>
Data from: Interrelationships of basal synapsids: cranial and postcranial morphological partitions suggest different topologies
Basal synapsids ('pelycosaurs') form the basalmost portion of the mammalian stem lineage and document the transition from primitive 'reptile-like' basal amniotes to derived, mammal-like therapsids. They dominated terrestrial ecosystems of the latest Carboniferous and Early Permian (∼300–271 million years ago), producing large-bodied terrestrial animals (3–6.5 metres long), high-fibre herbivores, and macropredators for the first time in vertebrate history, alongside an array of smaller-bodied forms. Despite numerous recent discoveries and reassessments of fossils collected over the past 250 years, and despite their importance for understanding the early diversification of terrestrial vertebrates, a comprehensive assessment of global relationships among basal synapsids has not been undertaken. A new phylogenetic dataset comprising 45 taxa (plus four outgroups and four therapsids) and 239 characters (147 cranial; 92 postcranial) reveals considerable uncertainty in the relationships of higher clades of basal synapsids. Although cranial data support the current consensus that Caseasauria is the most basal clade, postcranial data and the full dataset suggest that a clade of Ophiacodontidae + Varanopidae occupies this position. Although relationships within higher clades are well supported, relationships among those clades are poorly supported. The likely source of this uncertainty lies in the exceptionally poor early record of the group, which renders determinations of the plesiomorphic condition of higher clades speculative, although cranial data are generally represented by shorter ghost lineages and should perhaps be favoured. The new dataset suggests well-supported phylogenetic placements for several taxa of historically uncertain affinities: Trichasaurus is a caseid; Lupeosaurus is an edaphosaurid; and Basicranodon and Ruthiromia are varanopids.
Data from: High-dimensional variance partitioning reveals the modular genetic basis of adaptive divergence in gene expression during reproductive character displacement
Although adaptive change is usually associated with complex changes in phenotype, few genetic investigations have been conducted of adaptations that involve sets of high dimensional traits. Microarrays have supplied high-dimensional descriptions of gene expression, and phenotypic change resulting from adaptation often results in large-scale changes in gene expression. We demonstrate how genetic analysis of large-scale changes in gene expression generated during adaptation can be accomplished by determining by high-dimensional variance partitioning within classical genetic experimental designs. A microarray experiment conducted on a panel of recombinant inbred lines (RILs) generated from two populations of Drosophila serrata that have diverged in response to natural selection, revealed genetic divergence in 10.6% of 3762 gene products examined. Over 97% of the genetic divergence in transcript abundance was explained by only 12 genetic modules. The two most important modules, explaining 50% of the genetic variance in transcript abundance, were genetically correlated with the morphological traits that are known to be under selection. The expression of three candidate genes from these two important genetic modules was assessed in an independent experiment using qRT-PCR on 430 individuals from the panel of RILs, and confirmed the genetic association between transcript abundance and morphological traits under selection.
Data from: Functional niche partitioning in Therizinosauria provides new insights into the evolution of theropod herbivory
Dietary specialization is generally considered to be a crucial factor in driving morphological evolution across extant and extinct vertebrates. The ability to adapt to a specific diet and to exploit ecological niches is thereby influenced by functional morphology and biomechanical properties. Differences in functional behaviour and efficiency can therefore allow dietary diversification and the coexistence of similarly adapted taxa. Therizinosauria, a group of secondarily herbivorous theropod dinosaurs, is characterized by a suite of morphological traits thought to be indicative of adaptations to an herbivorous diet. Digital reconstruction, theoretical modelling and computer simulations of the mandibles of therizinosaur dinosaurs provides evidence for functional niche partitioning in adaptation to herbivory. Different mandibular morphologies present in therizinosaurians were found to correspond to different dietary strategies permitting coexistence of taxa. Morphological traits indicative of an herbivorous diet, such as a downturned tip of the lower jaw and an expanded postdentary region, were identified as having stress mitigating effects. The more widely distributed occurrence of these purported herbivorous traits across different dinosaur clades suggests that these features also could have played an important role in the evolution and acquisition of herbivory in other groups.
Data from: Accounting for uncertainty in the evolutionary timescale of green plants through clock-partitioning and fossil calibration strategies
Establishing an accurate evolutionary timescale for green plants (Viridiplantae) is essential to understanding their interaction and coevolution with the Earth's climate and the many organisms that rely on green plants. Despite being the focus of numerous studies, the timing of the origin of green plants and the divergence of major clades within this group remain highly controversial. Here, we infer the evolutionary timescale of green plants by analysing 81 protein-coding genes from 99 chloroplast genomes, using a core set of 21 fossil calibrations. We test the sensitivity of our divergence-time estimates to various components of Bayesian molecular dating, including the tree topology, clock models, clock-partitioning schemes, rate priors, and fossil calibrations. We find that the choice of clock model affects date estimation and that the independent-rates model provides a better fit to the data than the autocorrelated-rates model. Varying the rate prior and tree topology had little impact on age estimates, with far greater differences observed among calibration choices and clock-partitioning schemes. Our analyses yield date estimates ranging from the Paleoproterozoic to Mesoproterozoic for crown-group green plants, and from the Ediacaran to Middle Ordovician for crown-group land plants. We present divergence-time estimates of the major groups of green plants that take into account various sources of uncertainty. Our proposed timeline lays the foundation for further investigations into how green plants shaped the global climate and ecosystems, and how embryophytes became dominant in terrestrial environments.
Data from: Inference of genetic architecture from chromosome partitioning analyses is sensitive to genome variation, sample size, heritability and effect size distribution
Genomewide association studies have contributed immensely to our understanding of the genetic basis of complex traits. One major conclusion arising from these studies is that most traits are controlled by many loci of small effect, confirming the infinitesimal model of quantitative genetics. A popular approach to test for polygenic architecture involves so‐called "chromosome partitioning" where phenotypic variance explained by each chromosome is regressed on the size of the chromosome. First developed for humans, this has now been repeatedly used in other species, but there has been no evaluation of the suitability of this method in species that can differ in their genome characteristics such as number and size of chromosomes. Nor has the influence of sample size, heritability of the trait, effect size distribution of loci controlling the trait or the physical distribution of the causal loci in the genome been examined. Using simulated data, we show that these characteristics have major influence on the inferences of the genetic architecture of traits we can infer using chromosome partitioning analyses. In particular, small variation in chromosome size, small sample size, low heritability, a skewed effect size distribution and clustering of loci can lead to a loss of power and consequently altered inference from chromosome partitioning analyses. Future studies employing this approach need to consider and derive an appropriate null model for their study system, taking these parameters into consideration. Our simulation results can provide some guidelines on these matters, but further studies examining a broader parameter space are needed.
Data from: Homoplasy-based partitioning outperforms alternatives in Bayesian analysis of discrete morphological data
Bayesian analysis of morphological data is becoming increasingly popular mainly (but not only) because it allows for time-calibrated phylogenetic inference using relaxed morphological clocks and tip dating whenever fossils are available. As with molecular data, recent studies have shown that modeling among character rate variaton (ACRV) in morphological matrices greatly improves phylogenetic inference. In a likelihood framework this may be accomplished, for instance, by employing a hidden Markov model (HMM) to assign characters to rate categories drawn from a (discretized) Γ distribution and/or by partitioning datasets according to rate heterogeneity and estimating per-partition branch lengths, conditioned on a single topology. While the first approach is available in many phylogenetic analysis software, there is still no clear consensus on how to partition data, except perhaps in the simplest cases (e.g. "by codon" partitioning of coding sequences). Additionally, there is a trade-off between improvement in likelihood scores and the number of free parameters in the analysis, which rises quickly with the number of partitions. This trade-off may be dealt with by employing statistics that penalize overfitting of complex models, such as Akaike or Bayesian information criteria (AIC and BIC), or the more recently introduced stepping-stone (SS) method for marginal likelihood approximation. We applied the latter to three distinct matrices of discrete morphological data and demonstrated that sorting characters by homoplasy scores (obtained from implied weighting parsimony analysis) outperformed other partitioning strategies (anatomically-based and PartitionFinder2). The method was in fact so efficient in segregating characters by rates of evolution that no within-partition ACRV modeling was necessary, while among partition rate variation (APRV) was adequately accommodated by rate multipliers. We conclude that partitioning by homoplasy is a powerful and easy-to-implement strategy to address ACRV in complex datasets. We provide some guidelines focusing on morphological matrices, although this approach may be also applicable to molecular datasets.
Data from: Vertical partitioning between sister species of Rhizopogon fungi on mesic and xeric sites in an interior Douglas-fir forest
Understanding ectomycorrhizal fungal (EMF) community structure is limited by a lack of taxonomic resolution and autecological information. Rhizopogon vesiculosus and R. vinicolor (Basidiomycota) are morphologically and genetically related species. They are dominant members of interior Douglas-fir (Pseudotsuga menziesii var. glauca) EMF communities, but mechanisms leading to their coexistence are unknown. We investigated the microsite associations and foraging strategy of individual R. vesiculosus and R. vinicolor genets. Mycelia spatial patterns, pervasiveness and root colonization patterns of fungal genets were compared between Rhizopogon species and between xeric and mesic soil moisture regimes. Rhizopogon spp. mycelia were systematically excavated from the soil and identified using microsatellite DNA markers. Rhizopogon vesiculosus mycelia occurred at greater depth, were more spatially pervasive, and colonized more tree roots than R. vinicolor mycelia. Both species were frequently encountered in organic layers and between the interface of organic and mineral horizons. They were particularly abundant within microsites associated with soil moisture retention. The occurrence of R. vesiculosus shifted in the presence of R. vinicolor towards mineral soil horizons, where R. vinicolor was mostly absent. This suggests that competition and foraging strategy may contribute towards the vertical partitioning observed between these species. R. vesiculosus and R. vinicolor mycelia systems occurred at greater mean depths and were more pervasive in mesic plots compared to xeric plots. The spatial continuity and number of trees colonized by genets of each species did not significantly differ between soil moisture regimes.
Data from: Predation risk and resource abundance mediate foraging behaviour and intraspecific resource partitioning among consumers in dominance hierarchies
Dominance hierarchies and the resulting unequal resource partitioning among individuals are key mechanisms of population regulation. The strength of dominance hierarchies can be influenced by size-dependent trade-offs between foraging and predator avoidance whereby competitively inferior subdominants can access a larger proportion of limiting resources by accepting higher predation risk. Foraging-predation risk trade-offs also depend on resource abundance. Yet, few studies have manipulated predation risk and resource abundance simultaneously; consequently, their joint effect on resource partitioning within dominance hierarchies are not well understood. We addressed this gap by measuring behavioural responses of masu salmon (Oncorhynchus masou ishikawae) to experimental manipulations of predation risk and resource abundance in a natural temperate forest stream. Responses to predation risk depended on body size and social status such that larger fish (often social dominants) exhibited more risk-averse behaviour (e.g., lower foraging and appearance rates) than smaller subdominants after exposure to a simulated predator. The magnitude of this effect was lower when resources were elevated, indicating that dominant fish accepted a higher predation risk to forage on abundant resources. However, the influence of resource abundance did not extend to the population level, where predation risk altered the distribution of foraging attempts (a proxy for energy intake) from being skewed towards large individuals to being skewed towards small individuals after predator exposure. Our results imply that size-dependent foraging-predation risk trade-offs can weaken the strength of dominance hierarchies by allowing competitively inferior subdominants to access resources that would otherwise be monopolized.
Data from: Niche partitioning in a sympatric cryptic species complex
Competition theory states that multiple species should not be able to occupy the same niche indefinitely. Morphologically, similar species are expected to be ecologically alike and exhibit little niche differentiation, which makes it difficult to explain the co-occurrence of cryptic species. Here, we investigated interspecific niche differentiation within a complex of cryptic bumblebee species that co-occur extensively in the United Kingdom. We compared the interspecific variation along different niche dimensions, to determine how they partition a niche to avoid competitive exclusion. We studied the species B. cryptarum, B. lucorum, and B. magnus at a single location in the northwest of Scotland throughout the flight season. Using mitochondrial DNA for species identification, we investigated differences in phenology, response to weather variables and forage use. We also estimated niche region and niche overlap between different castes of the three species. Our results show varying levels of niche partitioning between the bumblebee species along three niche dimensions. The species had contrasting phenologies: The phenology of B. magnus was delayed relative to the other two species, while B. cryptarum had a relatively extended phenology, with workers and males more common than B. lucorum early and late in the season. We found divergent thermal specialisation: In contrast to B. cryptarum and B. magnus, B. lucorum worker activity was skewed toward warmer, sunnier conditions, leading to interspecific temporal variation. Furthermore, the three species differentially exploited the available forage plants: In particular, unlike the other two species, B. magnus fed predominantly on species of heather. The results suggest that ecological divergence in different niche dimensions and spatio-temporal heterogeneity in the environment may contribute to the persistence of cryptic species in sympatry. Furthermore, our study suggests that cryptic species provide distinct and unique ecosystem services, demonstrating that morphological similarity does not necessarily equate to ecological equivalence.
Data from: PartitionFinder: combined selection of partitioning schemes and substitution models for phylogenetic analyses.
In phylogenetic analyses of molecular sequence data, partitioning involves estimating independent models of molecular evolution for different sets of sites in a sequence alignment. Choosing an appropriate partitioning scheme is an important step in most analyses because it can affect the accuracy of phylogenetic reconstruction. Despite this, partitioning schemes are often chosen without explicit statistical justification. Here, we describe two new objective methods for the combined selection of best-fit partitioning schemes and nucleotide substitution models. These methods allow millions of partitioning schemes to be compared in realistic timeframes, and so permit the objective selection of partitioning schemes even for large multi-locus DNA datasets. We demonstrate that these methods significantly outperform previous approaches, including the ad hoc selection of partitioning schemes (e.g. partitioning by gene or codon position), and a recently proposed hierarchical clustering method. We have implemented these methods in an open-source program, PartitionFinder. This program allows users to select partitioning schemes and substitution models using a range of information-theoretic metrics (e.g. the BIC, AIC, and AICc). We hope that PartitionFinder will encourage the objective selection of partitioning schemes, and thus lead to improvements in phylogenetic analyses. PartitionFinder is written in Python and runs under Mac OSX 10.4 and above. The program, source code, and a detailed manual are freely available from .
Data from: Inferring the potentially complex genetic architectures of adaptation, sexual dimorphism, and genotype by environment interactions by partitioning of mean phenotypes.
Genetic architecture fundamentally affects the way that traits evolve. However, the mapping of genotype to phenotype includes complex interactions with the environment or even the sex of an organism that can modulate the expressed phenotype. Line cross analysis is a powerful quantitative genetics method to infer genetic architecture by analyzing the mean phenotype value of two diverged strains and a series of subsequent crosses and backcrosses. However, it has been difficult to account for complex interactions with the environment or sex within this framework. We have developed extensions to line cross analysis that allow for gene by environment and gene by sex interactions. Using extensive simulations studies and reanalysis of empirical data, we show that our approach can account for both unintended environmental variation when crosses cannot be reared in a common garden and can be used to test for the presence of gene by environment or gene by sex interactions. In analyses that fail to account for environmental variation between crosses we find that line cross analysis has low power and high false positive rates. However, we illustrate that accounting for environmental variation allows for the inference of adaptive divergence, and that accounting for sex differences in phenotypes allows practitioners to infer the genetic architecture of sexual dimorphism.
Data from: The relative importance of modeling site pattern heterogeneity versus partition-wise heterotachy in phylogenomic inference
Large taxa-rich genome-scale data sets are often necessary for resolving ancient phylogenetic relationships. But accurate phylogenetic inference requires that they are analyzed with realistic models that account for the heterogeneity in substitution patterns amongst the sites, genes and lineages. Two kinds of adjustments are frequently used: models that account for heterogeneity in amino acid frequencies at sites in proteins, and partitioned models that accommodate the heterogeneity in rates (branch lengths) among different proteins in different lineages (protein-wise heterotachy). Although partitioned and site-heterogeneous models are both widely used in isolation, their relative importance to the inference of correct phylogenies has not been carefully evaluated. We conducted several empirical analyses and a large set of simulations to compare the relative performances of partitioned models, site-heterogeneous models and combined partitioned site heterogeneous models. In general, site-homogeneous models (partitioned or not) performed worse than site heterogeneous, except in simulations with extreme protein-wise heterotachy. Furthermore, simulations using empirically-derived realistic parameter settings showed a marked long-branch attraction (LBA) problem for analyses employing protein-wise partitioning even when the generating model included partitioning. This LBA problem results from a small sample bias compounded over many single protein alignments. In some cases, this problem was ameliorated by clustering similarly-evolving proteins together into larger partitions using the PartitionFinder method. Similar results were obtained under simulations with larger numbers of taxa or heterogeneity in simulating topologies over genes. For an empirical Microsporidia test data set, all but one tested site-heterogeneous models (with or without partitioning) obtain the correct Microsporidia+Fungi grouping, whereas site-homogenous models (with or without partitioning) did not. The single exception was the fully partitioned site-heterogeneous analysis that succumbed to the compounded small sample LBA bias. In general unless protein-wise heterotachy effects are extreme, it is more important to model site-heterogeneity than protein-wise heterotachy in phylogenomic analyses. Complete protein-wise partitioning should be avoided as it can lead to a serious LBA bias. In cases of extreme protein-wise heterotachy, approaches that cluster similarly-evolving proteins together and coupled with site-heterogeneous models work well for phylogenetic estimation.
Data from: Ecological partitioning among parapatric cryptic species
Geographic range differences among species may result from differences in their physiological tolerances. In the intertidal zone, marine and terrestrial environments intersect to create a unique habitat, across which physiological tolerance strongly influences range. Traits to cope with environmental extremes are particularly important here because many species live near their physiological limits and environmental gradients can be steep. The snail Melampus bidentatus occurs in coastal salt marshes in the western Atlantic and the Gulf of Mexico. We used sequence data from one mitochondrial (COI) and two nuclear markers (histone H3 and a mitochondrial carrier protein, MCP) to identify three cryptic species within this broad-ranging nominal species, two of which have partially overlapping geographic ranges. High genetic diversity, low population structure, and high levels of migration within these two overlapping species suggest that historical range limitations do not entirely explain their different ranges. To identify microhabitat differences between these two species, we modeled their distributions using data from both marine and terrestrial environments. Although temperature was the largest factor setting range limits, other environmental components explained features of the ranges that temperature alone could not. In particular, the interaction of precipitation and salinity likely sets physiological limits that lead to range differences between these two cryptic species. This suggests that the response to climatic change in these snails will be mediated by changes to multiple environmental factors, and not just to temperature alone.
Data from: Below-ground resource partitioning alone cannot explain the biodiversity–ecosystem function relationship: a field test using multiple tracers
1. Belowground resource partitioning is among the most prominent hypotheses for driving the positive biodiversity-ecosystem function relationship. However, experimental tests of this hypothesis in biodiversity experiments are scarce, and the available evidence is not consistent. 2. We tested the hypothesis that resource partitioning in space, in time, or in both space and time combined drives the positive effect of diversity on both plant productivity and community resource uptake. At the community level, we predicted that total community resource uptake and biomass production above- and belowground will increase with increased species richness or functional group richness. We predicted that at the species level resource partition breadth will become narrower, and that overlap between the resource partitions of different species will become smaller with increasing species richness or functional group richness. 3. We applied multiple resource tracers (Li and Rb as potassium analogues, the water isotopologues - H218O and 2H2O, and 15N) in three seasons at two depths across a species and functional group richness gradient at a grassland biodiversity experiment. We used this multidimensional resource tracer study to test if plant species partition resources with increasing plant diversity across space, time, or both simultaneously. 4. At the community level, community resource uptake of nitrogen and potassium and above- and belowground biomass increased significantly with increasing species richness but not with increasing functional group richness. However, we found no evidence that resource partition breadth or resource partition overlap decreased with increasing species richness for any resource in space, time, or both space and time combined. Synthesis: These findings indicate that belowground resource partitioning may not drive the enhanced resource uptake or biomass production found here. Instead, other mechanisms such as facilitation, species-specific biotic feedback, or aboveground resource partitioning are likely necessary for enhanced overall ecosystem function.
Data from: Biomass partitioning of plants under soil pollution stress
<p><span>Polluted sites are ubiquitous worldwide but how plant partition their biomass between different organs in this context is unclear. </span><span>Here, we identified three possible drivers of biomass partitioning in our controlled study along pollution gradients: plant size reduction (pollution effect) combined with allometric scaling between organs; early deficit in root surfaces (pollution effect) inducing a decreased water uptake; increased biomass allocation to roots to compensate for lower soil resource acquisition consistent with the optimal partitioning theory (plant response). A complementary meta-analysis showed variation in biomass </span><span>partitioning</span><span> across published studies, with g</span><span>rass and woody species having distinct modifications of their root: shoot ratio. However, the modelling of biomass partitioning drivers showed that</span> <span>single harvest experiments performed in previous studies prevent identifying the main drivers at stake.</span> <span>The proposed distinction between pollution effects and plant response will help to improve our knowledge of plant allocation strategies in the context of pollution.</span></p>
Data from: Are all hosts created equal? Partitioning host species contributions to parasite persistence in multihost communities
[No abstract entered]
Supporting Data for Manuscript "Unraveling Thermally Regulated Gating Mechanisms in TPT Pore-Partitioned MOF-74: A Computational Endeavor"
<p>This dataset contains the supplementary computational data associated with the research article titled "Unraveling Thermally Regulated Gating Mechanisms in TPT Pore-Partitioned MOF-74: A Computational Endeavor" (currnetly under review). It encompasses trajectory files and calculated properties of TPT-X ligands in various modeling scenarios, providing insights into their dynamic behavior within the MOF-74 framework under different conditions.</p>
Data Sets for "Efficient Solution of the Number Partitioning Problem on a Quantum Annealer: A Hybrid Quantum-Classical Decomposition Approach"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.