Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
116
datasets available to search
ShareScore release 0.7.1
Dataset results
116 results for “phylogenetic signal”
Data from: Geometric morphometric character suites as phylogenetic data: extracting phylogenetic signal from gastropod shells
Open the record for dataset details and reuse information.
Conflicting phylogenetic signals in genomic data of the coffee family (Rubiaceae)
Open the record for dataset details and reuse information.
Data from: Clock gene evolution: seasonal timing, phylogenetic signal, or functional constraint?
Open the record for dataset details and reuse information.
Data from: Phylogenetic diversity and coevolutionary signals among trophic levels change across a habitat edge
Open the record for dataset details and reuse information.
Phylogenetic signal and bias in paleontology
Open the record for dataset details and reuse information.
Data from: Extracting phylogenetic signal from phylogenomic data: higher-level relationships of the nightbirds (Strisores)
A well-resolved phylogeny would facilitate study of adaptation to nocturnality in the avian superorder Strisores, a group that includes both nocturnal and diurnal lineages. Based on previous estimates, it could be hypothesized that there were multiple independent origins of nocturnality in this group. In order to refine the Strisores phylogeny, we generated genome-scale datasets of 2,289 – 4,243 ultra-conserved elements for 23 taxa representing all major living lineages in the group. Among the considerations for using genome-scale, molecular sequence data in phylogenomic analysis are issues related to GC content, GC variance and their effects on model selection. In this study, we employed a variety of analytical techniques to empirically investigate those issues in our data, as well as biases and errors resulting from alignment trimming, taxon selection and matrix completeness. Extensive analyses revealed conflict within the data, especially in regard to variation in GC content, that would not have been detected with more cursory study. Our results indicate that readily available models of molecular evolution are insufficient to encapsulate all phenomena present in genome-scale matrices, and that this problem may be at the root of many current issues in phylogenomic analysis. The analytical methods employed in this study are relevant to phylogenomic analysis of any large, heterogeneous matrix. In conclusion, we present a strongly supported estimate of the Strisores tree and discuss potential evolutionary pathways of nocturnality in this clade.
Figure 1 from: Antoł A, Kozłowski J (2020) Scaling of organ masses in mammals and birds: phylogenetic signal and implications for metabolic rate scaling. ZooKeys 982: 149-159. https://doi.org/10.3897/zookeys.982.55639
Figure 1 PGLS (solid lines) and OLS (dashed lines) interspecific scaling of tissue/organ masses in mammals with log fat-free-body mass as the independent variable. For the scaling with log body mass as an independent variable, see Suppl. material 1: Figure S3
Figure 2 from: Antoł A, Kozłowski J (2020) Scaling of organ masses in mammals and birds: phylogenetic signal and implications for metabolic rate scaling. ZooKeys 982: 149-159. https://doi.org/10.3897/zookeys.982.55639
Figure 2 PGLS (solid lines) and OLS (dashed lines) interspecific scaling of tissue/organ masses in birds with log fat-free-body mass as the independent variable. For the scaling with log body mass as an independent variable, see Suppl. material 1: Figure S5.
Data from: A generalized K statistic for estimating phylogenetic signal from shape and other high-dimensional multivariate data
Phylogenetic signal is the tendency for closely related species to display similar trait values due to their common ancestry. Several methods have been developed for quantifying phylogenetic signal in univariate traits and for sets of traits treated simultaneously, and the statistical properties of these approaches have been extensively studied. However, methods for assessing phylogenetic signal in high-dimensional multivariate traits like shape are less well developed, and their statistical performance is not well characterized. In this article, I describe a generalization of the K statistic of Blomberg et al. (2003) that is useful for quantifying and evaluating phylogenetic signal in highly-dimensional multivariate data. The method (Kmult) is found from the equivalency between statistical methods based on covariance matrices and those based on distance matrices. Using computer simulations based on Brownian motion, I demonstrate that the expected value of Kmult remains at 1.0 as trait variation among species is increased or decreased, and as the number of trait dimensions is increased. By contrast, estimates of phylogenetic signal found with a squared-change parsimony procedure for multivariate data change with increasing trait variation among species and with increasing numbers of trait dimensions, confounding biological interpretations. I also evaluate the statistical performance of hypothesis testing procedures based on Kmult and find that the method displays appropriate Type I error and high statistical power for detecting phylogenetic signal in high-dimensional data. Statistical properties of Kmult were consistent for simulations using bifurcating and random phylogenies, for simulations using different numbers of species, for simulations that varied the number of trait dimensions, and for different underlying models of trait covariance structure. Overall these findings demonstrate that Kmult provides a useful means of evaluating phylogenetic signal in high-dimensional multivariate traits. Finally, I illustrate the utility of the new approach by evaluating the strength of phylogenetic signal for head shape in a lineage of Plethodon salamanders.
Data from: Phylogenetic signal variation in the genomes of Medicago (Fabaceae)
Genome-scale data offer the opportunity to clarify phylogenetic relationships that are difficult to resolve with few loci, but they can also identify genomic regions with evolutionary history distinct from that of the species history. We collected whole-genome sequence data from 29 taxa in the legume genus Medicago, then aligned these sequences to the M. truncatula reference genome to confidently identify 87,596 variable homologous sites. We used this data set to estimate phylogenetic relationships among Medicago species, to investigate the number of sites needed to provide robust phylogenetic estimates, and to identify specific genomic regions supporting topologies in conflict with the genome-wide phylogeny. Our full genomic data set resolves relationships within the genus that were previously intractable. Sub-sampling the data reveals considerable variation in phylogenetic signal and power in smaller subsets of the data. Even when sampling 5,000 sites, no random sample of the data supports a topology identical to that of the genome-wide phylogeny. Phylogenetic relationships estimated from 500-site sliding windows revealed genome regions supporting several alternative species relationships among recently-diverged taxa, consistent with the expected effects of deep coalescence or introgression in the recent history of Medicago.
Data from: Weak phylogenetic signal in physiological traits of methane-oxidizing bacteria
The presence of phylogenetic signal is assumed to be ubiquitous. However, for microorganisms, this may not be true given that they display high physiological flexibility and have fast regeneration. This may result in fundamentally different patterns of resemblance, that is, in variable strength of phylogenetic signal. However, in microbiological inferences, trait similarities and therewith microbial interactions with its environment are mostly assumed to follow evolutionary relatedness. Here, we tested whether indeed a straightforward relationship between relatedness and physiological traits exists for aerobic methane-oxidizing bacteria (MOB). We generated a comprehensive data set that included 30 MOB strains with quantitative physiological trait information. Phylogenetic trees were built from the 16S rRNA gene, a common phylogenetic marker, and the pmoA gene which encodes a subunit of the key enzyme involved in the first step of methane oxidation. We used a Blomberg's K from comparative biology to quantify the strength of phylogenetic signal of physiological traits. Phylogenetic signal was strongest for physiological traits associated with optimal growth pH and temperature indicating that adaptations to habitat are very strongly conserved in MOB. However, those physiological traits that are associated with kinetics of methane oxidation had only weak phylogenetic signals and were more pronounced with the pmoA than with the 16S rRNA gene phylogeny. In conclusion, our results give evidence that approaches based solely on taxonomical information will not yield further advancement on microbial eco-evolutionary interactions with its environment. This is a novel insight on the connection between function and phylogeny within microbes and adds new understanding on the evolution of physiological traits across microbes, plants and animals.
Data from: Phylogenetic signal, feeding behaviour, and brain volume in Neotropical bats
Comparative correlational studies of brain size and ecological traits (e.g. feeding habits and habitat complexity) have increased our knowledge about the selective pressures on brain evolution. Studies conducted in bats as a model system assume that shared evolutionary history has a maximum effect on the traits. However, this effect has not been quantified. In addition, the effect of levels of diet specialization on brain size remains unclear. We examined the role of diet on the evolution of brain size in Mormoopidae and Phyllostomidae using two comparative methods. Body mass explained 89% of the variance in brain volume. The effect of feeding behaviour (either characterized as feeding habits, as levels of specialization on a type of item or as handling behaviour) on brain volume was also significant albeit not consistent after controlling for body mass and the strength of the phylogenetic signal (λ). Although the strength of the phylogenetic signal of brain volume and body mass was high when tested individually, λ values in phylogenetic generalized least squares models were significantly different from 1. This suggests that phylogenetic independent contrasts models are not always the best approach for the study of ecological correlates of brain size in New World bats.
Data from: Phylogenetic signal in mitochondrial and nuclear markers in sea anemones (Cnidaria, Actiniaria)
The mitochondrial genome of basal animals is generally more slowly evolving than that of bilaterians. This difference in rate complicates the study of relationships among members of these lineages and the discovery of cryptic species or the testing of morphological species concepts within them. We explore the properties of mitochondrial and nuclear ribosomal genes in the cnidarian order Actiniaria, using both an ordinal-scale and familyfamilial-scale sample of taxa. Although the markers do not show significant incongruence, they differ in their phylogenetic informativeness and the kinds of relationships they resolve. Among the markers studied here, the fragments of 12S rDNA and 18S rDNA most effectively recover well-supported nodes; those of 16S rDNA and 28S rDNA are less effective. The general patterns we observed are similar to those in other hexacorallians, although Actiniaria alone show saturation of transitions for orderordinal-scale analyses.
Data from: Breakdown of phylogenetic signal: a survey of microsatellite densities in 454 shotgun sequences from 154 non model eukaryote species
Microsatellites are ubiquitous in Eukaryotic genomes. A more complete understanding of their origin and spread can be gained from a comparison of their distribution within a phylogenetic context. Although information for model species is accumulating rapidly, it is insufficient due to a lack of species depth, thus intragroup variation is necessarily ignored. As such, apparent differences between groups may be overinflated and generalizations cannot be inferred until an analysis of the variation that exists within groups has been conducted. In this study, we examined microsatellite coverage and motif patterns from 454 shotgun sequences of 154 Eukaryote species from eight distantly related phyla (Cnidaria, Arthropoda, Onychophora, Bryozoa, Mollusca, Echinodermata, Chordata and Streptophyta) to test if a consistent phylogenetic pattern emerges from the microsatellite composition of these species. It is clear from our results that data from model species provide incomplete information regarding the existing microsatellite variability within the Eukaryotes. A very strong heterogeneity of microsatellite composition was found within most phyla, classes and even orders. Autocorrelation analyses indicated that while microsatellite contents of species within clades more recent than 200 Mya tend to be similar, the autocorrelation breaks down and becomes negative or non-significant with increasing divergence time. Therefore, the age of the taxon seems to be a primary factor in degrading the phylogenetic pattern present among related groups. The most recent classes or orders of Chordates still retain the pattern of their common ancestor. However, within older groups, such as classes of Arthropods, the phylogenetic pattern has been scrambled by the long independent evolution of the lineages.
Data from: Genomic repeat abundances contain phylogenetic signal
A large proportion of genomic information, particularly repetitive elements, is usually ignored when researchers are using next-generation sequencing. Here we demonstrate the usefulness of this repetitive fraction in phylogenetic analyses, utilising comparative graph-based clustering of next-generation sequence reads, which results in abundance estimates of different classes of genomic repeats. Phylogenetic trees are then inferred based on the genome-wide abundance of different repeat types treated as continuously varying characters; such repeats are scattered across chromosomes and in angiosperms can constitute a majority of nuclear genomic DNA. In six diverse examples, five angiosperms and one insect, this method provides generally well-supported relationships at interspecific and intergeneric levels that agree with results from more standard phylogenetic analyses of commonly used markers. We propose that this methodology may prove especially useful in groups where there is little genetic differentiation in standard phylogenetic markers. At the same time as providing data for phylogenetic inference, this method additionally yields a wealth of data for comparative studies of genome evolution.
Supplementary information for: A biased fossil record can preserve reliable phylogenetic signal
<p><span><span><span><span><span><span><span><span><span><span><span><span><i>Abstract.</i><b>––</b>The fossil record is notoriously imperfect and biased in representation, hindering our ability to place fossil specimens into an evolutionary context. For groups with fossil records mostly consisting of disarticulated parts (e.g., vertebrates, echinoderms, plants), the limited morphological information preserved sparks concerns about whether fossils retain reliable evidence of phylogenetic relationships, and lends uncertainty to analyses of diversification, paleobiogeography, and biostratigraphy in Earth history. To address whether a fragmentary past can be trusted, we need to assess whether incompleteness affects the quality of phylogenetic information contained in fossil data. Herein, we characterize skeletal incompleteness bias in a large dataset (6,585 specimens; 14,417 skeletal elements) of fossil squamates (lizards, snakes, amphisbaenians, and mosasaurs). We show that jaws + palatal bones, vertebrae, and ribs appear more frequently in the fossil record than other parts of the skeleton. This incomplete anatomical representation in the fossil record is biased against regions of the skeleton that contain the majority of morphological phylogenetic characters used to assess squamate evolutionary relationships. Despite this bias, parsimony- and model-based comparative analyses indicate that the most frequently-occurring parts of the skeleton in the fossil record retain similar levels of phylogenetic signal as parts of the skeleton that are rarer. These results demonstrate that the biased squamate fossil record contains reliable phylogenetic information, and support our ability to place incomplete fossils in the Tree of Life. </span></span></span></span></span></span></span></span></span></span></span></span></p>
FIGURE 1 in Mitochondrial genome of Poecilimon cretensis (Orthoptera: Tettigoniidae: Phaneropterinae): Strong phylogenetic signals in gene overlapping regions
FIGURE 1. The map of mitochondrial genome and habitus of Poecilimon cretensis
Data from: Phylogenetic signal in mitochondrial and nuclear markers in sea anemones (Cnidaria, Actiniaria)
Open the record for dataset details and reuse information.
Data from: Testing and quantifying phylogenetic signals and homoplasy in morphometric data
Open the record for dataset details and reuse information.
Data from: Components of phylogenetic signal in antagonistic and mutualistic networks
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.