Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “phylogenetic distance”
Fig. 5 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 5. Phylogenetic trees of Scarabaeidae generated using UCEs; node values indicate bootstrap support. A)The tree produced with the Scarab 3kv1 probe set. B) The topology produced with the Adephaga 2.9kv1 probe set.
Fig. 1 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 1. Phylogenetic relationships based on McKenna et al. (2019) among select Coleoptera taxa relevant to or included in UCE probe design. Color (online),
Fig. 3 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 3. UCE loci recovery during in silico testing plotted against different metrics: A) average genetic distance estimated based on common gene fragments used in phylogenetics; B) average genetic distance estimated using BUSCO genes; C) N50 assembly metrics; and D) BUSCO S values.
Fig. 2 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 2. Schematical overview of the workflow for the present study.Workflow proceeds from left to right and top to bottom.The tablets correspond to the three broader segments of the study: genomic resource generation, probe design, and in silico testing; within tablet boundaries can be found associated taxon sets, data, experimentation, and results. Arrows indicate the flow of data, associated results, and the location results can ultimately be found. Color corresponds to membership within a taxon set or probe set.The red boxes around Hydro 2.7kv1 and Scarab 3kv1 denote final optimized, tailored probe set design tailored for Hydrophiloidea and Scarabaeidae based on the results of this study. Length of 75CM box corresponds to alignment length. Abbreviations used: NCBI, National Center for Biotechnology Information; 75CM, 75% complete matrix; AMAS, alignment manipulation and summary statistics (Borowiec 2016); R-F, Robinson– Foulds distance (Robinson and Foulds 1981).
Data from: Untangling phylogenetical, geometrical and ornamental imprints on Early Triassic ammonoid biogeography: a similarity-distance decay study.
Open the record for dataset details and reuse information.
Data from: Integrating phylogenetic and ecological distances reveals new insights into parasite host specificity
Open the record for dataset details and reuse information.
Lpnet: Reconstructing phylogenetic networks from distances using integer linear programming
Open the record for dataset details and reuse information.
A global analysis of mosses reveals low phylogenetic endemism and highlights the importance of long-distance dispersal
Open the record for dataset details and reuse information.
Data from: Phylogenetic diversity of two geographically overlapping species in the lichen genus Sticta (Ascomycota: Peltigeraceae): Isolation by distance, environment, or fragmentation?
Open the record for dataset details and reuse information.
Data from: Minimizing the average distance to a closest leaf in a phylogenetic tree
When performing an analysis on a collection of molecular sequences, it can be convenient to reduce the number of sequences under consideration while maintaining some characteristic of a larger collection of sequences. For example, one may wish to select a subset of high-quality sequences that represent the diversity of a larger collection of sequences. One may also wish to specialize a large database of characterized "reference sequences" to a smaller subset that is as close as possible on average to a collection of "query sequences" of interest. Such a representative subset can be useful whenever one wishes to find a set of reference sequences that is appropriate to use for comparative analysis of environmentally-derived sequences, such as for selecting "reference tree" sequences for phylogenetic placement of metagenomic reads. In this paper we formalize these problems in terms of the minimization of the Average Distance to the Closest Leaf (ADCL) and investigate algorithms to perform the relevant minimization. We show that the greedy algorithm is not effective, show that a variant of the Partitioning Among Medoids (PAM) heuristic gets stuck in local minima, and develop an exact dynamic programming approach. Using this exact program we note that the performance of PAM appears to be good for simulated trees, and is faster than the exact algorithm for small trees. On the other hand, the exact program gives solutions for all numbers of leaves less than or equal to the given desired number of leaves, while PAM only gives a solution for the pre-specified number of leaves. Via application to real data, we show that the ADCL criterion chooses chimeric sequences less often than random subsets, while the maximization of phylogenetic diversity chooses them more often than random. These algorithms have been implemented in publicly available software.
Data from: Phylogenetic tree estimation with and without alignment: new distance methods and benchmarking
Phylogenetic tree inference is a critical component of many systematic and evolutionary studies. The majority of these studies are based on the two-step process of multiple sequence alignment followed by tree inference, despite persistent evidence that the alignment step can lead to biased results. Here we present a two-part study that first presents PaHMM-Tree, a novel neighbour joining-based method that estimates pairwise distances without assuming a single alignment. We then use simulations to benchmark its performance against a wide-range of other phylogenetic tree inference methods, including the first comparison of alignment-free distance-based methods against more conventional tree estimation methods. Our new method for calculating pairwise distances based on statistical alignment provides distance estimates that are as accurate as those obtained using standard methods based on the true alignment. Pairwise distance estimates based on the two-step process tend to be substantially less accurate. This improved performance carries through to tree inference, where PaHMM-Tree provides more accurate tree estimates than all of the pairwise distance methods assessed. For close to moderately divergent sequence data we find that the two-step methods using statistical inference, where information from all sequences is included in the estimation procedure, tend to perform better than PaHMM-Tree, particularly full statistical alignment, which simultaneously estimates both the tree and the alignment. For deep divergences we find the alignment step becomes so prone to error that our distance-based PaHMM-Tree outperforms all other methods of tree inference. Finally, we find that the accuracy of alignment-free methods tends to decline faster than standard two-step methods in the presence of alignment uncertainty, and identify no conditions where alignment-free methods are equal to or more accurate than standard phylogenetic methods even in the presence of substantial alignment error.
Data from: Ecosystem productivity is associated to bacterial phylogenetic distance in surface marine waters
Understanding the link between community diversity and ecosystem function is a fundamental aspect of ecology. Systematic losses in biodiversity are widely acknowledged but the impact this may exert on ecosystem functioning remains ambiguous. There is growing evidence of a positive relationship between species richness and ecosystem productivity for terrestrial macroorganisms, but similar links for marine microorganisms, which help drive global climate, are unclear. Community manipulation experiments show both positive and negative relationships for microbes. These previous studies rely, however, on artificial communities and any links between the full diversity of active bacterial communities in the environment, their phylogenetic relatedness, and ecosystem function remains hitherto unexplored. Here we test the hypothesis that productivity is associated to diversity in the metabolically active fraction of microbial communities. We show in natural assemblages of active bacteria that communities containing more distantly related members were associated with higher bacterial production. The positive phylogenetic diversity–productivity relationship was independent of community diversity calculated as the Shannon index. From our long-term (7-year) survey of surface marine bacterial communities we also found that similarly productive communities had greater phylogenetic similarity to each other, further suggesting that the traits of active bacteria are an important predictor of ecosystem productivity. Our findings demonstrate that the evolutionary history of the active fraction of a microbial community is critical for understanding their role in ecosystem functioning.
Data from: Non-linear effects of phylogenetic distance on early-stage establishment of experimentally introduced plants in grassland communities
1. The phylogenetic distance of an introduced plant species to a resident native community may play a role in determining its establishment success. While Darwin's naturalization hypothesis predicts a positive relationship, the preadaptation hypothesis predicts a negative relationship. Rigorous tests of this now so-called Darwin's naturalization conundrum require not only information on establishment successes but also of failures, which is frequently not available. Such essential information, however, can be provided by experimental introductions. 2. Here, we analysed three datasets from two field experiments in Germany and Switzerland. In the Swiss experiment, alien and native grassland species were introduced as seeds only with and without disturbance (tilling). In the German experiment, alien and native grassland species were introduced both as seeds and as seedlings with and without disturbance (tilling), and with and without fungicide application. For the seedling introduction experiment, there was an additional herbivore-exclusion treatment. 3. Phylogenetic distance affected establishment in the three datasets differently, with success peaking at intermediate distances for the seed datasets, but decreasing with increasing distances in the seedling dataset. Disturbance favored seedling survival, most likely by weakening the resident community. 4. Synthesis: By analyzing experimental introductions, we show that the relationship between phylogenetic distance and establishment, at least for seedling emergence, may actually be non-linear with an optimum at intermediate distances. Therefore, Darwin´s naturalization hypothesis and the preadaptation hypothesis need not be in conflict. Rather, the mechanisms underlying them can operate simultaneously or alternately depending on the life stage and on the environmental conditions of the resident community.
Fig. 4 in To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe
Fig. 4. The proportion of different types of loci targeted by scarab and hydrophiloid probe sets.
Data from: Ectoparasite fitness in auxiliary hosts: Phylogenetic distance from a principal host matters
Open the record for dataset details and reuse information.
Data from: Ecosystem productivity is associated to bacterial phylogenetic distance in surface marine waters
Open the record for dataset details and reuse information.
Data from: Likelihood-based parameter estimation for high-dimensional phylogenetic comparative models: overcoming the limitations of 'distance-based' methods
Open the record for dataset details and reuse information.
Data from: Minimizing the average distance to a closest leaf in a phylogenetic tree
Open the record for dataset details and reuse information.
Data from: Phylogenetic tree estimation with and without alignment: new distance methods and benchmarking
Open the record for dataset details and reuse information.
Data from: Non-linear effects of phylogenetic distance on early-stage establishment of experimentally introduced plants in grassland communities
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.