Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “phylogenetic distance”
Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group. in Distribution of extracellular enzyme-producing bacteria in the digestive tracts of 4 brackish water fish species
Figure. The phylogenetic tree showing the relationship among Brevibacillus parabrevis strains SA2.2 and TJ2.3, Bacillus licheniformis MG4.2, and their phylogenetically closest type strains. The GenBank accession numbers of the type strains and studied strains are shown following species names. Distance matrix was calculated by Kimura's 2-parameter model. The scale bar indicates 0.02 substitutions per nucleotide position. Alicyclobacillus pohliae AJ564766 served as an out-group.
Data for publication "Fast and Accurate Distance-based Phylogenetic Placement using Divide and Conquer" (APPLES-2)
<p>Data and scripts used in the paper "Fast and Accurate Distance-based Phylogenetic Placement using Divide and Conquer"</p>
To design, or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe design
<p>This repository contains Materials and designed UCE probe sets for the manuscript entitled "To design or not to design? Comparison of beetle ultraconserved element probe set utility based on phylogenetic distance, breadth, and method of probe design".</p>
Molecular phylogenetic analyses reveal multiple long-distance dispersal events and extensive cryptic speciation in Nervilia (Orchidaceae), an isolated basal Epidendroid genus
Open the record for dataset details and reuse information.
Effects of phylogenetic distance, niche overlap and habitat alteration on spatial co-occurrence patterns in Neotropical bats and birds
Open the record for dataset details and reuse information.
Fast and accurate distance‐based phylogenetic placement using divide and conquer
Open the record for dataset details and reuse information.
Table 2. Pairwise uncorrected p - distances for 16 S in Two new species of gymnophthalmid lizards of the genus Petracola (Squamata: Cercosaurinae) from the Andes of northeastern Peru, and their phylogenetic relationships
<p><b>Table 2.</b> Pairwise uncorrected <i>p</i> -distances for 16S rRNA between <i>Petracola</i> species. The asterisk (*) indicates type locality.</p><table><tbody><tr><th></th><th>1</th><th>2</th><th>3</th><th>4</th><th>5</th><th>6</th><th>7</th><th>8</th><th>9</th><th>10</th></tr></tbody><tbody><tr><th>(1) <i>P. ventrimaculatus</i> CORBIDI 9235</th><td>-</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><th>(2) <i>P. ventrimaculatus</i> KU 219838</th><td>0.024</td><td>-</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><th>(3) <i>P. waka</i> KU 212687</th><td>0.063</td><td>0.071</td><td>-</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><th>(4) <i>P. waka</i> MUBI 2603</th><td>0.073</td><td>0.091</td><td>0.063</td><td>-</td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><th>(5) <i>P. waka</i> MUBI 2605</th><td>0.073</td><td>0.091</td><td>0.063</td><td>0.000</td><td>-</td><td></td><td></td><td></td><td></td><td></td></tr><tr><th>(6) <i>P. waka</i> MUBI 2609*</th><td>0.069</td><td>0.082</td><td>0.066</td><td>0.031</td><td>0.031</td><td>-</td><td></td><td></td><td></td><td></td></tr><tr><th>(7) <i>P. waka</i> MUBI 2611*</th><td>0.069</td><td>0.082</td><td>0.066</td><td>0.031</td><td>0.031</td><td>0.000</td><td>-</td><td></td><td></td><td></td></tr><tr><th>(8) <i>P. shurugojalcapi</i> MUBI 17727</th><td>0.058</td><td>0.074</td><td>0.080</td><td>0.079</td><td>0.079</td><td>0.079</td><td>0.079</td><td>-</td><td></td><td></td></tr><tr><th>(9) <i>P. shurugojalcapi</i> PFAUNA 430</th><td>0.058</td><td>0.074</td><td>0.080</td><td>0.079</td><td>0.079</td><td>0.079</td><td>0.079</td><td>0.000</td><td>-</td><td></td></tr><tr><th>(10) <i>P. amazonensis</i> MUBI 11473</th><td>0.057</td><td>0.072</td><td>0.085</td><td>0.078</td><td>0.078</td><td>0.072</td><td>0.072</td><td>0.037</td><td>0.037</td><td>-</td></tr></tbody></table>
Data from: APPLES: Scalable distance-based phylogenetic placement with or without alignments
<p>Placing a new species on an existing phylogeny has increasing relevance to several applications. Placement can be used to update phylogenies in a scalable fashion and can help identify unknown query samples using (meta-)barcoding, skimming, or metagenomic data. Maximum likelihood (ML) methods of phylogenetic placement exist, but these methods are not scalable to reference trees with many thousands of leaves, limiting their ability to enjoy benefits of dense taxon sampling in modern reference libraries. They also rely on assembled sequences for the reference set and aligned sequences for the query. Thus, ML methods cannot analyze datasets where the reference consists of unassembled reads, a scenario relevant to emerging applications of genome-skimming for sample identification. We introduce APPLES, a distance-based method for phylogenetic placement. Compared to ML, APPLES is an order of magnitude faster and more memory efficient, and unlike ML, it is able to place on large backbone trees (tested for up to 200,000 leaves). We show that using dense references improves accuracy substantially so that APPLES on dense trees is more accurate than ML on sparser trees, where it can run. Finally, APPLES can accurately identify samples without assembled reference or aligned queries using kmer-based distances, a scenario that ML cannot handle. APPLES is available publically at github.com/balabanmetin/apples.</p>
The impacts of fine-tuning, phylogenetic distance, and sample size on big-data bioacoustics
<p>Vocalizations in animals, particularly birds, are critically important behaviors that influence their reproductive fitness. While recordings of bioacoustic data have been captured and stored in collections for decades, the automated extraction of data from these recordings has only recently been facilitated by artificial intelligence methods. These have yet to be evaluated with respect to accuracy of different automation strategies and features. Here, we use a recently published machine learning framework to extract syllables from ten bird species ranging in their phylogenetic relatedness from 1 to 85 million years, to compare how phylogenetic relatedness influences accuracy. We also evaluate the utility of applying trained models to novel species. Our results indicate that model performance is best on conspecifics, with accuracy progressively decreasing as phylogenetic distance increases between taxa. However, we also find that the application of models trained on multiple distantly related species can improve the overall accuracy to levels near that of training and analyzing a model on the same species. When planning big-data bioacoustics studies, care must be taken in sample design to maximize sample size and minimize human labor without sacrificing accuracy.</p>
Data from: Global variation in the relationship between avian phylogenetic diversity and functional distance is driven by environmental context and constraints
<p>Aim: If evolutionary distance is akin to evolutionary chance, then it follows that species assemblages that are distantly related will also be more disparate in terms of their traits, features and the niches they occupy. Yet, studies have found that the total phylogenetic distance of an assemblages, known as phylogenetic diversity, is an unreliable surrogate for functional diversity. We investigate global variation in the relationship between Faith's Phylogenetic Diversity (PD) and Mean Pairwise Functional Distance (MPFD) across latitude and the influence of migratory species on both these aspects of diversity.</p> <p>Location: Global.</p> <p>Time period: Present day.</p> <p>Major taxa studied: Birds.</p> <p>Methods: We measure PD and MPFD for over 9,000 species of bird across more than 17,000 globally distributed assemblages. We obtain standardised effect sizes for both indices by simulating assemblage composition under an ecologically informed null model. We employ path analysis to characterise variation in the relationship between PD's and MPFD across latitude, elevation and with proportion of migratory species.</p> <p>Results: Globally, assemblages that were phylogenetically diverse tended to be less functionally dispersed than expected; however this relationship showed considerable variation across latitude decreasing with distance from the equator. The proportion of migratory species in an assemblage was found to be an important predictor of functional diversity, with migrant rich assemblages generally showing less functional diversity than expected. We identify the Andes and Hengduan Mountains as regions of exceptional bird functional diversity.</p> <p>Main conclusions: The relationship between phylogenetic diversity and function diversity is context specific, varying across environmental gradients such as latitude, and influenced by ecological phenomena such a migration. Thus, care should be taken using phylogenetic diversity as a proxy for functional diversity, particularly in clades with sparse functional data. Instead we recommend that studies consider how phylogenetic diversity's surrogacy for functional diversity may be impacted by environmental context and evaluate empirical observations against biogeographically constrained and ecological informed null models.</p>
Data from: Global variation in the relationship between avian phylogenetic diversity and functional distance is driven by environmental context and constraints
Open the record for dataset details and reuse information.
Data from: APPLES: Scalable distance-based phylogenetic placement with or without alignments
Open the record for dataset details and reuse information.
Data from: Local adaptation, geographical distance and phylogenetic relatedness: assessing the drivers of siderophore-mediated social interactions in natural bacterial communities
Open the record for dataset details and reuse information.
The impacts of fine-tuning, phylogenetic distance, and sample size on big-data bioacoustics
Open the record for dataset details and reuse information.
Taxonomic, functional and phylogenetic beta diversity of upland forest birds in the Amazon: The relative importance of biogeographic regions, climate, and geographic distance
Open the record for dataset details and reuse information.
Pairwise distance demarcation of species in the family Coronaviridae. a, Diagonal matrix of PPDs of 2,505 viruses clustered according to 49 coronavirus species, 39 established and 10 pending or tentative, and ordered from the most to least populous species, from left to right; green and white, PPDs smaller and larger than the inter-species threshold, respectively. Areas of the green squares along the diagonal are proportional to the virus sampling of the respective species, and virus prototypes of the five most sampled species are specified to the left; asterisks indicate species that include viruses whose intra-species PPDs crossed the inter-species threshold (threshold 'violators'). b, Maximal intra-species PPDs (x axis, linear scale) plotted against virus sampling (y axis, log scale) for 49 species (green dots) of the Coronaviridae. Indicated are the acronyms of virus prototypes of the seven most sampled species. Green and blue plot sections represent intra-species and intra-subgenera PPD ranges. The vertical black line indicates the inter-species threshold. c, Shown are the PDs of non-identical residues (y axis) for four viruses representing three major phylogenetic lineages (clades) of the species Severe acute respiratorysyndrome-related coronavirus (panel b) and all pairs of the 256 viruses of this species ('all pairs'). The PD values were derived from pairwise distances in the MSA that were calculated using an identity matrix. Panels a and b were adopted from the DEmARC v.1.4 output. in The species Severe acute respiratory syndromerelated coronavirus: classifying 2019-nCoV and naming it SARS-CoV-2
Pairwise distance demarcation of species in the family Coronaviridae. a, Diagonal matrix of PPDs of 2,505 viruses clustered according to 49 coronavirus species, 39 established and 10 pending or tentative, and ordered from the most to least populous species, from left to right; green and white, PPDs smaller and larger than the inter-species threshold, respectively. Areas of the green squares along the diagonal are proportional to the virus sampling of the respective species, and virus prototypes of the five most sampled species are specified to the left; asterisks indicate species that include viruses whose intra-species PPDs crossed the inter-species threshold (threshold 'violators'). b, Maximal intra-species PPDs (x axis, linear scale) plotted against virus sampling (y axis, log scale) for 49 species (green dots) of the Coronaviridae. Indicated are the acronyms of virus prototypes of the seven most sampled species. Green and blue plot sections represent intra-species and intra-subgenera PPD ranges. The vertical black line indicates the inter-species threshold. c, Shown are the PDs of non-identical residues (y axis) for four viruses representing three major phylogenetic lineages (clades) of the species Severe acute respiratorysyndrome-related coronavirus (panel b) and all pairs of the 256 viruses of this species ('all pairs'). The PD values were derived from pairwise distances in the MSA that were calculated using an identity matrix. Panels a and b were adopted from the DEmARC v.1.4 output.
Data from: Phylogenetic diversity of two geographically overlapping species in the lichen genus Sticta (Ascomycota: Peltigeraceae): isolation by distance, environment, or fragmentation?
<p><span><b>Aim:</b> To test whether the degree of phylogenetic diversity differs in two congeneric, morphologically similar lichens that are both widespread and with a similar geographical range (Neotropics and Hawaii), but differ in altitudinal and habitat preferences, and whether the two species underwent isolation by distance (IBD), environment (IBE), or fragmentation (IBF).</span></p> <p><span><b>Location:</b> South and Central America, Caribbean, Hawaii, Azores.</span></p> <p><span><b>Taxon:</b> <i>Sticta</i> (Peltigeraceae).</span></p> <p><span><b>Methods:</b> Analysis of 395 specimens across the study area; ITS barcoding marker; maximum likelihood tree reconstruction within a broad taxonomic framework; TCS haplotype networks; Mantel test of genetic vs. geographic, environmental, and fragmentation distances; statistical comparison of BIOclim variables.</span></p> <p><span><b>Results:</b><b> </b><i>Sticta andina</i> exhibited high phenotypic variation and high reticulate phylogenetic diversity across its range, whereas the phenotypically more uniform <i>S. scabrosa</i> contained two main haplotypes, one unique to Hawaii (subsp. <i>hawaiiensis</i>). <i>Sticta andina</i> was restricted to well-preserved andine forests and paramos, habitats fragmented due to disruptive topology, whereas <i>S. scabrosa</i> was found in lowland to lower montane forests in rather exposed microsites, representing a more continuous habitat. These differences were statistically significant for several BIOclim variables. Mantel tests on genetic vs. geographic and environmental distances demonstrated that <i>S. scabrosa</i> followed a pattern of IBD across its full range but not within continental Central and South America. In contrast, <i>S. andina</i> did not exhibit IBD but showed weak, yet significant patterns of IBE at continental level and IBF in the northern Andes.</span></p> <p><b>Main Conclusions:</b> Autecology indirectly drives phylogenetic diversity in the two studied species. In the low altitude species, <i>S. scabrosa</i>, phylogenetic diversity is low and shows no correlation with geographic or environmental distances, except for the differentiation of the Hawaiian subspecies. We attribute this to rapid expansion and effective gene flow between populations across a more or less continuously distributed niche representing partially exposed microsites, including disturbed and anthropogenic vegetation, such as planted trees. In contrast, in the high altitude species, <i>S. andina</i>, phylogenetic diversity is high and correlated with both environmental niche differentiation (IBE) and fragmentation caused by the final Andean uplift (IBF). Therefore, an autoecological preference for high altitudes increases the likelihood for higher phylogenetic diversity.</p>
Data from: Untangling phylogenetical, geometrical and ornamental imprints on Early Triassic ammonoid biogeography: a similarity-distance decay study.
Ammonoids are diverse and widespread fossil shelly cephalopods that flourished in the world ocean during more than 300 million years before their total extinction, 65 million years ago. In spite of two centuries of intensive scientific studies, their mode(s) of life, and most particularly long-distance dispersal abilities remain poorly known. Here we address this question by looking at the latitudinal distribution of Early Triassic (~250 Myr) ammonoids through similarity-distance decay analyses. We examine and compare how rates of similarity-distance decay differ between various systematic, shell geometry and ornamentation groups, during the same ~3.5 myr Early Triassic time interval, in order to untangle phylogenetical, geometrical and ornamental imprints on the observed biogeographical pattern. Our data do not support any phylogenetical and shell ornamentation control on the similarity-distance decay, but rather evidence a significant effect of (sub-)adult shell geometry: most evolute morphs tend to have been more endemic than most involute ones. This result contrasts with the classical hypothesis that long-distance ammonoid dispersal mainly occurred during the earliest planktonic young stages, and thus that (sub-)adult morphological characteristics should not constrain large-scale biogeographical patterns of ammonoids. While a direct control by Sea Surface Temperature can be discarded, this result may indicate that at least some adult Triassic ammonoid morphs were good active swimmers able to achieve long-distance migrations, as observed for some present-day coleoid cephalopods.
A global analysis of mosses reveals low phylogenetic endemism and highlights the importance of long-distance dispersal
<p><span><span><span><span><span><span><span><span><span><span><span><u>Aim:</u> </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>Digitization of herbarium specimens and DNA sequencing efforts in the past decade have enabled integrative analyses of patterns of diversity and endemism in a phylogenetic context. Here, we compare the best available floristic databases to a comprehensive specimen database to examine spatial patterns of moss phylogenetic assembly. We test the hypotheses that 1) mosses exhibit phylogenetic regionalization, 2) islands contain significantly high phylogenetic diversity, and 3) that moss phylogenetic endemism is low on a global scale.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><u>Location:</u> </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>Global</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><u>Taxon:</u> </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>Mosses</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><u>Methods:</u> </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>We developed a phylogeny of 3,654 moss species using 25 markers and compiled a global specimen database from online repositories. We calculated floristic and phylogenetic measures of diversity and endemism and performed randomizations to test for significant deviations from expectations. We use rarefaction and extrapolation to alleviate substantial differences in sampling effort across the globe. We used both phylogenetic and floristic methods to test for spatial regionalization. We compare our specimen-based results to those obtained using a floristic dataset. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><u>Results:</u></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span> Phylogenetic diversity is more robust to missing data than species richness. Mean phylogenetic distance was significantly higher than expected in areas with high species richness, indicating that reported richness in these areas is likely a product of repeated colonization. Phylogenetic endemism is low globally. Phylogenetic regionalizations cluster into a Holarctic/Holantarctic temperate region, a pantropical region, and a region composed of Australia, New Zealand, and South Africa.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><u>Main Conclusions:</u></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span> Future efforts for collecting, sequencing, and databasing moss species should focus on the tropics, particularly Africa and Southeast Asia. We provide further evidence to support several important theories developed in moss biogeography, including the role of long-distance dispersal in shaping floristic patterns, the dominance of anagenesis in driving patterns of island diversity, and the role of climatic instability in driving patterns of assembly in the Holarctic.</span></span></span></span></span></span></span></span></span></span></span></p>
Lpnet: Reconstructing phylogenetic networks from distances using integer linear programming
<p>We present Lpnet, a variant of the widely used Neighbor-net method that approximates pairwise distances between taxa by a circular phylogenetic network. We first apply standard methods to construct a binary phylogenetic tree and then use integer linear programming to compute optimal circular orderings that agree with all tree splits. This approach achieves an improved approximation of the input distance for the clear majority of experiments that we have run for simulated and real data. We release an implementation in R that can handle up to 94 taxa and usually needs about one minute on a standard computer for 80 taxa. For larger taxa sets, we include a top-down heuristic which also tends to perform better than Neighbor-net.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.