Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

44

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

44 results for “species tree estimation”

Learn how ShareScore rates datasets ↗
zenodo48/100

Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024

<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (&ldquo;full&rdquo; matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (&ldquo;reduced&rdquo; matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the &ldquo;new technology&rdquo; search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
edi44/100

Dalton and Nenana study site data including: invasive plant density estimates, invasive plant density, soil data, seedling estimates for dominant tree species and ground cover estimates for sites

This dataset contains invasive plant and stand level data for study sites along the Dalton and Parks highways in interior Alaska in the summer of 2012. Study sites were situated in burned and mature black spruce forests to compare invasive plant colonization patterns. Invasive plant density estimates along the road adjacent to each site are included, as well as invasive plant density within study sites. Other data includes ground cover estimates for dominant ground cover types, estimates of seedling abundance for dominant tree species, soil paramters (mineral soil pH and mineral soil moisture, residual organic layer/ organic layer depths, and active layer depths).

openOpenFeb 2016View details →
dryad40/100

The implications of incongruence between gene tree and species tree topologies for divergence time estimation

<p>Phylogenetic analyses are increasingly being performed with datasets that incorporate hundreds of loci. Due to incomplete lineage sorting, hybridization, and horizontal gene transfer, the gene trees for these loci may often have topologies that differ from each other and from the species tree. The effect of these topological incongruences on divergence time estimation has not been fully investigated. Using a series of simulation experiments and empirical analyses, we demonstrate that when topological incongruence between gene trees and the species tree is not accounted for, the temporal duration of branches in regions of the species tree that are affected by incongruence is underestimated, whilst the duration of other branches is considerably overestimated. This effect becomes more pronounced with higher levels of topological incongruence. We show that this pattern results from erroneous estimation of the number of substitutions along branches in the species tree, although the effect is modulated by the assumptions inherent to divergence time estimation, such as those relating to the fossil record or among-branch-substitution-rate variation. By only analysing loci with gene trees that are topologically congruent with the species tree, or only taking into account the branches from each gene tree that are topologically congruent with species tree, we demonstrate that the effects of topological incongruence can be ameliorated. Nonetheless, even when topologically congruent gene trees or topologically congruent branches are selected, error in divergence time estimates remains. This stems from temporal incongruences between divergence times in species trees and divergence times in gene trees, and more importantly, the difficulty of incorporating necessary assumptions for divergence time estimation.</p>

opencc-zeroMar 2022View details →
zenodo40/100

Polymorphism-aware estimation of species trees and evolutionary forces from genomic sequences with RevBayes

<p>Supplementary files of Polymorphism-aware estimation of species trees and evolutionary forces from genomic sequences with RevBayes by Borges, Boussau, H&ouml;hna, Pereira and Kosiol<br> &nbsp;</p>

opencc-by-4.0May 2022View details →
dryad40/100

Data from: Improving quartet graph construction for scalable and accurate species tree estimation from gene trees

<p>Summary methods are one of the dominant approaches for estimating species trees from genome-scale data. However, they can fail to produce accurate species trees when the input gene trees are highly discordant due to gene tree estimation error as well as biological processes, like incomplete lineage sorting. Here, we introduce a new summary method TREE-QMC that offers improved accuracy and scalability under these challenging scenarios. TREE-QMC builds upon the algorithmic framework of QMC (Snir and Rao 2010) and its weighted version wQMC (Avni et al. 2014). Their approach takes weighted quartets (four-leaf trees) as input and builds a species tree in a divide-and-conquer fashion, at each step constructing a graph and seeking its max cut. We improve upon this methodology in two ways. First, we address scalability by providing an algorithm to construct the graph directly from the input gene trees. By skipping the quartet weighting step, TREE-QMC has a time complexity of O(n^3 k) with some assumptions on subproblem sizes, where n is the number of species and k is the number of gene trees. Second, we address accuracy by normalizing the quartet weights to account for "artificial taxa," which are introduced during the divide phase so that solutions on subproblems can be combined during the conquer phase. Together, these contributions enable TREE-QMC to outperform the leading methods (ASTRAL-III, FASTRAL, wQFM) in an extensive simulation study. We also present the application of these methods to an avian phylogenomics data set.</p>

opencc-zeroJul 2022View details →
dryad40/100

Data from: Species tree branch length estimation despite incomplete lineage sorting, duplication, and loss

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad40/100

Data from: Improving quartet graph construction for scalable and accurate species tree estimation from gene trees

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad40/100

The implications of incongruence between gene tree and species tree topologies for divergence time estimation

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad40/100

ASTRAL-II: coalescent-based species tree estimation with many hundreds of taxa and thousands of genes

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad40/100

QuCo: quartet-based co-estimation of species trees and gene trees

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad36/100

Data from: ASTRAL: genome-scale coalescent-based species tree estimation

<p>Species trees provide insight into basic biology, including the mechanisms of evolution and how it modifies biomolecular function and structure, biodiversity and co-evolution between genes and species. Yet, gene trees often differ from species trees, creating challenges to species tree estimation. One of the most frequent causes for conflicting topologies between gene trees and species trees is incomplete lineage sorting (ILS), which is modelled by the multi-species coalescent. While many methods have been developed to estimate species trees from multiple genes, some which have statistical guarantees under the multi-species coalescent model, existing methods are too computationally intensive for use with genome-scale analyses or have been shown to have poor accuracy under some realistic conditions.</p> <p>Results: We present ASTRAL, a fast method for estimating species trees from multiple genes. ASTRAL is statistically consistent, can run on datasets with thousands of genes and has outstanding accuracy—improving on MP-EST and the population tree from BUCKy, two statistically consistent leading coalescent-based methods. ASTRAL is often more accurate than concatenation using maximum likelihood, except when ILS levels are low or there are too few gene trees.</p>

opencc-zeroJan 2024View details →
dryad36/100

Data for: Theoretical and practical considerations when using retroelement insertions to estimate species trees in the anomaly zone

<p>A potential shortcoming of concatenation methods for species tree estimation is their failure to account for incomplete lineage sorting. Coalescent methods address this problem but make various assumptions that, if violated, can result in worse performance than concatenation. Given the challenges of analyzing DNA sequences with both concatenation and coalescent methods, retroelement insertions (RIs) have emerged as powerful phylogenomic markers for species tree estimation. Here, we show that two recently proposed quartet-based methods, SDPquartets and ASTRAL_BP, are statistically consistent estimators of the unrooted species tree topology under the coalescent when RIs follow a neutral infinite-sites model of mutation and the expected number of new RIs per generation is constant across the species tree. The accuracy of these (and other) methods for inferring species trees from RIs has yet to be assessed on simulated data sets, where the true species tree topology is known. Therefore, we evaluated eight methods given RIs simulated from four model species trees, all of which have short branches and at least three of which are in the anomaly zone. In our simulation study, ASTRAL_BP and SDPquartets always recovered the correct species tree topology when given a sufficiently large number of RIs, as predicted. A distance-based method (ASTRID_BP) and Dollo parsimony also performed well in recovering the species tree topology. In contrast, unordered, polymorphism, and Camin-Sokal parsimony (as well as an approach based on MDC) typically fail to recover the correct species tree topology in anomaly zone situations with more than four ingroup taxa. Of the methods studied, only ASTRAL_BP automatically estimates internal branch lengths (in coalescent units) and support values (i.e., local posterior probabilities). We examined the accuracy of branch length estimation, finding that estimated lengths were accurate for short branches but upwardly biased otherwise. This led us to derive the maximum likelihood (branch length) estimate for when RIs are given as input instead of binary gene trees; this corrected formula produced accurate estimates of branch lengths in our simulation study, provided that a sufficiently large number of RIs were given as input. Lastly, we evaluated the impact of data quantity on species tree estimation by repeating the above experiments with input sizes varying from 100 to 100,000 parsimony-informative RIs. We found that, when given just 1,000 parsimony-informative RIs as input, ASTRAL_BP successfully reconstructed major clades (i.e clades separated by branches &gt;0.3 CUs) with high support and identified rapid radiations (i.e., shorter connected branches), although not their precise branching order. The local posterior probability was effective for controlling false positive branches in these scenarios.</p>

opencc-zeroNov 2021View details →
dryad36/100

Data from: Species tree estimation and the impact of gene loss following whole-genome duplication

<p>Whole-genome duplication (WGD) has been demonstrated to occur broadly and repeatedly in the evolutionary history of eukaryotes, and is recognized as a prominent evolutionary force, especially in plants. Immediately following WGD, most genes are present in two copies as paralogs. Due to this redundancy, one copy of a paralog pair commonly undergoes pseudogenization and is eventually lost. When speciation occurs shortly after WGD, however, differential loss of paralogs may lead to spurious phylogenetic inference resulting from the inclusion of pseudoorthologs – paralogous genes mistakenly identify as orthologs because they are present in single copes within each sampled species. The influence and impact of including pseudoorthologs versus true orthologs as result of gene extinction (or incomplete laboratory sampling) in a phylogenetic context is only recently starting to gain empirical attention. Moreover, few of these studies have yet to investigate this phenomenon in an explicit coalescent framework. Here, using mathematical models, numerous simulated data sets, and two newly assembled empirical data sets, we assess the effect of pseudoorthologs on species tree estimation under varying levels of incomplete lineage sorting (ILS) and different patterns of gene loss following WGD. When gene loss occurs in the terminal branches of the species tree, the alignment-based (BPP) and gene-tree-based (ASTRAL, MP-EST, and STAR) coalescent methods are adversely affected as the level of ILS increases. This can be greatly improved by sampling a sufficiently large number of genes. Under the same circumstances, however, concatenation methods consistently estimate incorrect species trees as the number of sampled genes increases. Furthermore, pseudoorthologs can mislead species tree inference if gene loss occurs in the internal branches of the species tree, where both coalescent and concatenation methods are prone to produce inconsistent results. However, pseudoorthologs are problematic when filtering only for single-copy genes in phylogenomic data sets. Pruning orthologs or even randomly selecting a copy from multi-copy genes can avoid most of those pseudoorthologs. These results underscore the importance of understanding the influence of pseudoorthologs in the phylogenomics era.</p>

opencc-zeroJun 2022View details →
dryad36/100

Supplementary material for: Impact of ghost introgression on coalescent-based species tree inference and estimation of divergence time

<p><span>The species studied in any evolutionary investigation generally constitute a small proportion of all the species currently existing or that have gone extinct. It is therefore likely that introgression, which is widespread across the tree of life, involves "ghosts," i.e., unsampled, unknown, or extinct lineages. However, the impact of ghost introgression on estimations of species trees has rarely been studied and is poorly understood. Here, we use mathematical analysis and simulations to examine the robustness of species tree methods based on the multispecies coalescent model to introgression from a ghost or extant lineage. We found that many results originally obtained for introgression between extant species can easily be extended to ghost introgression, such as the strongly interactive effects of incomplete lineage sorting (ILS) and introgression on the occurrence of anomalous gene trees (AGTs). The relative performance of the summary species tree method (ASTRAL) and the full-likelihood method (*BEAST) varies under different introgression scenarios, with the former being more robust to gene flow between non-sister species whereas the latter performing better under certain conditions of ghost introgression. When an outgroup ghost (defined as a lineage that diverged before the most basal species under investigation) acts as the donor of the introgressed genes, the time of root divergence among the investigated species generally was overestimated, whereas ingroup introgression, as commonly perceived, can only lead to underestimation. In many cases of ingroup introgression that may or may not involve ghost lineages, the stronger the ILS, the higher the accuracy achieved in estimating the time of root divergence, although the topology of the species tree is more prone to be biased by the effect of introgression.</span></p>

opencc-zeroJun 2022View details →
dryad36/100

Data from: Species tree estimation using ddRADseq data from historical specimens confirms the monophyly of highly disjunct species of Chloropyron (Orobanchaceae)

Sequence data exist for only about 1/5 of plant species; therefore we are at risk of losing many branches of the tree of life even before they are placed into a molecular evolutionary context. This necessitates methods for phylogeny estimation of understudied, rare, and threatened taxa, which often forces researchers to utilize historical collections. The restriction site-associated DNA sequencing (RADseq) family of reduced representation sequence generation has provided a flexible and efficient method for the rapid generation of hundreds to tens of thousands of loci, and has recently seen adoption for phylogeny estimation. However, these methods have been primarily utilized with freshly collected or well preserved tissue. Here we sample all taxa of a genus of rare flowering plants, Chloropyron (Orobanchaceae), from herbarium sheets dating up to 25 yr and use double digest restriction site-associated DNA sequencing (ddRADseq) to resolve intraspecific relationships. We find all species in Chloropyron to be monophyletic, with the inland taxon C. maritimum ssp. canescens sister to the rest of the coastal C. maritimum (ssp. maritimum + ssp. palustre), and the two distinct subspecies of C. molle to be each other's closest relative with strong support. In addition, we demonstrate the utility of reduced representation libraries to address phylogenomic problems in a group of rare species and address pitfalls of accurately inferring relationships when the amount of missing data is large, as is often the case when using historical specimens and rare taxa.

opencc-zeroDec 2017View details →
zenodo36/100

Parameter estimation and species tree rooting using ALE and GeneRax

<p>Some output files associated with the analyses in an article (currently) entitled &quot;Parameter estimation and species tree rooting using ALE and GeneRax&quot;</p> <p>The outputs were obtained by running ALEml_undated either with default (maximum likelihood) estimation of the delta, tau, and lambda parameters, or with delta:tau fixed to a pre-defined ratio. The option to fix delta:tau (or other parameter combinations) was recently added to ALE (https://github.com/ssolo/ALE).</p> <p>The .ale files used as input for these analyses (for the bacterial dataset) are available from the data repository for the original 2021 paper here:&nbsp;&nbsp;https://doi.org/10.6084/m9.figshare.12651074.v12</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

Data for: Theoretical and practical considerations when using retroelement insertions to estimate species trees in the anomaly zone

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad36/100

Data from: Species tree estimation using ddRADseq data from historical specimens confirms the monophyly of highly disjunct species of Chloropyron (Orobanchaceae)

Open the record for dataset details and reuse information.

publicJul 2019View details →
dryad36/100

Data from: ASTRAL: genome-scale coalescent-based species tree estimation

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad36/100

Supplementary material for: Impact of ghost introgression on coalescent-based species tree inference and estimation of divergence time

Open the record for dataset details and reuse information.

publicJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record