Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.7.1
Dataset results
23 results for “Supermatrix”
Fig. 1 in A novel supermatrix approach improves resolution of phylogenetic relationships in a comprehensive sample of danthonioid grasses
Fig. 1. Geologic map of Florissant Fossil Beds National Monument in central Colorado, USA, modified from Evanoff et al. (2001: fig. 1). Areal extent of the Monument is outlined by thick gray line.
Fig. 2 in A novel supermatrix approach improves resolution of phylogenetic relationships in a comprehensive sample of danthonioid grasses
Fig. 2. Talpid mammal Oreotalpa florissantensis gen. et sp. nov., FLFO 5813 (holotype), right dentary with m1–m3 from UCM locality 92179, Florissant Formation, Florissant Fossil Beds National Monument, Colorado, USA; latest Eocene (Chadronian). SEM micrographs; in lingual (A), labial (B), and occlusal (C) views, and explanatory drawing of occlusal view (D). Anterior is to the right. Original drawing by Leigh Anne McConnaughey. For B, C, and D, anterior is to the right.
A supermatrix phylogeny of the world's bees (Hymenoptera: Anthophila)
<p>The increasing availability of large phylogenies based on genetic data has provided new opportunities to study the evolution of species traits, their origins and diversification, and biogeography; yet there are limited attempts to synthesise existing phylogenetic information for major insect groups. Bees (Hymenoptera: Anthophila) are a large group of insect pollinators that have a worldwide distribution, and a wide variation in ecology, morphology, and life-history traits, including sociality. For these reasons, as well as their major economic importance as pollinators, numerous molecular phylogenetic studies of family and genus-level relationships have been published, providing an opportunity to assemble a bee 'tree-of-life'. We used publicly available genetic sequence data, including phylogenomic data, reconciled to a taxonomic database, to produce a concatenated supermatrix phylogeny for the Anthophila comprising 4,586 bee species, representing 23% of species and 82% of genera. At family, subfamily, and tribe levels, support for expected relationships was robust, but between and within some genera, relationships remain uncertain. Within families, sampling of genera ranged from 67–100% but species coverage was lower (17–41%). Our phylogeny mostly reproduces the relationships found in recent phylogenomic studies with a few exceptions. We provide a summary of these differences and the current state of molecular data available and its gaps. We discuss the advantages and limitations of this bee supermatrix phylogeny (available online at beetreeoflife.org), which may enable new insights into long-standing questions about evolutionary drivers in bees, and potentially insects.</p>
A supermatrix phylogeny of the world’s bees (Hymenoptera: Anthophila)
Open the record for dataset details and reuse information.
Genome-wide supermatrix analyses of maples (Acer, Sapindaceae) reveal recurring inter-continental migration, mass extinction, and rapid lineage divergence
Open the record for dataset details and reuse information.
Dipsacales four locus supermatrix assembled with PyNCBIminer
Open the record for dataset details and reuse information.
Data from: A supermatrix phylogeny of corvoid passerine birds (Aves: Corvides)
The Corvides (previously referred to as the core Corvoidea) are a morphologically diverse clade of passerine birds comprising nearly 800 species. The group originated some 30 million years ago in the proto-Papuan archipelago, to the north of Australia, from where lineages have dispersed and colonized all of the world's major continental and insular landmasses (except Antarctica). During the last decade multiple species-level phylogenies have been generated for individual corvoid families and more recently the inter-familial relationships have been resolved, based on phylogenetic analyses using multiple nuclear loci. In the current study we analyse eight nuclear and four mitochondrial loci to generate a dated phylogeny for the majority of corvoid species. This phylogeny includes 667 out of 780 species (85.5%), 141 out of 143 genera (98.6%) and all 31 currently recognized families, thus providing a baseline for comprehensive macroecological, macroevolutionary and biogeographical analyses. Using this phylogeny we assess the temporal consistency of the current taxonomic classification of families and genera. By adopting an approach that enforces temporal consistency by causing the fewest possible taxonomic changes to currently recognized families and genera, we find the current familial classification to be largely temporally consistent, whereas that of genera is not.
Data from: Bayesian analysis of a morphological supermatrix sheds light on controversial fossil hominin relationships
The phylogenetic relationships of several hominin species remain controversial. Two methodological issues contribute to the uncertainty—use of partial, inconsistent datasets and reliance on phylogenetic methods that are ill-suited to testing competing hypotheses. Here, we report a study designed to overcome these issues. We first compiled a supermatrix of craniodental characters for all widely accepted hominin species. We then took advantage of recently developed Bayesian methods for building trees of serially sampled tips to test among hypotheses that have been put forward in three of the most important current debates in hominin phylogenetics—the relationship between Australopithecus sediba and Homo, the taxonomic status of the Dmanisi hominins, and the place of the so-called hobbit fossils from Flores, Indonesia, in the hominin tree. Based on our results, several published hypotheses can be statistically rejected. For example, the data do not support the claim that Dmanisi hominins and all other early Homo specimens represent a single species, nor that the hobbit fossils are the remains of small-bodied modern humans, one of whom had Down syndrome. More broadly, our study provides a new baseline dataset for future work on hominin phylogeny and illustrates the promise of Bayesian approaches for understanding hominin phylogenetic relationships.
Fig. 4 in Pitfalls in supermatrix phylogenomics
Fig. 4. Distribution of missing data in phylogenomic datasets. The raw dimensions of a supermatrix (in numbers of genes and taxa) should always be accompanied by an estimate of its level (and ideally distribution) of missing data (A). Indeed, some datasets were advertised as very large when published but required a lot of useless computing power due to a large proportion of missing data (B). More importantly, such datasets are more exposed to data errors, systematic error and phylogenetic artefacts, as their effective number of species is actually quite low for the largest part of their width. In other cases, a targeted completion (in the outgroups) was enough to improve the phylogenetic accuracy (D–E). In contrast, datasets assembled with the optimal phylogenomic zone in mind combine a suffcient number of genes and taxa while featuring a low proportion of missing data (C). These are the most appropriate datasets to produce accurate phylogenetic relationships.
Fig. 2 in Pitfalls in supermatrix phylogenomics
Fig. 2. Example of a frameshift-aware alignment produced by MACSE. Trpc2(-like) sequences of bats were aligned at the nucleotide level (A) and amino acid level (B), unravelling several frameshifts (indicated by "!") and stop codons (indicated by "*") in the pseudogenes. Hence, MACSE automatically provided an alignment that would otherwise require a lot of tedious manual work. It can also be used to replace these frameshifts and stop codons by standard codons (e.g., "NNN" or "---") in order to obtain an alignment suitable for further analysis with standard tools (e.g., PhyloBayes or PAML).
Fig. 1 in Pitfalls in supermatrix phylogenomics
Fig. 1. Evolution of phylogenetic reconstruction over time. Different zones can be delimited based on the number of genes (X axis) and number of taxa (Y axis) composing the supermatrix at hand. Generally speaking, the (upper) left part of the map has more to do with identifying species (as in barcoding studies), while the right part of the map corresponds to multigene and phylogenomic datasets assembled for recovering large-scale phylogenetic relationships. Each of these zones suffers from its own combination of issues (stochastic error, systematic error, data errors, computational requirements and missing data). Interestingly, the "optimal zone" in phylogenomics is not the one corresponding to the highest number of genes and taxa, because this computationally "intractable zone" is also the one where data errors and missing data are the most abundant. The latter aspect is due to the continuous shrinking of the number of orthologous genes when considering increasingly more species, owing to gene loss, gene duplication and gene transfer events.
Fig. 3 in Pitfalls in supermatrix phylogenomics
Fig. 3. Typical output of Phylo-MCOA. A matrix containing as many rows as the number of species and as many columns as the number of genes was computed, in which complete (black arrows) and cell-bycell (dashed circles) outliers can easily be detected. Cells with a high value (dark grey) represent species whose position in a given gene is not concordant with their position in all the other genes. It is thus a measure of distance to the common signal present in the data.
Data from: Bayesian analysis of a morphological supermatrix sheds light on controversial fossil hominin relationships
Open the record for dataset details and reuse information.
Data from: Patterns of macroevolution among Primates inferred from a supermatrix of mitochondrial and nuclear DNA.
Open the record for dataset details and reuse information.
Data from: Taxon sampling to address an ancient rapid radiation: a supermatrix phylogeny of early brachyceran flies (Diptera)
Open the record for dataset details and reuse information.
Data from: Combining phylogenomic and supermatrix approaches, and a time-calibrated phylogeny for squamate reptiles (lizards and snakes) based on 52 genes and 4162 species
Open the record for dataset details and reuse information.
Data from: A supermatrix phylogeny of corvoid passerine birds (Aves: Corvides)
Open the record for dataset details and reuse information.
Data from: A comparison of supermatrix and supertree methods for multilocus phylogenetics using organismal datasets
It has been proposed that supertree approaches should be applied to large multilocus sequence datasets to achieve computational tractability. Large datasets such as those derived from phylogenomics studies can be broken into many locus-specific tree searches and the resulting trees can be stitched together via a supertree method. Using simulated data, workers have reported that they can rapidly construct a supertree that is comparable to the results of heuristic tree search on the entire dataset. To test this assertion with organismal data, we compared tree length under the parsimony criterion and computational time for twenty multilocus datasets using supertree (SuperFine and SuperTriplets) and supermatrix (heuristic search in TNT) approaches. Tree length and computational times were compared among methods using the Wilcoxon matched-pairs signed rank test. Supermatrix searches produce significantly shorter trees than either supertree approach (SuperFine or SuperTriplets; p < 0.0002 in both cases). Moreover, the processing time of supermatrix search was significantly lower than SuperFine+locus-specific search (p < 0.01) but roughly equivalent to that of SuperTriplets+locus-specific search (p > 0.4, not significant). In conclusion, we show by using real rather than simulated data, that there is no basis, either in time tractability or tree length, for use of supertrees over heuristic tree search using a supermatrix for phylogenomics.
Data from: Building the avian tree of life using a large-scale, sparse supermatrix
Birds are the most diverse tetrapod class, with about 10,000 extant species that represent a remarkable evolutionary radiation in which most taxa arose during a short period of time. There has been a tremendous increase in the amount of molecular data available from birds, and more than two-thirds of these species have some sequence data available. Here we assembled these available sequence data from birds to estimate a large-scale avian phylogeny. We performed an unconstrained maximum likelihood analysis of a sparse supermatrix comprising 22 nuclear loci and seven mitochondrial regions from 6714 species. We inferred a phylogeny with a backbone remarkably similar to that obtained by detailed analyses of multigene datasets, yet with the addition of thousands of more taxa. All orders were monophyletic with generally high support. While most families and genera were well supported, a number of them, especially within the oscine passerines, had little or no support. This likely reflects problems with the circumscription of these genera and families. Our results indicate that the amount of sequence data currently available is sufficient to produce a robust estimate of the avian tree of life using current methods of inference. The availability of a tree that is unconstrained by prior information, with branch lengths that have a direct connection to the underlying data, should be useful for comparative methods, taxonomic revisions, and prioritizing taxa that should be targeted for additional data collection.
Data from: Phylogenetic relationships and character evolution analysis of Saxifragales using a supermatrix approach
Premise of the study: We sought novel evolutionary insights for the highly diverse Saxifragales by constructing a large phylogenetic tree encompassing 36.8% of the species-level biodiversity. Methods: We built a phylogenetic tree for 909 species of Saxifragales and used this hypothesis to examine character evolution for: annual or perennial habit, woody or herbaceous habit, ovary position, petal number, carpel number, and stamen: petal ratio. We employed likelihood approaches to investigate the effect of habit and life history on speciation and extinction within this clade. Key results: Two major shifts occurred from a woody ancestor to the herbaceous habit, with multiple secondary changes from herbaceous to woody. Transitions among superior, subinferior, and inferior ovaries appear equiprobable. A major increase in petal number is correlated with a large increase in carpel number; these increases have co-occurred multiple times in Crassulaceae. Perennial or woody lineages have higher rates of speciation than annual or herbaceous ones, but higher probabilities of extinction offset these differences. Hence, net diversification rates are highest for annual, herbaceous lineages and lowest for woody perennials. The shift from annuality to perenniality in herbaceous taxa is frequent. Conversely, woody perennial lineages to woody annual transitions are infrequent; if they occur, the woody annual state is left immediately. Conclusions: The large tree provides new insights into character evolution that are not obvious with smaller trees. Our results indicate that in some cases the evolution of angiosperms might be conditioned by constraints that have been so far overlooked.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.