Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “probabilistic morphology”
Evaluating the performance of probabilistic algorithms for phylogenetic analysis of big morphological datasets: a simulation study
<p>Reconstructing the tree of life is an essential task in evolutionary biology. It demands accurate phylogenetic inference for both extant and extinct organisms, the latter being almost entirely dependent on morphological data. While parsimony methods have traditionally dominated the field of morphological phylogenetics, a rapidly growing number of studies are now employing probabilistic methods (maximum likelihood and Bayesian inference). The present-day toolkit of probabilistic methods offers varied software with distinct algorithms and assumptions for reaching global optimality. However, benchmark performance assessments of different software packages for the analyses of morphological data, particularly in the era of big data, are still lacking. Here, we test the performance of four major probabilistic software under variable taxonomic sampling and missing data conditions: the Bayesian inference-based programs MrBayes and RevBayes, and the maximum likelihood-based IQ-TREE and RAxML. We evaluated software performance by calculating the distance between inferred and true trees using a variety of metrics, including Robinson-Foulds (RF), Matching Splits (MS), and Kuhner-Felsenstein (KF) distances. Our results show that increased taxonomic sampling improves accuracy, precision, and resolution of reconstructed topologies across all tested probabilistic software applications and all levels of missing data. Under the RF metric, Bayesian inference applications were the most consistent, accurate, and robust to variation in taxonomic sampling in all tested conditions, especially at high levels of missing data, with little difference in performance between the two tested programs. The MS metric favored more resolved topologies that were generally produced by IQ-TREE. Adding more taxa dramatically reduced performance disparities between programs. Importantly, our results suggest that the RF metric penalizes incorrectly resolved nodes (false positives) more severely than the MS metric, which instead tends to penalize polytomies. If false positives are to be avoided in systematics, Bayesian inference should be preferred over maximum likelihood for the analysis of morphological data.</p>
Evaluating the performance of probabilistic algorithms for phylogenetic analysis of big morphological datasets: a simulation study
Open the record for dataset details and reuse information.
Tempo and mode in karyotype evolution revealed by a probabilistic model incorporating both chromosome number and morphology
Open the record for dataset details and reuse information.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Fossil taxa are critical to inferences of historical diversity and the origins of modern biodiversity, but realizing their evolutionary significance is contingent on restoring fossil species to their correct position within the tree of life. For most fossil species, morphology is the only source of data for phylogenetic inference; this has traditionally been analysed using parsimony, the predominance of which is currently challenged by the development of probabilistic models that achieve greater phylogenetic accuracy. Here, based on simulated and empirical datasets, we explore the relative efficacy of competing phylogenetic methods in terms of clade support. We characterize clade support using bootstrapping for parsimony and Maximum Likelihood, and intrinsic Bayesian posterior probabilities, collapsing branches that exhibit less than 50% support. Ignoring node support, Bayesian inference is the most accurate method in estimating the tree used to simulate the data. After assessing clade support, Bayesian and Maximum Likelihood exhibit comparable levels of accuracy, and parsimony remains the least accurate method. However, Maximum Likelihood is less precise than Bayesian phylogeny estimation, and Bayesian inference recaptures more correct nodes with higher support compared to all other methods, including Maximum Likelihood. We assess the effects of these findings on empirical phylogenies. Our results indicate probabilistic methods should be favoured over parsimony.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.