Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
193
datasets available to search
ShareScore release 0.9.0
Dataset results
193 results for “parsimony”
Bayesian morphological clock versus parsimony: An insight into the relationships and dispersal events of postvacuum Cricetidae (Rodentia, Mammalia)
Open the record for dataset details and reuse information.
Code, data and results for manuscript "A parsimonious empirical approach to streamflow recession analysis and forecasting"
<p>This repository hosts the supplementary materials associated with the paper:<br> > Delforge, D., Muñoz-Carpena, R., Van Camp, M. Vanclooster, M. (2020), A parsimonious empirical approach to streamflow recession analysis and forecasting (accepted at Water Resources Research - 29-01-2020).</p> <p>This data set contains streamflow and recession data, a python code file and a Jupyter notebook illustrating how to apply the EDM-Simplex method to forecast the recession, and the outputs of the global sensitivity analysis. All files are documented in the readme.md Markdown files. </p> <p>Streamflow data were obtained from the Aqualim portal (<a href="http://aqualim.environnement.wallonie.be/">http://aqualim.environnement.wallonie.be/</a>) of the "Service Public de Wallonie" and shared with their kind permission. This work is part of a Ph.D. supported by a FRIA grant from the Fund for Scientific Research (FSR-FNRS, Belgium). The authors acknowledge University of Florida Research Computing for providing computational resources and support that have contributed to the research results stored in this repository. URL: <a href="http://researchcomputing.ufl.edu">http://researchcomputing.ufl.edu</a>.</p>
Verifiability of genus-level classification under quantification and parsimony theories: a case study of follicucullid radiolarians
<p>The classical taxonomy of fossil invertebrates is based on subjective judgments of morphology, which can cause confusion since there are no codified standards for the classification of genera. Here, we explore the validity of the genus taxonomy of 75 species and morphospecies of the Follicucullidae, a Late Paleozoic family of radiolarians, using a new method, <a name="_Hlk41571699">Hayashi's quantification theory II</a> (HQT-II), a general multivariate statistical method for categorical datasets relevant to discriminant analysis. We identify a scheme of ten genera rather than the currently accepted three genera (<i>Follicucullus</i>, <i>Ishigaconus</i> and <i>Parafollicucullus</i>). As HQT-II cannot incorporate stratigraphic data, a phylogenetic tree of Follicucullidae was reconstructed for 38 species using maximum parsimony. Six lineages emerged, roughly in concordance with the results of HQT-II. Combined with parsimony ancestral state reconstruction, the ancestral group of this family is <i>Haplodiacanthus</i>. Five other groups were discriminated, the <i>Parafollicucullus</i>, <i>Curvalbaillella</i>, <i>Pseudoalbaillella</i>, <i>Longtanella</i>, and <i>Follicucullus</i>–<i>Cariver</i> lineages. The morphological evolution of these lineages comprises a minimum essential list of eight states of four traits. <a name="_Hlk41488876">HQT-II is a novel discriminant analytical multivariate method that may be of value in other taxonomic problems of paleobiology.</a></p>
Dataset and codes for "BaHSYM: parsimonious Bayesian Hierarchical Model to predict river Sediment Yield"
<p>This folder contains:</p> <ul> <li>R project file</li> <li>R code for Best Fit model</li> <li>R code for temporal cross-validation</li> <li>R code for spatial cross-validation</li> <li>R code for cluster analysis</li> <li>dataset containing all input variables for the river gauges (and catchments) used for the development and testing of the BaHSYM model in Austria</li> </ul> <p>It also contains the same codes and datasets adapted to reproduce the model by de Vente et al. (2011), i.e. with the same structure but with the variables used in such model.</p>
Data from: Maximum parsimony inference of phylogenetic networks in the presence of polyploid complexes
<p>Phylogenetic networks provide a powerful framework for modeling and analyzing reticulate evolutionary histories. While polyploidy has been shown to be prevalent not only in plants but also in other groups of eukaryotic species, most work done thus far on phylogenetic network inference assumes diploid hybridization. These inference methods have been applied, with varying degrees of success, to data sets with polyploid species, even though polyploidy violates the mathematical assumptions underlying these methods. Statistical methods were developed recently for handling specific types of polyploids and so were parsimony methods that could handle polyploidy more generally yet while excluding processes such as incomplete lineage sorting.</p> <p>In this paper, we introduce a new method for inferring most parsimonious phylogenetic networks on data that include polyploid species. Taking gene trees as input, the method seeks a phylogenetic network that minimizes deep coalescences while accounting for polyploidy. The method could also infer trees, thus potentially distinguishing between auto- and allo-polyploidy. We demonstrate the performance of the method on both simulated and biological data. The inference method as well as a method for evaluating given phylogenetic networks are implemented and publicly available in the PhyloNet software package.</p>
Data from: Parsimonious inference of hybridization in the presence of incomplete lineage sorting
Hybridization plays an important evolutionary role in several groups of organisms. A phylogenetic approach to detect hybridization entails sequencing multiple loci across the genomes of a group of species of interest, reconstructing their gene trees, and taking their differences as indicators of hybridization. However, methods that follow this approach mostly ignore population effects, such as incomplete lineage sorting (ILS). Given that hybridization occurs between closely related organisms, ILS may very well be at play and, hence, must be accounted for in the analysis framework. To address this issue, we present a parsimony criterion for reconciling gene trees within the branches of a phylogenetic network, and a local search heuristic for inferring phylogenetic networks from collections of gene-tree topologies under this criterion. This framework enables phylogenetic analyses while accounting for both hybridization and ILS. Further, we propose two techniques for incorporating information about uncertainty in gene-tree estimates. Our simulation studies demonstrate the good performance of our framework in terms of identifying the location of hybridization events, as well as estimating the proportions of genes that underwent hybridization. Also, our framework shows good performance in terms of efficiency on handling large data sets in our experiments. Further, in analysing a yeast data set, we demonstrate issues that arise when analysing real data sets. Although a probabilistic approach was recently introduced for this problem, and although parsimonious reconciliations have accuracy issues under certain settings, our parsimony framework provides a much more computationally efficient technique for this type of analysis. Our framework now allows for genome-wide scans for hybridization, while also accounting for ILS.
Data from: Phylogenetic inference using discrete characters: performance of ordered and unordered parsimony and of three-item statements
The cladistic literature does not always specify the kind of multistate character treatment that is applied for an analysis. Characters can be treated either as unordered transformation series or as rooted [three-item analysis (3ia)] or unrooted state trees (ordered characters). We aimed to measure the impact of these character treatments on phylogenetic inference. Discrete characters can be represented either as rows or columns in matrices (e.g. for parsimony) or as hierarchies for 3ia. In the present study, we use simulated and empirical examples to assess the relative merits of each method considering both the character treatment and representation. We measure two parameters (resolving power and artefactual resolution) using a new tree comparison metric, ITRI (inter-tree retention index). Our results suggest that the hierarchical character representation not only results (with our simulation settings) in the greatest resolving power, but also in the highest artefactual resolution. Our empirical examples provide equivocal results. Parsimony unordered states yield less resolving power and more artefactual resolutions than parsimony ordered states, both with our simulated and empirical data. Relationships between three operational taxonomic units (OTUs), irrespective of their relationships with other OTUs, are called three-item statements (3is). We compare the intersection tree (which reconstructs a single tree from all of the common 3is of source trees) with the traditional strict consensus and show that the intersection tree retains more of the information contained in the source trees.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Fossil taxa are critical to inferences of historical diversity and the origins of modern biodiversity, but realizing their evolutionary significance is contingent on restoring fossil species to their correct position within the tree of life. For most fossil species, morphology is the only source of data for phylogenetic inference; this has traditionally been analysed using parsimony, the predominance of which is currently challenged by the development of probabilistic models that achieve greater phylogenetic accuracy. Here, based on simulated and empirical datasets, we explore the relative efficacy of competing phylogenetic methods in terms of clade support. We characterize clade support using bootstrapping for parsimony and Maximum Likelihood, and intrinsic Bayesian posterior probabilities, collapsing branches that exhibit less than 50% support. Ignoring node support, Bayesian inference is the most accurate method in estimating the tree used to simulate the data. After assessing clade support, Bayesian and Maximum Likelihood exhibit comparable levels of accuracy, and parsimony remains the least accurate method. However, Maximum Likelihood is less precise than Bayesian phylogeny estimation, and Bayesian inference recaptures more correct nodes with higher support compared to all other methods, including Maximum Likelihood. We assess the effects of these findings on empirical phylogenies. Our results indicate probabilistic methods should be favoured over parsimony.
Data from: Inferring complex phylogenies using parsimony: an empirical approach using three large DNA data sets for angiosperms
To explore the feasibility of parsimony analysis for large data sets, we conducted heuristic parsimony searches and bootstrap analyses on separate and combined DNA data sets for 190 angiosperms and three outgroups. Separate data sets of 18S rDNA (1,855 bp), rbc L (1,428 bp), and atp B (1,450 bp) sequences were combined into a single matrix 4,733 bp in length. Analyses of the combined data set show great improvements in computer run times compared to those of the separate data sets and of the data sets combined in pairs. Six searches of the 18S rDNA rbc L atp B data set were conducted; in all cases TBR branch swapping was completed, generally within a few days. In contrast, TBR branch swapping was not completed for any of the three separate data sets, or for the pairwise combined data sets. These results illustrate that it is possible to conduct a thorough search of tree space with large data sets, given sufficient signal. In this case, and probably most others, sufficient signal for a large number of taxa can only be obtained by combining data sets. The combined data sets also have higher internal support for clades than the separate data sets, and more clades receive bootstrap support of 50% in the combined analysis than in analyses of the separate data sets. These data suggest that one solution to the computational and analytical dilemmas posed by large data sets is the addition of nucleotides, as well as taxa.
Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 in A new genus and three new species of mangrove slugs from the Indo-West Pacific (Mollusca: Gastropoda: Euthyneura: Onchidiidae)
Fig. 4. Maximum parsimony consensus tree within Paromoionchis gen. nov., performed with ITS2 DNA sequences from 80 individuals (including 7 outgroups). Numbers by the branches are the bootstrap values (only numbers>50% are indicated). Numbers for each individual correspond to unique identifiers for DNA extraction. All sequences for specimens of Paromoionchis gen. nov. are new. Information on specimens can be found in the lists of material examined and in Table 1. The letter A corresponds to a clade referred to in the text. The color used for each (mitochondrial) unit is the same as that used in Figs 1–3 and 5–6.
FIGURE 1 Maximum Parsimony phylogenetic analyses. A in Dating the origin and diversiFIcation of Pan-Chelidae (Testudines, Pleurodira) under multiple molecular clock approaches
FIGURE 1 Maximum Parsimony phylogenetic analyses. A: Morphological phylogeny. B: Molecular phylogeny. C: Total-evidence phylogeny. Bootstrap supports are coded in grayscale. Australasian species are shown in red; South American species are shown in green. Abbreviations: A, Acanthochelys; B, Bonapartemys; Ch, Chelodina; El, Elseya; H, Hydromedusa; L, Lomalatachelys; M, Mesoclemmys; Me, Mendozachelys; My, Myuchelys; Pa, Palaeophrynops; Ph, Phrynops; Pl, Platemys; Pr, Prochelidella; Ps, Pseudemydura; Ri, Rionegrochelys; Y, Yaminuechelys. †, extinct taxa.
Data from: Bayesian methods outperform parsimony but at the expense of precision in the estimation of phylogeny from discrete morphological data
Open the record for dataset details and reuse information.
Data from: Turning the crown upside down: gene tree parsimony roots the eukaryotic tree of life
Open the record for dataset details and reuse information.
Verifiability of genus-level classification under quantification and parsimony theories: a case study of follicucullid radiolarians
Open the record for dataset details and reuse information.
Data from: Inferring complex phylogenies using parsimony: an empirical approach using three large DNA data sets for angiosperms
Open the record for dataset details and reuse information.
Data from: Phylogeography and Molecular Systematics of the Peromyscus Aztecus Species Group (Rodentia: Muridae) Inferred Using Parsimony and Likelihood
Open the record for dataset details and reuse information.
Data from: Probabilistic methods surpass parsimony when assessing clade support in phylogenetic analyses of discrete morphological data
Open the record for dataset details and reuse information.
Data from: Phylogenetic inference using discrete characters: performance of ordered and unordered parsimony and of three-item statements
Open the record for dataset details and reuse information.
Data from: Parsimony, not Bayesian analysis, recovers more stratigraphically congruent phylogenetic trees
Open the record for dataset details and reuse information.
Data from: Probabilistic methods outperform parsimony in the phylogenetic analysis of data simulated without a probabilistic model
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.