Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
193
datasets available to search
ShareScore release 0.9.0
Dataset results
193 results for “parsimony”
Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany
<p>The dataset consists of particulate matter pollution concentration, measured in three localities - Hermsdorf, Charlottenburg and Adlershof, in Berlin, Germany.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer_rd_30s.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer_rd_30s.geojson</a> shows the observed PM2.5 concentration in a 30 second interval.</p> <p><a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> shows the concentrations shown is the local concentration (observed concentration - background concentration) in a 30 second interval. The background concentration is calculated as the lowest 5 percentile of the measured concentration for each measurement round. </p> <p><a href="../api/records/10076056/draft/files/PM2.5_lc_max.geojson/content" target="_blank" rel="noopener noreferrer">PM2.5_lc_max.geojson</a> contains the information from <a href="../api/records/10076056/draft/files/pm25_summer.geojson/content" target="_blank" rel="noopener noreferrer">pm25_summer.geojson</a> in a 25m resolution. Additionally, it contains the land use information for each coordinate.</p> <p>The original publication providing all necessary background information on study sites, methodology and data processing is the following: Venkatraman Jagatha, J., T. Sauter, C. Schneider (2024): Parsimonious Random-Forest-Based Land-Use Regression Model Using Particulate Matter Sensors in Berlin, Germany. MDPI Sensors, 24(13), 4193, DOI: 10.3390/s24134193. The paper is fully open access and can be downloaded at <a href="https://doi.org/10.3390/s24134193">https://doi.org/10.3390/s24134193</a>.</p> <p>Information on working with geojson file can be found under <a href="https://geojson.readthedocs.io/en/latest/">GeoJSON</a> .</p>
Parsimonious machine learning for the global mapping of aboveground biomass density
<p>This repository hosts data and code presented in the article "Parsimonious machine learning for the global mapping of aboveground biomass potential". The repository contains a compressed file containing all the code needed to reproduce the methodology that we developed and to analyse its results. We did not upload all the temporary and intermediate data files that are created during the execution of the method. We rather uploaded "milestone" data, i.e. final results or important intermediate ones. This includes the final training dataset, model calibration data, the final trained model, the global data for prediction, the final global map of potential aboveground biomass density (AGBD) at present times (raster files at 1km2 and 10km2 resolution), maps depicting regions where climatic conditions are outside of the training range of positive AGBD instances and maps depicting world regions without trees. </p> <p><strong>Files:</strong></p> <p><a href="../api/records/11580414/draft/files/code.zip/content" target="_blank" rel="noopener noreferrer">code.zip</a> : Compressed directory with all the code needed to reproduce the methodology presented in the manuscript. Contains a README file. Also contains temporary data generated in the process, the training dataset, the trained model, and model calibration data.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_1km_present_climate_1980_2010.tif</a> : the predicted global potential AGBD under contemporary climate conditions and at a resolution of 1 squared kilometer.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_10km_</a><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">present_climate_1980_2010.tif</a> : the predicted global potential AGBD under contemporary climate conditions downsampled at a resolution of 10 squared kilometers.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_10km_model_difference.tif</a> : the difference between our prediction of potential AGBD and the prediction from a complex state-of-the-art model from Walker et al. (2022). </p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGB_Mg_1km_</a><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">present_climate_1980_2010.tif</a> : the predicted global potential pixel-level AGB under contemporary climate conditions downsampled at a resolution of 1 squared kilometers.</p> <p><a href="../api/records/11580414/draft/files/number_predictors_out_of_range.zip/content" target="_blank" rel="noopener noreferrer">number_predictors_out_of_range.zip</a> : tiled maps representing the number of climatic predictors outside of the training range before including 0 AGBD instances in the training dataset. </p> <p><a href="../api/records/11580414/draft/files/tree_absence_map.zip/content" target="_blank" rel="noopener noreferrer">tree_absence_map.zip</a> : tiled maps representing world regions without trees. Based on Crowther et al. (2015) (https://elischolar.library.yale.edu/yale_fes_data/1/).</p> <p><a href="../api/records/11580414/draft/files/potential_agbd_Mgha_climate_envelope.pkl/content" target="_blank" rel="noopener noreferrer">inference_pipeline_potential_agbd_Mgha_climate.pkl</a> : Calibrated model for the prediction of potential AGBD given bioclimatic conditions. </p> <p><a href="../api/records/11580414/draft/files/predictors_data_global.zip/content" target="_blank" rel="noopener noreferrer">predictors_data_global.zip</a> : Global predictors data to apply the model on.</p>
Accuracy of phylogenetic reconstructions from continuous characters analyzed under parsimony and its parametric correlates
<p>Quantitative traits are a source of evolutionary information often difficult to handle in cladistics. Tools exist to analyze this kind of data without subjective discretization, avoiding biases in the delimitation of categorical states. Nonetheless, the ability of continuous characters to accurately infer relationships is incompletely understood, particularly under parsimony analysis. This study evaluates the accuracy of phylogenetic reconstructions from simulated matrices of continuous characters evolving under alternative evolutionary processes and analyzed by parsimony. We sampled 100 empirical trees to simulate 9,000 matrices, each containing between 25 and 50 taxa and 50 and 150 continuous characters evolving under three evolutionary processes: Brownian-Motion (BM), Ornstein-Uhlenbeck (OU) and Early-Burst (EB) with variable parametrizations. Our cladogram comparisons revealed that continuous character matrices, when discretized objectively and analyzed by parsimony in TNT, carry phylogenetic signals to infer species relationships, regardless of the evolutionary models and parameterization schemes. Interestingly, implementing Equal Weighting (EW) or Implied Weighting (IW) with varying penalization strengths against homoplasies did not affect cladogram reconstructions on the basis of continuous characters. Finally, the accuracy of continuous characters in resolving species relationships is skewed toward apical nodes of the recovered trees. Our findings provide general insights of the utility of quantitative traits in cladistics and demonstrate that their effectiveness in estimating shallower nodes is independent of the underlying evolutionary model, parameters and weighting schemes.</p>
Code and data: Exploring congruent diversification histories with flexibility and parsimony
<p>This repository contains the code and data for the article "Exploring congruent diversification histories with flexibility and parsimony" (abstract bellow).</p> <p>Data :</p> <ul> <li><strong>4705sp_mammal-time.tree</strong>: Species-level calibrated mammalian phylogeny from <em>Alvarez-Carretero et al. </em>(<a href="https://doi.org/10.6084/m9.figshare.14885691">https://doi.org/10.6084/m9.figshare.14885691</a>)</li> <li><strong>mammals_samplingfraction.csv</strong> : Clade-specific sampling fractions from <em>Quintero et al.</em>(<a href="https://www.biorxiv.org/content/10.1101/2022.08.09.503355v1.full">https://www.biorxiv.org/content/10.1101/2022.08.09.503355v1.full</a>).</li> </ul> <p>Code :</p> <ul> <li><strong>CRABS-v1.1.0.9004.zip</strong>: Archived version of the CRABS package with our extension.</li> <li><strong>Mammalian_rates_EBD_HSMRF.rev</strong>: <em>Rev</em> script for the mammalian diversification analysis in RevBayes with regularized priors on diversification rates.</li> <li><strong>Mammalian_rates_EBD_independent.rev</strong>: <em>Rev</em> script for the mammalian diversification analysis in RevBayes with independent diversification rates at each interval.</li> <li><strong>Mammals_proccess_RevBayes_outputs.Rmd</strong>: R notebook for processing the outputs from the RevBayes mammalian diversification analysis, plotting the rates through time, and saving the median trajectories used for further analyses.</li> <li><strong>Exploring_congruent_diversification_histories_with_flexibility_and_parsimony.Rmd</strong>: R notebook for comparing the initial CRABS features and our new extensions. It enables replicating the figures in the article.</li> </ul> <p>Outputs :</p> <ul> <li><strong>output_inferredIntervals_fixedRhp_HSMRF.zip</strong> & <strong>output_inferredIntervals_fixedRhp_independent.zip</strong>: The raw traces from the RevBayes analysis, and the resulting median rate trajectories that are used to construct the congruence class illustrated in the article.</li> </ul> <p><br> Abstract</p> <ol> <li>Using phylogenies of present-day species to estimate diversification rate trajectories -- speciation and extinction rates over time -- is a challenging task due to non-identifiability issues. Given a phylogeny, there exists an infinite set of trajectories that result in the same likelihood; this set has been coined a congruence class. Previous work has developed approaches for sampling trajectories within a given congruence class, with the aim to assess the extent to which congruent scenarios can vary from one another. Based on this sampling approach, it has been suggested that rapid changes in speciation or extinction rates are conserved across the class. Reaching such conclusions requires to sample the broadest possible set of distinct trajectories.</li> <li>We introduce a new method for exploring congruence classes, that we implement in the R package CRABS. Whereas existing methods constrain either the speciation rate or the extinction rate trajectory, ours provides more flexibility by sampling congruent speciation and extinction rate trajectories simultaneously. This allows covering a more representative set of distinct diversification rate trajectories. We also implement a filtering step that allows selecting the most parsimonious trajectories within a class.</li> <li>We demonstrate the utility of our new sampling strategy using a simulated scenario. Next, we apply our approach to the study of mammalian diversification history. We show that rapid changes in speciation and extinction rates need not be conserved across a congruence class, but that selecting the most parsimonious trajectories shrinks the class to concordant scenarios.</li> <li>Our approach opens new avenues both to truly explore the myriad of potential diversification histories consistent with a given phylogeny, embracing the uncertainty inherent to phylogenetic diversification models, and to select among these different histories. This should help refining our inference of diversification trajectories from extant data.</li> </ol>
FIG. 5. — Strict consensus tree from eight most parsimonious trees recovered for Molossus E. Geoffroy, 1805 in Diversity, morphological phylogeny, and distribution of bats of the genus Molossus E. Geoffroy, 1805 (Chiroptera, Molossidae) in Brazil
FIG. 5. — Strict consensus tree from eight most parsimonious trees recovered for Molossus E. Geoffroy, 1805. Numbers above the branches indicate Bootstrap values and bottom numbers indicate Bremer support values.
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Text-fig. 9. Phylogenetic tree indicating the number of required character state changes (steps) under parsimony for various positions of Miranthus gen. nov. in a molecular based backbone tree (see material and methods for additional details). in Early Flowers Of Primuloid Ericales From The Late Cretaceous Of Portugal And Their Ecological And Phytogeographic Implications
Text-fig. 9. Phylogenetic tree indicating the number of required character state changes (steps) under parsimony for various positions of Miranthus gen. nov. in a molecular based backbone tree (see material and methods for additional details).
Figure 1. Most parsimonious phylogeny among 12 in Monophyly and taxonomy of the Neotropical seasonal killifish genus Leptolebias (Teleostei: Aplocheiloidei: Rivulidae), with the description of a new genus
Figure 1. Most parsimonious phylogeny among 12 species of the Rivulidae (tree length, L = 118; consistency index, CI = 0.82; retention index, RI = 0.85). Numbers above branches are bootstrap values.
Figure 6. Most parsimonious result from a in Exploring phylogenetic relationships of Pteraspidiformes heterostracans (stem-gnathostomes) using continuous and discrete characters
Figure 6. Most parsimonious result from a phylogenetic analysis of discrete (1—64) and discretized continuous characters identified through gap coding (88—100). A, strict consensus of 30 most parsimonious trees with equally weighted characters (tree length 346). B, most parsimonious solution with implied weighted characters (k = 3) (tree length 27.86). Psammosteidae taxa in bold (for which quantitative characters have been treated as inapplicable, i.e. non-homologous).
FIGURE 37. Phylogenetic results from parsimony analysis using the cranial dataset. A in A reappraisal of the cranial and mandibular osteology of the spinosaurid Irritator challengeri (Dinosauria: Theropoda)
FIGURE 37. Phylogenetic results from parsimony analysis using the cranial dataset. A, strict consensus tree of 153 MPTs retained from an equal weighting analysis (see Methods for details); B, reduced consensus tree, pruning wild card taxa from the strict consensus. Wild card taxa are highlighted with coloured boxes in A, and their possible topological positions are shown with same coloured squares in B. Important clades are labelled.
FIGURE 36. Phylogenetic results from parsimony analysis using the full dataset. A in A reappraisal of the cranial and mandibular osteology of the spinosaurid Irritator challengeri (Dinosauria: Theropoda)
FIGURE 36. Phylogenetic results from parsimony analysis using the full dataset. A, strict consensus tree of 8184 MPTs retained from an equal weighting analysis (see methods for details); B, partial reduced consensus tree, showing the clade Spinosauridae after removal of the taxon Vallibonavenatrix; C, strict consensus tree of 406 MPTs retained from an implied weighting analysis using a concavity constant of k=10 (see Methods for details). Important clades are labelled. Irritator as the main focus of our study is highlighted in bold face within the clade Spinosauridae.
◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West) in Bumps on the back: An unusual morphology in phylogenetically distinct Peridinium aff. cinctum (= Peridinium tuberosum; Peridiniales, Dinophyceae)
◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West)
◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae) in Morphological and molecular variability of Peridinium volzii Lemmerm. (Peridiniaceae, Dinophyceae) and its relevance for infraspecific taxonomy
◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae)
Fig. 2. Maximum parsimony tree inferred from 18S in Novel piroplasmid and Hepatozoon organisms infecting the wildlife of two regions of the Brazilian Amazon
Fig. 2. Maximum parsimony tree inferred from 18S rRNA gene sequences of Hepatozoon spp., with Babesia sp. as outgroup (488 characters; 52 parsimony-informative sites). Numbers at nodes are the support values for the major branches (bootstrap over 500 replicates). The sequences obtained in this study are in bold. Numbers in brackets are GenBank accession numbers.
Fig. 1. Maximum parsimony tree inferred from 18S in Novel piroplasmid and Hepatozoon organisms infecting the wildlife of two regions of the Brazilian Amazon
Fig. 1. Maximum parsimony tree inferred from 18S rRNA gene sequences of piroplasmids (Babesia spp., Theileria spp., Cytauxzoon spp.), with Plasmodium ovale as outgroup (316 characters; 65 parsimony-informative sites). Numbers at nodes are the support values for the major branches (bootstrap over 500 replicates). The sequences obtained in this study are in bold. Numbers in brackets are GenBank accession numbers.
Fig. 5. Parsimony splits network constructed from a per and ITS2 concatenated sequence data set. Heterozygous specimens are indicated with A and B in Ecological and geographical speciation in Lucilia bufonivora: The evolution of amphibian obligate parasitism
Fig. 5. Parsimony splits network constructed from a per and ITS2 concatenated sequence data set. Heterozygous specimens are indicated with A and B. 'bufonivora_EUROPE_A' represents a consistent haplotype present in all 12 samples from Europe (Table 1), of which just two were heterozygous ('bufonivora_frog' and 'bufonivora_NLWi'). 'bufonivora_CAN' and 'elongata_CAN' are represented by two samples each, none of which were heterozygous. Scale bar represents expected changes per site.
Text-fig. 6. Most parsimonious tree obtained after addition of Acaciaephyllum to the data set of Doyle (2008), with modifications discussed in the text, and with relationships of other taxa fixed with a backbone constraint tree based on results of Doyle (2008). Relative parsimony of alternative positions of Acaciaephyllum is indicated as in Text-fig. 2. Gnet = Gnetales. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 6. Most parsimonious tree obtained after addition of Acaciaephyllum to the data set of Doyle (2008), with modifications discussed in the text, and with relationships of other taxa fixed with a backbone constraint tree based on results of Doyle (2008). Relative parsimony of alternative positions of Acaciaephyllum is indicated as in Text-fig. 2. Gnet = Gnetales.
Text-fig. 4. Most parsimonious trees obtained after addition of Virginianthus (with "Liliacidites" minutus pollen) to the (A) D&E and (B) J/M trees. Relative parsimony of alternative positions of Virginianthus is indicated as in Text-fig. 2; abbreviations as in Text-fig. 1. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 4. Most parsimonious trees obtained after addition of Virginianthus (with "Liliacidites" minutus pollen) to the (A) D&E and (B) J/M trees. Relative parsimony of alternative positions of Virginianthus is indicated as in Text-fig. 2; abbreviations as in Text-fig. 1.
Text-fig. 2. Representative most parsimonious trees obtained after addition of Liliacidites to (A) the D&E tree (Text-fig. 1) and (B) the J/M tree, with relationships among major clades based on the plastid genome analyses of Jansen et al. (2007) and Moore et al. (2007). Thicker lines indicate all most parsimonious (MP), one step less parsimonious (MP+1), and two step less parsimonious (MP+2) positions for Liliacidites. Abbreviations as in Text-fig. 1. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 2. Representative most parsimonious trees obtained after addition of Liliacidites to (A) the D&E tree (Text-fig. 1) and (B) the J/M tree, with relationships among major clades based on the plastid genome analyses of Jansen et al. (2007) and Moore et al. (2007). Thicker lines indicate all most parsimonious (MP), one step less parsimonious (MP+1), and two step less parsimonious (MP+2) positions for Liliacidites. Abbreviations as in Text-fig. 1.
Text-fig. 3. One of two most parsimonious trees obtained after addition of Anacostia (with Similipollis pollen) to the D&E tree. Relative parsimony of alternative positions of Anacostia is indicated as in Text-fig. 2; abbreviations as in Text-fig. 1. in Early Cretaceous Monocots: A Phylogenetic Evaluation
Text-fig. 3. One of two most parsimonious trees obtained after addition of Anacostia (with Similipollis pollen) to the D&E tree. Relative parsimony of alternative positions of Anacostia is indicated as in Text-fig. 2; abbreviations as in Text-fig. 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.