Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Simulated trees”
Input data for article "Large eddy simulation of the optimal street-tree layout for pedestrian-level aerosol particle concentrations"
<p>Input dataset used when performing LES simulations for journal article "Large eddy simulation of the optimal street-tree layout for pedestrian-level aerosol particle concentrations" (Karttunen et al., in preparation). The dataset was used with the PALM model system revision 3698 and most likely it won't work on older or newer versions.</p> <p>Instructions for use:<br> A precursor run must be run first. Output data (BINOUT) of it should be linked into a BININ directory of the actual scenario runs. You'll most likely have to tweak the CPU grid settings in ENVPAR and PARIN files in order to fit them to your computational resources. For more information on usage please refer to the PALM model documentation available online in <a href="https://palm.muk.uni-hannover.de/trac/wiki/doc">https://palm.muk.uni-hannover.de/trac/wiki/doc</a>.</p>
Evolutionary adaptation of trees and modelled future larch forest extent in Siberia. Code and simulation data
<p>Code and datset used for the publication: "Evolutionary adaptation of trees and modelled future larch forest extent in Siberia" 2023 Gloy et al.</p>
Data from: Genetic relationships, structure and parentage simulation among the olive tree (Olea europaea L. subsp. europaea) cultivated in Southern Italy revealed by SSR markers
Open the record for dataset details and reuse information.
Data from: Long term impacts of selective logging on two Amazonian tree species with contrasting ecological and reproductive characteristics: inferences from Eco-gene model simulations
Open the record for dataset details and reuse information.
Data from: Disentangling the formation of contrasting tree-line physiognomies combining model selection and Bayesian parameterization for simulation models
Open the record for dataset details and reuse information.
Wildland-urban interface fire dynamics simulator input files for pyric tree spatial patterning interactions in historical and contemporary mixed conifer forests, California, USA
Open the record for dataset details and reuse information.
Data from: Simulating local adaptation to climate of forest trees with a Physio-Demo-Genetics model
Open the record for dataset details and reuse information.
R script to simulate the phenology of the box tree moth, Cydalima perspectalis
<p>This folder contains :<br> - ReadMe file with the following explanation<br> - the R script to simulate the phenology of the box tree moth (ProgR_phenology_box_tree_moth.r),<br> - the flight curve of the box tree moth in Orleans (France) in 2017 (see published dataset: https://doi.org/10.5281/zenodo.3719293 )<br> - temperature dataset (temperature.txt; col 1 = year, col 2 = month, col 3 = day, col 4 = Tmin, col 5 = Tmax)<br> - photoperiod dataset (photoperiod.txt; col 1 = year, col 2 = month, col 3 = day, col 4 = julian day, col 5 = day length in hours)<br> Note that these two former datasets are provided because they are required to make simulations.<br> Warning: since we are not allowed to redistribute the temperature and photoperiod datasets, the related files have been modified and are not those actually used in our paper in preparation.</p> <p>The R script was used in R version 3.6.1 (2019-07-05).</p> <p>To download the software R, please visit: https://www.r-project.org/</p> <p>This work was done in the frame of a regional project called INCA (grant from Région Centre Val de Loire).</p>
Model trees and associated simulated nucleotide sequences for testing phylogenetic inference methods
<p>This repository contains 142 tar.gz archive files, each containing nucleotide sequence data that have been simulated using <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> for testing alignment-free phylogenetic inference methods. These datasets were generated by using the results (trees and model parameters) of 142 phylogenomic analyses of real-case data as model (available <a href="https://zenodo.org/record/4034261">here</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>Each archive contains the following files/directories:</p> <ul> <li><code>GTR.params.trees.tsv </code> a tab-delimited file summarizing the real-case GTR+Γ model parameters and the phylogenetic tree used to simulate the sequence dataset (gathered from <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>)</li> <li><code>tax.tsv </code> a tab-delimited file containing the initial (col 1) and simplified (col 2) taxon names</li> <li><code>model.nwk </code> a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing the initial model tree (gathered from <code>GTR.params.trees.tsv</code>) with simplified leaf names (following <code>tax.tsv</code>)</li> <li><code>control.txt </code> the <em>INDELible</em> input file used to simulate the evolution of a sequence along the tree in <code>model.nwk</code></li> <li><code>seq/ </code> a directory containing the simulated sequences (one FASTA file per leaf in the tree in <code>model.nwk</code>)</li> </ul> <p>___</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Data from: SimPhy: phylogenomic simulation of gene, locus and species trees
We present a fast and flexible software package—SimPhy—for the simulation of multiple gene families evolving under incomplete lineage sorting, gene duplication and loss, horizontal gene transfer—all three potentially leading to species tree/gene tree discordance—and gene conversion. SimPhy implements a hierarchical phylogenetic model in which the evolution of species, locus, and gene trees is governed by global and local parameters (e.g., genome-wide, species-specific, locus-specific), that can be fixed or be sampled from a priori statistical distributions. SimPhy also incorporates comprehensive models of substitution rate variation among lineages (uncorrelated relaxed clocks) and the capability of simulating partitioned nucleotide, codon, and protein multilocus sequence alignments under a plethora of substitution models using the program INDELible. We validate SimPhy's output using theoretical expectations and other programs, and show that it scales extremely well with complex models and/or large trees, being an order of magnitude faster than the most similar program (DLCoal-Sim). In addition, we demonstrate how SimPhy can be useful to understand interactions among different evolutionary processes, conducting a simulation study to characterize the systematic overestimation of the duplication time when using standard reconciliation methods. SimPhy is available at https://github.com/adamallo/SimPhy, where users can find the source code, precompiled executables, a detailed manual and example cases.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.
Data from: Using genomic location and coalescent simulation to investigate gene tree discordance in Medicago L.
Several well-documented evolutionary processes are known to cause conflict between species-level phylogenies and gene-level phylogenies. Three of the most challenging processes for species tree inference are incomplete lineage sorting, hybridization and gene duplication, which may result in unwarranted comparisons of paralogous genes. Several existing methods have dealt with these processes but none has yet been able to untangle all three at once. Here, we propose a stepwise method by which these processes can be discerned using information on genomic location coupled with coalescent simulations. In the first step, highly discordant genes within genomic blocks (putative paralogs) are identified and excluded from the data set and, in the second step, blocks of linked genes are grouped according to their hybrid history. Existing multispecies coalescent software can then be applied to recover the principal tree(s) that make up the species tree/network without violating the underlying model. The potential of the approach is evaluated on simulated data derived from a species network composed of nine species, of which one is of hybrid origin, and displaying a single-gene duplication that leads to paralogous comparisons. We apply our method to an empirical set of 12 genes from 7 species sampled in the plant genus Medicago that display phylogenetic discordance. We identify the causes of the discordance and demonstrate that the Medicago orbicularis lineage experienced an episode of ancient hybridization. Our results show promise as a new way to explore phylogenetic sequence data that can significantly improve species tree inference in presence of hybridization and undetected paralogy or other causes leading to extremely discordant gene trees.
SimPhy configuration scripts for simulations reported in the study titled: Species tree inference methods intended to deal with incomplete lineage sorting are robust to the presence of paralogs
<p>Many recent phylogenetic methods have focused on accurately inferring species trees when there is gene tree discordance due to incomplete lineage sorting (ILS). For almost all of these methods, and for phylogenetic methods in general, the data for each locus is assumed to consist of orthologous, single-copy sequences. Loci that are present in more than a single copy in any of the studied genomes are excluded from the data. These steps greatly reduce the number of loci available for analysis. The question we seek to answer in this study is: What happens if one runs such species tree inference methods on data where paralogy is present, in addition to or without ILS being present? Through simulation studies and analyses of two large biological data sets, we show that running such methods on data with paralogs can still provide accurate results. We use multiple different methods, some of which are based directly on the multispecies coalescent (MSC) model, and some of which have been proven to be statistically consistent under it. We also treat the paralogous loci in multiple ways: from explicitly denoting them as paralogs, to randomly selecting one copy per species. In all cases the inferred species trees are as accurate as equivalent analyses using single-copy orthologs. Our results have significant implications for the use of ILS-aware phylogenomic analyses, demonstrating that they do not have to be restricted to single-copy loci. This will greatly increase the amount of data that can be used for phylogenetic inference.</p>
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Open the record for dataset details and reuse information.
Estimating accurate gene trees in the presence of intra-locus recombination: A simulation study
Open the record for dataset details and reuse information.
Data from: SimPhy: phylogenomic simulation of gene, locus and species trees
Open the record for dataset details and reuse information.
Data from: Using genomic location and coalescent simulation to investigate gene tree discordance in Medicago L.
Open the record for dataset details and reuse information.
Data from: The influence of gene flow on species tree estimation: a simulation study
Open the record for dataset details and reuse information.
SimPhy configuration scripts for simulations reported in the study titled: Species tree inference methods intended to deal with incomplete lineage sorting are robust to the presence of paralogs
Open the record for dataset details and reuse information.
Macrosystems VIDA Tree Growth Simulation - 100 by 100 World 30 Species
Patterns of biodiversity, such as the increase toward the tropics and the peaked curve during ecological succession, are fundamental phenomena for ecology. Such patterns have multiple, interacting causes, but temperature emerges as a dominant factor across organisms from microbes to trees and mammals, and across terrestrial, marine, and freshwater environments. However, there is little consensus on the underlying mechanisms, even as global temperatures increase and the need to predict their effects becomes more pressing. The purpose of this project is to generate and test theory for how temperature impacts biodiversity through its effect on biochemical processes and metabolic rate. A combination of standardized surveys in the field and controlled experiments in the field and laboratory measure diversity of three taxa -- trees, invertebrates, and microbes -- and key biogeochemical processes of decomposition in seven forests distributed along a geographic gradient of increasing temperature from cold temperate to warm tropical. This dataset was based on simulations run by VIDA, a software suite that attempts to model the growth of individual trees using empirically derived--or randomly chosen--values for use with allometric relationships. By modeling the behavior of an individual tree, it is possible to model population dynamics in a spatially explicit simulation space. The modeling was done by Sean Hammond at The Brown Lab (PI, Jim Brown) at the University of New Mexico as part of a macrosystems biodiversity and latitude project supported by the National Science Foundation under Cooperative Agreement DEB#1065836.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.