Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
248
datasets available to search
ShareScore release 0.9.0
Dataset results
248 results for “tree of life”
Fruit, seed dispersal, and life history traits of tropical rainforest trees of the Anamalai Hills, Western Ghats, India
<p>This dataset contains compiled Fruit, seed dispersal, and life history traits of tropical rainforest trees of the Anamalai Hills, Western Ghats, India. The list of species included are mainly from the following two related publications:<br>- Muthuramkumar, S., Ayyappan, N., Parthasarathy, N., Mudappa, D., Raman, T.R.S., Selwyn, M.A. and Pragasan, L.A. (2006), <a href="https://doi.org/10.1111/j.1744-7429.2006.00118.x">Plant Community Structure in Tropical Rain Forest Fragments of the Western Ghats, India</a>. <em>Biotropica</em>, 38: 143-160. https://doi.org/10.1111/j.1744-7429.2006.00118.x<br>- Osuri, A., Chakravarthy, D., Mudappa, D., Raman, T., Ayyappan, N., Muthuramkumar, S., & Parthasarathy, N. (2017). <a href="http://httpd//doi.org/10.1017/S0266467417000219">Successional status, seed dispersal mode and overstorey species influence tree regeneration in tropical rain-forest fragments in Western Ghats, India</a>. <em>Journal of Tropical Ecology</em>, 33(4), 270-284. doi:10.1017/S0266467417000219<br>The present dataset is an expanded and updated version of the related dataset available at <a href="https://doi.org/10.5061/dryad.vd0nn">https://doi.org/10.5061/dryad.vd0nn</a><br> <br>Species traits information was collated from <a href="http://www.biotik.org/">BIOTIK (http://www.biotik.org/</a>), <a href="http://www.flowersofindia.net/">Flowers of India (http://www.flowersofindia.net/)</a>, India Biodiversity Portal (http://indiabiodiversity.org/), <a href="https://doi.org/10.5061/dryad.234/1">Global wood density database (https://doi.org/10.5061/dryad.234/1)</a> and <a href="https://doi.org/10.1017/S0266467417000219">Osuri et al. (2014): https://doi.org/10.1017/S0266467417000219</a>. We also referred to the following previous studies that provided information on the successional status of rain-forest species in the Western Ghats (Chetana 2013, Pascal 1988, Raman et al. 2009, Sreejith 2005).</p> <p><strong>References:</strong><br>CHETANA, H. C. 2013. Assessing the ecological processes in abandoned tea plantations and its implication for ecological restoration in the Western Ghats, India. PhD thesis, Manipal University.<br>OSURI, A. M., KUMAR, V. S. & SANKARAN, M. 2014. Altered stand structure and tree allometry reduce carbon storage in evergreen forest fragments in India’s Western Ghats. <em>Forest Ecology and Management </em>329: 375–383.<br>PASCAL, J. P. 1988. <em>Wet evergreen forests of the Western Ghats of India: Ecology, structure, floristic composition and succession</em>. Institut Français de Pondichéry, Pondicherry.<br>RAMAN, T. R. S., MUDAPPA, D. & KAPOOR, V. 2009. Restoring rainforest fragments: survival of mixed-native species seedlings under contrasting site conditions in the Western Ghats, India. <em>Restoration Ecology</em> 17:137–147.<br>SREEJITH, K. A. 2005. Ecological and ecophysiological studies on the successional status of tree seedlings in tropical wet evergreen and semi-evergreen forests of Kerala. PhD thesis, Forest Research Institute, Dehradun.</p> <p><strong>Geographic Coverage:</strong><br>1. Location/Study Area: Valparai Plateau, Tamil Nadu, India; Anamalai Tiger Reserve, Tamil Nadu, India<br>2. GPS coordinates: Valparai Plateau (10°15'- 10°22'N, 76°52' - 76°59'E); Anamalai Tiger Reserve (10°12' - 10°35'N, 76°49' - 77°24'E)</p> <p><strong>Temporal Coverage:</strong><br>1. Begins: 2003-03-01 (Year, Month, Day)<br>2. Ends: 2024-02-10 (Year, Month, Day)</p> <p>Besides the <strong>README.txt</strong> file, the dataset includes the following comma-delimited text (csv) file with the data in columns as explained below:</p> <p><strong>Anamalai_tree_traits_2024.csv</strong></p> <p><strong>spec_name_ORIG:</strong> Scientific name of the species used during the data collection<br><strong>genus:</strong> Genus of the taxon<br><strong>specificEpithet:</strong> Specific epithet of the taxon in the Latin binomial name<br><strong>Accept_name_WFO:</strong> Updated scientific name of the species as in Plants of the World Online (POWO, https://powo.science.kew.org/)<br><strong>Habit:</strong> life form of the species(tree/shrub/cane/palm)<br><strong>Distribution:</strong> Distribution of the species in the study area (Native/Endemic/Introduced)<br><strong>IUCN_status:</strong> IUCN status of the species (CR-Critically Endangered,DD-Data deficient,EN-Endangered,LC-Least Concern,NT-Near Threatened,VU-Vulnerable,NA-Unknown)<br><strong>Wden_final:</strong> Wood density value assigned for the species (g cm^-3); NA - not available; sourced from Global wood density database (https://doi.org/10.5061/dryad.234/1)<br><strong>wd_level:</strong> Level in which the wood density value belongs (Species - wood density value is from species level; genus - wood density value assigned is the genus level average value)<br><strong>fruit_type:</strong> Morphological type of fruit<br><strong>fleshy_dry:</strong> Whether fruit is a dry fruit or fleshy, with aril or other parts <br><strong>seed_size:</strong> Species seed size: L = Large (>3 cm); M = Medium (1-3 cm); S = Small (<1 cm)<br><strong>disperser:</strong> Categories indicating seed dispersal mode: Bird, mammal, bird and mammal (Mammal_bird), gravity, wind, or unknown<br><strong>habitat:</strong> Habitat affinity category: EG_edg - evergreen forest edge; EG_for - evergreen forest; Dec_for - deciduous forest; Int – Introduced species; Unknown – Unknown<br><strong>habt_new:</strong> Habitat affinity new category: Mature – mature forest; Secondary – secondary forest, NA - unknown/Introduced species<br><strong>ad_ht:</strong> Species maximum adult height (m)</p>
Life History Traits of Resprouting Puerto Rican Tropical Dry Forest Trees, Guánica Forest, 1981-2018
This dataset provides trait and demographic data for 44 tropical dry forest tree species from the Guánica State Forest in southwest Puerto Rico. The study area spans 4,500 ha of semi-deciduous TDF, where the sampled species represent over 90% of all individuals with a diameter at breast height (dbh) ≥2.5 cm. The dataset integrates ten functional traits, combining newly collected measurements (2017–2018) with previously published data (Vargas et al. 2021b). Previously published data includes xylem-specific hydraulic conductivity (ks), Huber value (hv), and hydraulic safety margin (HSM), with species-level data availability ranging from 19 to 44 species, except for HSM, which was measured for six species. Trait measurements were primarily collected during the wet season (August–November), except stomatal behaviour traits (psimax, psidv, and gsmax), which were assessed during the winter dry season before leaf fall. Demographic data encompass species-specific growth rates and annual survival rates for adult trees, derived from four permanent census plots (625 m² to 10,000 m²) distributed across the forest. These plots, established in mature upland TDF on limestone substrates with mollisol soils, were monitored between 1992 and 2019. Growth rate estimates are based on diameter increments recorded at regular censuses over 20.4–26.4 years. Survival rates were calculated over a 21-year period (1998–2019), mitigating the influence of extreme drought events. Standardised measurement protocols ensured data consistency, including repeated diameter assessments at multiple stem locations and the exclusion of wet-season measurements to prevent water-related swelling artifacts. Growth rates were derived from the regression slope of dbh against time, incorporating a minimum of two dbh measurements per individual (following Poorter et al. 2010). Annual survival rate was calculated over a 21-year timespan (1998–2019) to avoid bias introduced by an intense drought in 1997. The followin
MATEdb2, a Collection of High-Quality Metazoan Proteomes across the Animal Tree of Life to Speed Up Phylogenomic Studies
<p>Recent advances in high-throughput sequencing have exponentially increased the number of genomic data available for animals (Metazoa) in the last decades, with high-quality chromosome-level genomes being published almost daily. Nevertheless, generating a new genome is not an easy task due to the high cost of genome sequencing, the high complexity of assembly, and the lack of standardized protocols for genome annotation. The lack of consensus in the annotation and publication of genome files hinders research by making researchers lose time in reformatting the files for their purposes but can also reduce the quality of the genetic repertoire for an evolutionary study. Thus, the use of transcriptomes obtained using the same pipeline as a proxy for the genetic content of species remains a valuable resource that is easier to obtain, cheaper, and more comparable than genomes. In a previous study, we presented the Metazoan Assemblies from Transcriptomic Ensembles database (MATEdb), a repository of high-quality transcriptomic and genomic data for the two most diverse animal phyla, Arthropoda and Mollusca. Here, we present the newest version of MATEdb (MATEdb2) that overcomes some of the previous limitations of our database: (i) we include data from all animal phyla where public data are available, and (ii) we provide gene annotations extracted from the original GFF genome files using the same pipeline. In total, we provide proteomes inferred from high-quality transcriptomic or genomic data for almost 1,000 animal species, including the longest isoforms, all isoforms, and functional annotation based on sequence homology and protein language models, as well as the embedding representations of the sequences. We believe this new version of MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open, greener, and collaborative science.</p>
The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life. Data file underpinning Figure 2A and Figure 2B
<div>These datasheets accompany the article "The Earth BioGenome Project Phase II: Illuminating the Eukaryotic Tree of Life" in Frontiers in Science</div> <div>This file contains data processed from Catalog of Life on 31 December 2023. The catalog was downloaded and post-processed to</div> <div>remove prokaryotic taxa</div> <div>remove extinct and fossil taxa</div> <div>remove taxon names that were listed as junior synonyms</div> <div>remove taxon names listed as "invalid"</div> <div>Total living, valid eukaryotic genera 167,085</div> <div>The taxa were sorted by the nomenclatorial Code under which they were declared (to avoid namespace clashes)</div> <div>International Code for Algae, Fungi and Plants https://www.iapt-taxon.org/nomen/main.php</div> <div>Algal, Fungal, Plant code genera 31,076</div> <div>International Code of Zoological Nomenclature https://www.iczn.org/the-code/the-code-online/</div> <div>Zoological code genera 136,009</div> <div>The Code-sorted taxa were aggregated by the generic portion of their names, and two plots were generated:</div> <div>a plot aggregating the cumulative number of species in genera sorted by species number (Figure 2A)</div> <div>a plot illustrating the distribution of the size of genera (Figure 2B)</div> <div>This data file gives access to these processed data for</div> <div>Figure 2 A Data</div> <div>Figure 2 B Data</div> <div>The original data including the intermediate calculations of values, and the plotted graphs, are available as a GoogleDoc at https://docs.google.com/spreadsheets/d/1V-bTtWjIRasC3AgID0jGlyToKqI-H9h1aPeSxUpNrjk/edit?usp=sharing</div>
Data for: The structure of evolutionary model space for proteins across the tree of life
<p>Supporting data for "The structure of evolutionary model space for proteins across the tree of life," submitted by GE Scolaro and EL Braun. The data files correspond to three gzipped tarballs including protein multiple sequence alignments, PAML format models of protein evolution, and model fit data; see included README for details.</p>
NEXUS file describing the taxonomic relationships of the 466 species for which genome sequencing was underway at Tree of Life, Wellcome Sanger Institute, at 31 December 2020
<p>This NEXUS file shows the taxonomic relationships of 466 species of eukaryote. The taxonomy derives from the NCBI TaxonomyDB. The species are those for which genome sequencing is underway at the Tree of Life programme, Wellcome Sanger Institute, as of 31st Decemnber 2020. The NEXUS file includes a figtree block generated in FigTree [<strong><a href="https://github.com/rambaut/figtree">https://github.com/rambaut/figtree</a>] </strong>that informs display of the data as a circular tree with species coloured by taxonomic Family, and Families with more than one species represented as triangles. The figure is used in publications and presentations describing the activities of the Tree of Life programme and the projects in which Tree of Life is involved, especially the Darwin Tree of Life project [https://darwintreeoflife.org].</p>
Data from: Enriching the ant tree of life: enhanced UCE bait set for genome-scale phylogenetics of ants and other Hymenoptera
1. Targeted enrichment of conserved genomic regions (e.g., ultraconserved elements or UCEs) has emerged as a promising tool for inferring evolutionary history in many organismal groups. Because the UCE approach is still relatively new, much remains to be learned about how best to identify UCE loci and design baits to enrich them. 2. We test an updated UCE identification and bait design workflow for the insect order Hymenoptera, with a particular focus on ants. The new strategy augments a previous bait design for Hymenoptera by (a) changing the parameters by which conserved genomic regions are identified and retained, and (b) increasing the number of genomes used for locus identification and bait design. We perform in vitro validation of the approach in ants by synthesizing an ant-specific bait set that targets UCE loci and a set of "legacy" phylogenetic markers. Using this bait set, we generate new data for 84 taxa (16/17 ant subfamilies) and extract loci from an additional 17 genome-enabled taxa. We then use these data to examine UCE capture success and phylogenetic performance across ants. We also test the workability of extracting legacy markers from enriched samples and combining the data with published data sets. 3. The updated bait design (hym-v2) contained a total of 2,590-targeted UCE loci for Hymenoptera, significantly increasing the number of loci relative to the original bait set (hym-v1; 1,510 loci). Across 38 genome-enabled Hymenoptera and 84 enriched samples, experiments demonstrated a high and unbiased capture success rate, with the mean locus enrichment rate being 2,214 loci per sample. Phylogenomic analyses of ants produced a robust tree that included strong support for previously uncertain relationships. Complementing the UCE results, we successfully enriched legacy markers, combined the data with published Sanger data sets, and generated a comprehensive ant phylogeny containing 1,060 terminals. 4. Overall, the new UCE bait design strategy resulted in an enhanced bait set for genome-scale phylogenetics in ants and likely all of Hymenoptera. Our in vitro tests demonstrate the utility of the updated design workflow, providing evidence that this approach could be applied to any organismal group with available genomic information.
Supplementary Data for: A time-calibrated 'Tree of Life' of aquatic insects for knitting historical patterns of evolution and measuring extant phylogenetic biodiversity across the world
<p>This compendium of files includes the dated phylogenetic tree in Newick format (<strong>Data S1</strong>), the list of statistical routines used for the three empirical case studies (<strong>Data S2</strong>), and the high-resolution version of the figures in the supplementary materials and main text (<strong>Data S3</strong>) for the <em>Earth-Science Reviews</em> paper "A time-calibrated ‘Tree of Life’ of aquatic insects for knitting historical patterns of evolution and measuring extant phylogenetic biodiversity across the world", which is under consideration. The best-scoring molecular tree (<strong>Data S1</strong>) can be opened using freely available programs like R (R Development Core Team, 2021), Dendroscope (Huson and Scornavacca, 2012), and FigTree (Rambaut, 2018).</p> <p>Please, feel free to send an email to the maintainer Dr. Jorge García Girón (jogarg@unileon.es OR Jorge.Garcia-Giron@oulu.fi) if you face any trouble downloading, opening, or using these files.</p> <ul> <li>Huson, D. H., & Scornavacca, C. (2012). Dendroscope 3: An interactive tool for rooted phylogenetic trees and networks. <em>Systematic Biology</em>, <em>61(6)</em>, 1061–1067.</li> <li>R Development Core Team (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/</li> <li>Rambaut, A. (2018). FigTree. Institute of Evolutionary Biology, University of Edinburgh, Edinburgh, UK. http://tree.bio.ed.ac.uk/software/figtree/</li> </ul>
Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages
<p><strong>Abstract</strong></p> <p>Advances in phylogenetic methods and high-throughput sequencing have allowed the reconstruction of deep phylogenetic relationships in the evolutionary history of eukaryotes. Yet, the root of the eukaryotic tree of life remains elusive. The most ‘popular’ (i.e. in textbooks and reviews) hypothesis for the root is between Unikonta (Opisthokonta + Amoebozoa) and Bikonta (all other eukaryotes), which emerged from analyses of a single gene fusion and a limited sampling of eukaryotic lineages. Subsequent highly-cited studies based on concatenation of genes supported this hypothesis with some variations or proposed a root within the Excavata. However, concatenation of genes neither considers phylogenetically-informative events (i.e. gene duplications and losses), nor provides an estimate of the root. A more recent study using gene tree-species tree reconciliation methods suggested the root lies between Opisthokonta and all other eukaryotes, but only including 59 taxa and 20 genes. Here we apply a gene tree – species tree reconciliation approach to a gene-rich and taxon-rich dataset (i.e. 2,786 gene families from two sets of ~158 diverse eukaryotic lineages) to assess the root, and we iterate each analysis 100 times to quantify tree space uncertainty. Our results estimate a root between Fungi and all other eukaryotes, or between Opisthokonta and all other eukaryotes, and reject alternative popular roots from the literature. Based on further analysis of genome size we propose Opisthokonta + others as the most likely root. Finding the root of the eukaryotic tree of life is critical for the field of comparative biology as it allows us to understand the timing and mode of evolution of characters across the evolutionary history of eukaryotes.</p> <p>Methods<br> Here we provide the alignments, gene trees, inputs, and outputs from our project entitled "Phylogenomic Analyses of 2,786 Genes in 158 Lineages Support a Root of The Eukaryotic Tree of Life Between Opisthokonts and All Other Lineages". Sequences and alignments were produced using the phylogenomic pipeline PhyloToL, which contains a taxon- and gene-rich database (including eukaryotes, archaea, and bacteria). These data were then used for 1) assessing the root of the eukaryotes and 2) for comparison with other previously published hypotheses. In both cases, we used the species tree - gene tree reconciliation tool iGTP. We also did a comparison of hypotheses using the likelihood-based tool SpeciesRax</p> <p><strong>iGTP</strong></p> <p>Input data</p> <p>These data are divided into four datasets based on taxa selection. For dataset SEL+, taxa were selected based on their taxonomy; for RAN+, taxa were selected randomly among the major eukaryotic clades Opisthokonta, Amoebozoa, Archaeplastida, Excavata, SAR, and some orphan lineages. Datasets SEL- and RAN- are the same as SEL+ and RAN+, but exclude microsporidians in order to account for and avoid long branch attraction due to microsporidians fast-evolutionary rates. We chose the gene families that contain at least 25 taxa representing at least four of the five major eukaryotic clades. Additionally, at least 2 of the major clades had to contain at least 2 minor clades (e.g. Glaucophytes and Rhodophyta are minor clades in the major clade Archaeplastida). In a pilot analysis, we produced an alignment and a phylogenetic tree for each gene family using the default settings of a previous version of PhyloToL (GUIDANCE V1.3.1 sequence cutoff = 0.3 and column cutoff = 0.4; RAxML quick tree with model PROTGAMMALG and no bootstraps). Then, we kept the gene families that are exclusive of eukaryotes or the ones in which eukaryotes were monophyletic. From a total of 3,002 gene families that met our criteria, 2786 passed the initial steps of PhyloToL when including only the data from the dataset SEL+. These 2,786 gene families were used for further analyses with all datasets.</p> <p>MSAs were produced with PhyloToL (GUIDANCE V2.02 sequence cutoff = 0.3, column cutoff = 0.4, number of iterations = 5; Sela, et al. 2015). The default parameters of PhyloToL include up to five iterations of GUIDANCE V2.02 with 10 bootstraps and MAFFT V7 with algorithm E-INS-i for less than 200 sequences or “auto” option if more than 200 sequences, and maxiterate = 1000. Instead, here we run up to five iterations of GUIDANCE with 20 bootstraps and the simple MAFFT algorithm FFT-NS-2. Then, we perform an additional GUIDANCE run with 100 bootstraps and the default MAFFT parameters for PhyloToL.</p> <p>Gene trees were inferred with RAxML v.8.2.4 with 10 ML searches for best-ML tree (option "-# 10"), using the rapid hill-climbing algorithm (option "-f d") and no bootstrap replicates. The protein evolution model used was evaluated during the gene tree inference (option "-m PROTCATAUTO") by testing all models available in RAxML (e.g. JTT, LG, WAG, etc) with optimization of substitution rates and of site-specific evolutionary rates which were categorized into four distinct rate categories for greater computational efficiency.</p> <p>We ran 100 repetitions of iGTP analyses per dataset. But, given the complexity of the datasets and the heuristic nature of some key steps of the iGTP algorithm (e.g. gene tree rooting and initial starting species tree generation), in a preliminary analysis, we faced two systematic challenges with iGTP as the inferred species tree was affected by: 1) the order of the leaves in the input unrooted gene tree Newick strings (i.e. the input trees were treated as rooted even though we specified that they were not); and 2) the input gene order in the 100 replicates. Therefore, we randomly shuffled the order of the leaves in the unrooted gene trees (keeping the same topology), and randomly shuffled the order of the input gene trees in each of the 100 replicates per dataset. Here we provide the 100 input files generated for those iGTP analyses. </p> <p>Output data</p> <p>Here we also share the data generated after two analyses: 1) root assessment and 2) hypothesis testing. For the former, we allowed iGTP to calculate the more parsimonious root given our input files. For the latter, we allowed iGTP to calculate the reconciliation cost of the gene trees given the input files and constraints in the species trees to reflect previously published root hypotheses. The constraints are explained in the README file.</p> <p>SpeciesRax</p> <p>Since we removed LGT and contamination from our dataset using a series of filters, we applied the model UndatedDL instead of UndatedDTL, which implies that we only took into consideration duplications and losses and ignored the transferences. Then, the command used for SpeciesRax was...</p> <p>./generax --families forGeneRax/famFile --species-tree forGeneRax/spsTreeAn.newick --strategy SKIP --rec-model UndatedDL --per-family-rates --prefix An_DS1 --si-strategy EVAL </p> <p>input </p> <p>Here, we are sharing all the necessary files to run SpeciesRax, including the mapping file (famFile), the species trees (spsTrees; the best iGTP constrained species trees per hypothesis), and their underlying gene trees (trees_r)</p> <p>output</p> <p>Output folder from SpeciesRax, which includes log files, events (duplications, losses) counts, and statistics (i.e., reconciliation likelihood values)</p> <p>Note: <br> * As in the iGTP analyses, for the SpeciesRax files, the words Op, Fu, Di, Un, An, refer to the five root hypotheses compared: Opisthokonta-others, Fungi-others, Discoba-others, Unikonta-Bikonta, and (Ancyromonadida + Metamonada)-others, respectively. <br> * For the SpeciesRax analyses, the Un word refers to Ut (Cavalier-Smith 2003) <br> </p>
DateLife: leveraging databases and analytical tools to reveal the dated Tree of Life
<p>Achieving a high-quality reconstruction of a phylogenetic tree with branch lengths proportional to absolute time (chronogram) is a difficult and time-consuming task. But the increased availability of fossil and molecular data, and time-efficient analytical techniques has resulted in many recent publications of large chronograms for a large number and wide diversity of organisms. Knowledge of the evolutionary time frame of organisms is key for research in the natural sciences. It also represent valuable information for education, science communication, and policy decisions. When chronograms are shared in public, open databases, this wealth of expertly-curated and peer-reviewed data on evolutionary timeframe is exposed in a programatic and reusable way, as intensive and localized efforts have improved data sharing practices, as well as incentivizited open science in biology. Here we present DateLife, a service implemented as an R package and an R Shiny website application available at www.datelife.org, that provides functionalities for efficient and easy finding, summary, reuse, and reanalysis of expert, peer-reviewed, public data on time frame of evolution. The main DateLife workflow constructs a chronogram for any given combination of taxon names by searching a local chronogram database constructed and curated from the Open Tree of Life Phylesystem phylogenetic database, which incorporates phylogenetic data from the TreeBASE database as well. We implement and test methods for summarizing time data from multiple source chronograms using supertree and congruification algorithms, and using age data extracted from source chronograms as secondary calibration points to add branch lengths proportional to absolute time to a tree topology. DateLife will be useful to increase awareness of the existing variation in alternative hypothesis of evolutionary time for the same organisms, and can foster exploration of the effect of alternative evolutionary timing hypotheses on the results of downstream analyses, providing a framework for a more informed interpretation of evolutionary results.</p>
Supporting Information for: Single-fly genome assemblies fill major phylogenomic gaps across the Drosophilidae Tree of Life
<p>This data repository contains supporting information, data, and code for figures and analysis pipelines the PLOS Biology article: "Single-fly genome assemblies fill major phylogenomic gaps across the Drosophilidae Tree of Life."</p> <ul> <li><strong>4d_full.treefile</strong>: Data underlying Figure 1 (note: tree was plotted as a cladogram and key groups collapsed on iToL; the treefile was not modified) and Figure S1.</li> <li><strong>S2_data.csv</strong>: Data underlying Figure 2.</li> <li><strong>S3_data.csv</strong>: Data underlying Figure 3.</li> <li>Data underlying Figure 4 is found in Table S4 of supplementary_tables.xlsx in the main manuscript</li> <li><strong>S5_data.csv</strong> Data underlying Figure 5</li> <li><strong>S6_data.csv</strong> Data underlying Figure S2 </li> <li><strong>illumina_only_assms.tar.gz</strong>: Archive of Illumina-only assemblies (FASTA) based on publicy available data that we did not generate. Assemblies generated from our own short-read data have been submitted to NCBI GenBank.</li> <li><strong>illumina_vcfs.tar.gz</strong>: Illumina-based variant calls and BED tracks of masked bases.</li> <li><strong>genomes.tar.gz</strong>: Genome files, for archival purposes.</li> <li><strong>repeatModeler-lib.tar.gz</strong>: RepeatModeler2 libraries.</li> <li><strong>diploid_genomes.tar.gz</strong>: diploid genomes and BED tracks of phased regions.</li> <li><strong>trees.tar.gz</strong>: phylogenies</li> </ul>
Data from: Meta-analytical evidence for frequency-dependent selection across the tree of life
<p>Explaining the maintenance of genetic variation in fitness related traits within populations is a fundamental challenge in ecology and evolutionary biology. Frequency-dependent selection (FDS) is one mechanism that can maintain such variation, especially when selection favours rare variants (negative FDS). However, our general knowledge about the occurrence of FDS, its strength and direction remain fragmented, limiting general inferences about this important evolutionary process. We systematically reviewed the published literature on FDS and assembled a database of 747 effect sizes from 101 studies to analyse the occurrence, strength, and direction of FDS, and the factors that could explain heterogeneity in FDS. Using a meta-analysis, we found that overall, FDS is more commonly negative, although not significantly when accounting for phylogeny. An analysis of absolute values of effect sizes, however, revealed the widespread occurrence of modest FDS. However, negative FDS was only significant in laboratory experiments and non-significant in mesocosms and field-based studies. Moreover, negative FDS was stronger in studies measuring fecundity and involving resource competition over studies using other fitness components or focused on other ecological interactions. Our study unveils key general patterns of FDS and points in future promising research directions that can help us understand a long-standing fundamental problem in evolutionary biology and its consequences for demography and ecological dynamics.</p>
Eye morphology contributes to the ecology and evolution of the avian tree of life
<p>The avian eye is the single most important external anatomical trait for interpreting light environments by birds and varies widely in size and shape across the avian tree of life. The attached dataset provides measurements on eye size taken from preserved museum specimens for roughly one third of the avian tree of life (N = 3,475 species). The original dataset was collected by Stanley Ritland and Alice Hutchinson and archived as appendices in Stanley Ritland's Dissertation from the University of Chicago (1982): "The Allometry of the Vertebrate Eye".</p>
Fig. 40 in The Amphibian Tree Of Life
Fig. 40. Maximumlikelihood tree of ranoids of Delorme et al. (2004), based on sequences 12S and 16S rRNA for a total of 1198 bp. Alignment was made using the program SeAl (1995; cost functions not provided) and by comparison with models of secondary structure. treated as missing data. The maximumlikelihood nucleotide substitution model accepted was
Fig. 66. A in The Amphibian Tree Of Life
Fig. 66. A simplied tree of our results (fig. 50) tree showing families. Numbers on branches branch lengths, Bremer, and jackknife values, as well as molecular synapomorphies to be appendices 4 and 5. See table 5 for taxon names associated with internal numbered branches
Fig. 20 in The Amphibian Tree Of Life
Fig. 20. Consensus of 12 equally parsimonious trees of selected members of Megophryidae Delorme and Dubois (2001), rooted on Scaphiopus and Pelodytes. Underlying data were 54
Fig. 10 in The Amphibian Tree Of Life
Fig. 10. Parsimony tree of Plethodontidae by Macey (2005), a reanalysis of entire mt DNA sequence data provided by Mueller et al. (2004). On right are the traditional taxonomy and
Fig. 9 in The Amphibian Tree Of Life
Fig. 9. Tree of Plethodontidae by Mueller et al. (2004), with the traditional taxonomic (Desmognathinae 1 tribes of Plethodontinae; Wake, 1966) placed on the right, with taxonomic
Fig. 7 in The Amphibian Tree Of Life
Fig. 7. Relationships of salamanders suggested by Wiens et al. (2005). Families are noted Results reflect a parsimony analysis of 326 character transformations of morphology (221 informative), and DNA sequences from nu rRNA (212 bp from Larson, 1991; 147
Fig. 47. A in The Amphibian Tree Of Life
Fig. 47. A, Rhacophorid and mantellid Liem (1970) based on 36 direct to dendritic phological transformation series, rooted pothetical generalized ranid ancestor.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.