Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18,140

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

18,140 results for “Phylogenetics”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area"

<p>This repository archives the dataset of the article &quot;Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area&quot;. The cognate annotation of Tshangla, Kho-Bwa, Hrusish, Mishmic, and Tani languages were done by us. The cognate decision on the other languages was annotated by Sagart et al. (2019).&nbsp;&nbsp;Please use the following information to cite our work:&nbsp;<br> Wu, M.-S, Bodt, T. A, Tresoldi, T. (2022). &nbsp;Bayesian phylogenetics illuminate shallower relationships Trans-Himalayan languages in the Tibet-Arunachal area. Linguistics of the Tibeto-Burman Area. [forthcoming]</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Darwin: an amino acid sequence collection of complete proteomes from eukaryotes with different phylogenetic affinities (v. 03_2020_137)

<p><strong>Background</strong></p> <p>Every time we find an interesting gene in an organism of interest, the first question is often &ldquo;how widely is this gene distributed in the eukaryotic kingdom?&rdquo;. Naturally, one could use NCBI BLAST search against the non-redundant sequence database provided by GenBank to answer this question. However, it can be cumbersome to parse the results and assign them to taxonomic units. It is also not straightforward to get an overview of which eukaryotic groups are represented in the results. Top BLAST hits can be crowded with sequences from closely-related organisms making it difficult gain an overview of the overall distribution across eukaryotes. To streamline this process, we developed an in-house database of complete eukaryotic proteomes. We tagged each sequence with a eukaryotic group handle (two-character symbol) and combined them into a single data set searchable by standalone BLAST on one&rsquo;s own computer. We named this data set &ldquo;Darwin&rdquo; to reflect the diverse nature of the sequences it contains.&nbsp;</p> <p><strong>Methods</strong></p> <p>We downloaded predicted proteomes in FASTA format from different sources such as GenBank, Joint Genome Institute (Depart of Energy, USA), Broad Institute (Massachusetts Institute of Technology, USA), Phytozome and a number of other specialized websites catering for a specific organism such as the Arabidopsis Information Resource (TAIR), or the Saccharomyces Genome Database (SGD). All the organisms we included in Darwin are listed in Table 1. To reduce redundancy, we took care not to include the same species more than once unless subspecies were known to show wide diversity. Each sequence header was tagged with a eukaryotic group handle composed of two-character symbols (based on Keeling&nbsp;<em>et al</em>., 2005). These handles clearly appear in BLAST output and can be parsed easily. We combined sequences from all proteomes into a single data set and named it &ldquo;Darwin&rdquo;.</p> <p><strong>Results</strong></p> <p>The current version of Darwin (v. 03_2020_137) contains 2,601,132 amino acid sequences from 137 eukaryotes (Table 1, Data file 1). The sizes of the proteomes were diverse, ranging from ~4000 sequences in some alveolates to 60,000-76,000 in plants. Darwin represents most of the supergroups of eukaryotic kingdom described in Keeling&nbsp;<em>et al.,</em>&nbsp;(2005) except those in Rhizaria whose genomes were not available at the time of data set construction. The data set contains larger numbers of proteomes from fungi and plants reflecting areas of interest in our group.&nbsp;</p> <p><strong>Conclusions</strong></p> <p>Darwin is provided as a text fasta file that can be formatted for BLAST searches on standalone computers. The results from the BLAST searches can be parsed to determine how widely a gene of interest is distributed among different eukaryotes. Simple counting of the eukaryotic group handles would also yield an overview of the distribution across taxa. Darwin is also useful for rapidly finding out whether a gene is missing in particular taxa.</p> <p><strong>Reference</strong></p> <p>Keeling PJ, Burger G, Durnford DG, Lang BF, Lee RW, Pearlman RE, Roger AJ, Gray MW (2005) The tree of eukaryotes.&nbsp;<em>Trends Ecol. Evol.</em>&nbsp;<strong>20:</strong>&nbsp;670-676</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Dataset for: Evaluating phylogenetic methods for quantifying risks and opportunities presented by forks in open source software (master dissertation).

<p>This is the data for my master dissertation [1]. If you wish to get a copy, download it from Zenodo and open docs/master.pdf.</p> <p>Data acquisition and encoding techniques are described in paragraph 3.1.1 (table 3.1).</p> <p>The data is described in more detail in paragraph 4.1 (table 4.2).</p> <p>* fork1_all.csv: MySQL server / MariaDB server<br> * fork2_all.csv: Linux kernel / Android kernel<br> * fork3_all.csv: Apache OpenOffice / LibreOffice</p> <p>==Cite==<br> [1] A. Ortiz-Troncoso. Evaluating phylogenetic methods for quantifying risks and opportunities presented<br> by forks in open source software (master dissertation). Zenodo, 2018. doi: http://doi.org/10.5281/zenodo.1158292</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024

<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (&ldquo;full&rdquo; matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (&ldquo;reduced&rdquo; matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the &ldquo;new technology&rdquo; search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Criteria for prioritizing selection of Mexican maize landrace accessions for conservation in situ or ex situ based on phylogenetic analysis

<p>Data for processed SSR markers in maize accessions. A database in Structured Query Language (SQL) is provided. Please see the text file &quot;READMEmaizeSSR.pdf&quot;.</p>

opencc-by-4.0Dec 2022View details →
zenodo48/100

Aligned bam files for "Phylogenetic modeling of enhancer shifts in mole-rats reveals regulatory changes associated with tissue-specific traits"

<p>Aligned bam files used for analysis in&nbsp;&quot;Phylogenetic modeling of enhancer shifts in mole-rats reveals regulatory changes associated with tissue-specific traits&quot;.</p> <p>This is an accompanying dataset to&nbsp;Datasets and code for &quot;Phylogenetic modeling of enhancer shifts in mole-rats reveals regulatory changes associated with tissue-specific traits&quot; (https://zenodo.org/record/7442105).</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Phylogenetic and epidemiologic data relating to age-specific HIV incidence and transmission in Rakai, Uganda, 2003-2018.

<p>This repository contains the data for the analyses presented in the paper Growing gender inequity in HIV infection in Africa: sources and policy implications by M. Monod, A. Brizzi, R. Galiwango, R. Ssekubugu, Y. Chen, X. Xi et al. available in the pre-print&nbsp;<a href="https://doi.org/10.1101/2023.03.16.23287351">https://doi.org/10.1101/2023.03.16.23287351</a>&nbsp;</p> <p>We thank all contributors, program staff and participants to the Rakai Community Cohort Study; all members of the PANGEA-HIV consortium, the <a href="https://www.rhsp.org/index.php">Rakai Health Sciences Program</a>, and CDC Uganda for comments on an earlier version of the manuscript.</p> <p>We also extend our gratitude to the <a href="https://doi.org/10.14469/hpc/2232">Imperial College Research Computing Service</a> and the <a href="https://www.bdi.ox.ac.uk/about/biomedical-research-computing">Biomedical Research Computing Cluster</a> at the University of Oxford for providing the computational resources to perform this study. Additionally, we thank the Office of Cyberinfrastructure and Computational Biology at the <a href="https://www.niaid.nih.gov/">National Institute for Allergy and Infectious Diseases</a> for data management support; and Zulip for sponsoring team communications through the Zulip Cloud Standard chat app.&nbsp;</p> <p>All analysis code is available from <a href="https://github.com/MLGlobalHealth/phyloSI-RakaiAgeGender">https://github.com/MLGlobalHealth/phyloSI-RakaiAgeGender</a>.</p>

opencc-by-4.0Mar 2023View details →
edi48/100

The role of riparian functional and phylogenetic diversity on leaf litter processing in rivers

While taxonomic diversity mediates changes in ecosystem function is well-studied, how deeper dimensions of biodiversity, specifically phylogenetic and functional, independent of taxonomic diversity, drive important processes is understudied. The overarching goal of this work was to determine the role of these dimensions of biodiversity independently and/or interactively explain carbon processing in rivers. Here, we explicitly link community structure and subsequent traits of riparian forests to adjacent ecosystem processing of carbon (e.g., leaf litter). This was accomplished by examining how forests are actually structured in addition to experimental manipulations of phylogenetic and functional diversities of riparian forest community inputs of leaf litter to streams. Experimental field manipulations were carried out in three Piedmont headwater streams to answer the following questions: (1) Does existing variation in taxonomic, functional and phylogenetic diversity of riparian communities differentially drive decomposition in rivers? And (2) Independent of taxonomic diversity, how does functional and phylogenetic diversity of leaf litter assemblages influence rates of decomposition in rivers? We observed significant interspecific variation in breakdown among 30 riparian tree species, in addition to significant relationships between breakdown rate and important foliar tissue chemistries. Breakdown of mixtures that reflected the composition of the riparian species composition did not vary with functional nor phylogenetic diversity, but breakdown of litter mixtures was higher than that of single species. In a separate study, when manipulated independently, functional and phylogenetic diversity were positively related to breakdown, and explained similar degrees of variation. These results are important to understand in light of deepening knowledge of the role different dimensions of biodiversity take in explaining ecosystem function, as well as how these measures can b

openCC (other)Mar 2022View details →
edi48/100

Inventory of High-resolution phylogenetic profiles of the planktonic microbial communities (via 16S and 18S rRNA gene amplicons) from Shark River Slough and Taylor Slough, Everglades National Park (FCE LTER), Florida, USA, 2017 - ongoing

Planktonic microbial communities mediate many vital biogeochemical processes in wetland ecosystems, yet compared to other aquatic ecosystems, like oceans, lakes, rivers, or estuaries, they remain relatively underexplored. Our study site, the Florida Everglades (USA)—a vast iconic wetland consisting of a slow-moving system of shallow rivers connecting freshwater marshes with coastal mangrove forests and seagrass meadows—is a highly threatened model ecosystem for studying salinity and nutrient gradients, as well as the effects of sea level rise and saltwater intrusion. This dataset provides the first high-resolution phylogenetic profiles of planktonic bacterial and eukaryotic microbial communities (using 16S and 18S rRNA gene amplicons) from these environments. The dataset contains 16S and 18S rRNA data from 2017, and contains 16S rRNA data for monthly (2019) and quarterly water samples (2020-ongoing). The 2017 data are published in Laas et al. 2022. A detailed list of sequence data and their accession numbers in GenBank is provided and will be updated as more data are published. This data package is an inventory of sequence read archive (SRA) entries available through GenBank BioProject PRJNA525456 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA525456) and BioProject PRJNA1018945 (at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1018945). This data package is associated with the following publication: Laas, P., Ugarelli, K., Travieso, R., Stumpf, S., Gaiser, E. E., Kominoski, J. S., & Stingl, U. (2022). Water column microbial communities vary along salinity gradients in the Florida Coastal Everglades wetlands. Microorganisms, 10(2), 215. https://doi.org/10.3390/microorganisms10020215 Instead of citing this package, which is an inventory, please cite the original GenBank data or journal article, as appropriate. Citation guidance for the journal article is available on the respective publisher's website.

openCC (other)Feb 2024View details →
zenodo44/100

Patterns and drivers of species diversity in the Indo-Pacific red seaweed Portieria: phylogenetic data

<p>Alignments, trees and Biogeobears analyses related to the study: Leliaert F, Payo DA, Gurgel CFD, Schils T, Draisma SGA, Saunders GW, Kamiya M, Sherwood AR, Lin S-M, Huisman John&nbsp;M, Le Gall L, Anderson RJ, Bolton John&nbsp;J, Mattio L, Zubia M, Spokes T, Vieira C, Payri CE, Coppejans E, D&#39;hondt S, Verbruggen H, De Clerck O. Patterns and drivers of species diversity in the Indo-Pacific red seaweed Portieria. Journal of Biogeography. 2018;45(10):2299-313. doi:10.1111/jbi.13410</p> <p>Abstract: Biogeographical processes underlying Indo-Pacific biodiversity patterns have been relatively well studied in marine shallow water invertebrates and fishes, but have been explored much less extensively in seaweeds, despite these organisms often displaying markedly different patterns. Using the marine red alga Portieria as a model, we aim to gain understanding of the evolutionary processes generating seaweed biogeographical patterns. Our results will be evaluated and compared with known patterns and processes in animals. Species diversity estimates were inferred using DNA-based species delimitation methods. Historical biogeographical patterns were inferred based on a six-gene time-calibrated phylogeny, distribution data of 802 specimens, and probabilistic modelling of geographic range evolution. The importance of geographic isolation for speciation was further evaluated by population genetic analyses at the intraspecific level. We delimited 92 candidate species, most with restricted distributions, suggesting low dispersal capacity. Highest species diversity was found in the Indo-Malay Archipelago (IMA). Our phylogeny indicates that Portieria originated during the late Cretaceous in the area that is now the Central Indo-Pacific. The biogeographical history of Portieria includes repeated dispersal events to peripheral regions, followed by long-term persistence and diversification of lineages within those regions, and limited dispersal back to the IMA. Our results suggest that the long geological history of the IMA played an important role in shaping Portieria diversity. High species richness in the IMA resulted from a combination of speciation at small spatial scales, possibly as a result of increased regional habitat diversity from the Eocene onwards, and species accumulation via dispersal and/or island integration through tectonic movement. Our results are consistent with the biodiversity feedback model, in which biodiversity hotspots act as both &lsquo;centres of origin&rsquo; and &lsquo;centres of accumulation&rsquo;, and corroborate previous findings for invertebrates and fish that there is no single unifying model explaining the biological diversity within the IMA.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Supplementary material accompanying "Factoring lexical and phonetic phylogenetic characters from word lists"

<p>This repository contains the scripts and the data that were used to run the analyses for the paper &quot;Factoring lexical and phonetic phylogenetic characters from word lists&quot;. For details, please refer to the README.md file provided along with the dataset. If you run into problems replicating the analysis, please do not hesitate to contact the authors.</p>

opencc-by-4.0Nov 2015View details →
zenodo44/100

Supplement for "Using Phylogenetic Networks to Model Chinese Dialect History"

<p>This is the supplementary material accompanying the paper &quot;Using Phylogenetic Networks to Model Chinese Dialect History&quot;, which appeared in 2014 in &quot;Language Dynamics and Change&quot; (volume 4, issue 2).</p>

opencc-zeroAug 2014View details →
zenodo44/100

Phlorest phylogeny derived from Zhang et al 2019 'Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Zhang M, Yan S, Pan W, &amp; Jin L. 2019. Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic. Nature, 569, 112–115.</p> </blockquote>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Phlorest phylogeny derived from Kolipakam et al. 2018 'A Bayesian phylogenetic study of the Dravidian language family'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Kolipakam V, Jordan FM, Dunn M, Greenhill SJ, Bouckaert R, Gray RD &amp; Verkerk A. 2018. A Bayesian phylogenetic study of the Dravidian language family. R. Soc. Open Sci. 5: 171504.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Phlorest phylogeny derived from Chang et al. 2015 'Ancestry-constrained phylogenetic analysis supports the Indo-European steppe hypothesis'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Chang W, Cathcart C, Hall D, &amp; Garrett A. 2015. Ancestry-constrained phylogenetic analysis supports the Indo-European steppe hypothesis. Language, 91(1):194-244.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Phlorest phylogeny derived from Bowern & Atkinson 2012 'Computational phylogenetics and the internal structure of Pama-Nyungan'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Bowern C &amp; Atkinson QD. 2012. Computational phylogenetics and the internal structure of Pama-Nyungan. Language, 88(4), 817-845.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Phlorest phylogeny derived from Birchall et al. 2016 'A combined comparative and phylogenetic analysis of the Chapacuran language family'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall, Joshua, Michael Dunn, and Simon J. Greenhill. 2016. A combined comparative and phylogenetic analysis of the Chapacuran language family. International Journal of American Linguistics 82 (3): 255–84. doi: 10.1086/687383</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Phlorest phylogeny derived from Kitchen et al. 2009 'Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Kitchen A, Ehret C, Assefa S &amp; Mulligan CJ. 2009. Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Proceedings of the Royal Society B: Biological Sciences, 270(1668), 2703-2710.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Phlorest phylogeny derived from Lee & Hasegawa 2011 'Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee S, Hasegawa T (2011) Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proceedings of the Royal Society B: Biological Sciences, 278(1725):3662–9.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Supporting Data: ontophylo: Reconstructing the evolutionary dynamics of phenomes using new ontology-informed phylogenetic methods

<p>This dataset contains all scripts and data for reproducing the analyses of the paper. The README files contain additional information.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record