Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

88

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

88 results for “phylogenetic datasets”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area"

<p>This repository archives the dataset of the article &quot;Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area&quot;. The cognate annotation of Tshangla, Kho-Bwa, Hrusish, Mishmic, and Tani languages were done by us. The cognate decision on the other languages was annotated by Sagart et al. (2019).&nbsp;&nbsp;Please use the following information to cite our work:&nbsp;<br> Wu, M.-S, Bodt, T. A, Tresoldi, T. (2022). &nbsp;Bayesian phylogenetics illuminate shallower relationships Trans-Himalayan languages in the Tibet-Arunachal area. Linguistics of the Tibeto-Burman Area. [forthcoming]</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Dataset for: Evaluating phylogenetic methods for quantifying risks and opportunities presented by forks in open source software (master dissertation).

<p>This is the data for my master dissertation [1]. If you wish to get a copy, download it from Zenodo and open docs/master.pdf.</p> <p>Data acquisition and encoding techniques are described in paragraph 3.1.1 (table 3.1).</p> <p>The data is described in more detail in paragraph 4.1 (table 4.2).</p> <p>* fork1_all.csv: MySQL server / MariaDB server<br> * fork2_all.csv: Linux kernel / Android kernel<br> * fork3_all.csv: Apache OpenOffice / LibreOffice</p> <p>==Cite==<br> [1] A. Ortiz-Troncoso. Evaluating phylogenetic methods for quantifying risks and opportunities presented<br> by forks in open source software (master dissertation). Zenodo, 2018. doi: http://doi.org/10.5281/zenodo.1158292</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024

<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (&ldquo;full&rdquo; matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (&ldquo;reduced&rdquo; matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the &ldquo;new technology&rdquo; search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

CLDF dataset derived from Birchall et al.'s "A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family" from 2016

<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall J, Dunn M, &amp; Greenhill SJ. 2016. A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family. International Journal of American Linguistics 82(3). 255–284.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF Dataset derived from the Bahnaric data in Sidwell's "Austroasiatic dataset for phylogenetic analysis" from 2015

<p>Cite the source of the dataset as:</p> <blockquote> <p>Sidwell, Paul. 2015. Austroasiatic dataset for phylogenetic analysis: 2015 version. Mon-Khmer Studies (Notes, Reviews, Data-Papers) 44. lxviii-ccclvii.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Lee and Hasegawa's "Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages" from 2011

<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee, Sean and Hasegawa, Toshikazu (2011). Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proceedings of the Royal Society B: Biological Sciences, 278(1725), 3662–3669. doi:10.1098/rspb.2011.0518.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Gerardi and Reichert's "The Tupí-Guaraní Language Family: A Phylogenetic Classification" from 2021

<p>Cite the source of the dataset as:</p> <blockquote> <p>Ferraz Gerardi, Fabrício and Reichert, Stanislav (2021) The Tupí-Guaraní Language Family: A Phylogenetic Classification. Diachronica 38(2). 151--188. DOI: https://doi.org/10.1075/dia.18032.fer.</p> </blockquote>

opencc-by-4.0Sep 2024View details →
zenodo44/100

CLDF dataset derived from Satterthwaite-Phillips' "Phylogenetic Inference of the Tibeto-Burman Languages" from 2011

<p>Cite the source of the dataset as:</p> <blockquote> <p>Satterthwaite-Phillips, Damian (2011) Phylogenetic inference of the Tibeto-Burman languages or on the usefuseful of lexicostatistics (and &quot;megalo&quot;-comparison) for the subgrouping of Tibeto-Burman. Stanford: Stanford University.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Dataset for "Phylogenetic structure of European forest vegetation" - Journal of Biogeography (DOI: 10.1111/jbi.14046)

<p>This dataset contains the list of plant occurrences and geographical and environmental attributes of the vegetation-plots analyzed in the paper titled &ldquo;Phylogenetic structure of European forest vegetation&rdquo; by Padull&eacute;s Cubino et al. (2021; Journal of Biogeography; DOI: 10.1111/jbi.14046).&nbsp;</p> <p>The dataset contains 3 tables:</p> <ol> <li>&ldquo;Table_taxa.csv&rdquo;: It includes the list of angiosperm plant taxa in selected vegetation plots.</li> <li>&ldquo;Table_sites.csv&rdquo;: It includes data on the environmental variables of plots, their classification into different forest types, their location in 1<sup>o</sup>&nbsp;&times; 1<sup>o</sup>&nbsp;grid cells, and the reference to the original datasets archived in the European Vegetation Archive (EVA; http://euroveg.org/eva-database-participating-databases).</li> <li>&ldquo;Metadata.csv&rdquo;: It includes a description of the fields found in the two previous tables.</li> </ol>

opencc-by-4.0Jan 2021View details →
zenodo40/100

CLDF dataset derived from Kitchen et al.'s "Bayesian phylogenetic analysis of Semitic languages" from 2009

<p>Cite the source of the dataset as:</p> <blockquote> <p>Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Andrew Kitchen, Christopher Ehret, Shiferaw Assefa, Connie J. Mulligan. Proc. R. Soc. B 2009 -; DOI: 10.1098/rspb.2009.0408. Published 29 April 2009</p> </blockquote>

opencc-by-nc-4.0Jul 2023View details →
dryad40/100

Supplementary datasets, data analysis code, and R tutorials for: Phylogenetic analysis of adaptation in comparative physiology and biomechanics: overview and a case study of thermal physiology in treefrogs

<p>Comparative phylogenetic studies of adaptation are uncommon in biomechanics and physiology. Such studies require collecting data from many species, a challenge when data collection is experimentally intensive. Moreover, researchers struggle to employ the most biologically appropriate phylogenetic tools for identifying adaptive evolution. Here, we detail an established but greatly underutilized phylogenetic comparative framework—the Ornstein-Uhlenbeck process—that explicitly models long-term adaptation. We discuss challenges in implementing and interpreting the model, and we outline potential solutions. We demonstrate use of the model through studying the evolution of thermal physiology in treefrogs. Frogs of the family Hylidae have twice colonized the temperate zone from the tropics, and such colonization likely involved a fundamental change in physiology due to colder and more seasonal temperatures. However, which traits changed to allow colonization is unclear. We measured cold-temperature tolerance and characterized thermal performance curves in jumping for twelve species of treefrogs distributed from the Neotropics to temperate North America. We then conducted phylogenetic comparative analyses to examine how tolerances and performance curves evolved and to test whether that evolution was adaptive. We found that tolerance to low temperatures increased with the transition to the temperate zone. In contrast, jumping well at colder temperatures was unrelated to biogeography and thus did not adapt during dispersal. Overall, our paper shows how comparative phylogenetic methods can be leveraged in biomechanics and physiology to test the evolutionary drivers of variation among species.</p>

opencc-zeroOct 2021View details →
zenodo40/100

Fig. 3. Phylogenetic trees obtained from a concatenated dataset with a in Molecular Systematics and Morphological Analyses of the Subgenus Setihenricia (Echinodermata: Asteroidea: Henricia) from Japan

Fig. 3. Phylogenetic trees obtained from a concatenated dataset with a total length of 1,277 bp, consisting of seven mitochondrial genes (16S, tRNA-Ala, tRNA-Leu, tRNA-Asn, tRNA-Gln, tRNA-Pro, and COI). The trees were built based on maximum likelihood (ML, left) and Bayesian inference (BI, right). Values at nodes indicate bootstrap scores from ML and posterior probabilities from BI. Outgroups are only shown in the ML tree with both the support values. Scale bars indicate the number of nucleotide substitutions per site. OTUs sequenced in this study are in bold face. Each letter in parentheses after non-bold OTUs denotes the source: C, Chichvarkhin (2017b); F, Foltz and Rocha- Olivares (unpublished); K, Knott et al. (2018); L, Lopes et al. (2016); M, Matsubara et al. (2004); W, Wada et al. (1996). Circles indicate species listed as Setihenricia in Chichvarkhin and Chichvarkhina (2017). Triangles indicate species morphologically identified as Setihenricia in this study (see Fig. 4A).

opencc-by-4.0Jul 2019View details →
zenodo40/100

Dataset of Rainy years counteract negative effects of drought on taxonomic, functional, and phylogenetic diversity: resilience in annual plant communities

<p>Data used in the article:&nbsp;</p> <p><strong>Rainy years counteract negative effects of drought on taxonomic, functional, and phylogenetic diversity: resilience in annual plant communities</strong></p> <p><strong>Abstract</strong></p> <p>1- Climate models forecast changes in the amounts and distribution of rain, which may affect ecosystems worldwide, especially in drylands where water is already the limiting factor for plant life. Annual plant communities are common in drylands where they can complete their entire life cycle during the rainy period while avoiding the dry season. Moreover, seed dormancy allows them to disperse over time by remaining in the seed bank for long periods. However, the extent to which these communities will be able to tolerate increasing drought is uncertain.</p> <p>2- We performed a five-year rainfall reduction treatment under field conditions and determined its effects on annual plant communities in a Mediterranean gypsum ecosystem. We assessed the taxonomic, functional, and phylogenetic diversity of these communities each year for five years.</p> <p>3-The taxonomic and functional diversity decreased under the rainfall reduction treatment whereas the phylogenetic diversity increased. Moreover, the relative importance of species with drought-resistant functional designs increased in the community assemblages. However, after a rainy season with above average rainfall, all of the diversity values recovered completely even under the rainfall reduction treatment.</p> <p>4- Our results provide important insights into the responses of these plant communities under a climate change scenario, where they indicate high losses of diversity during drought events but rapid recovery in milder years.</p> <p><em>Synthesis</em> Our findings highlight the great resilience of annual plant communities in drylands, which may allow them to tolerate increased drought under the present climate change scenario.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

FIGURE 37. Phylogenetic results from parsimony analysis using the cranial dataset. A in A reappraisal of the cranial and mandibular osteology of the spinosaurid Irritator challengeri (Dinosauria: Theropoda)

FIGURE 37. Phylogenetic results from parsimony analysis using the cranial dataset. A, strict consensus tree of 153 MPTs retained from an equal weighting analysis (see Methods for details); B, reduced consensus tree, pruning wild card taxa from the strict consensus. Wild card taxa are highlighted with coloured boxes in A, and their possible topological positions are shown with same coloured squares in B. Important clades are labelled.

opencc-by-4.0Dec 2023View details →
zenodo40/100

FIGURE 36. Phylogenetic results from parsimony analysis using the full dataset. A in A reappraisal of the cranial and mandibular osteology of the spinosaurid Irritator challengeri (Dinosauria: Theropoda)

FIGURE 36. Phylogenetic results from parsimony analysis using the full dataset. A, strict consensus tree of 8184 MPTs retained from an equal weighting analysis (see methods for details); B, partial reduced consensus tree, showing the clade Spinosauridae after removal of the taxon Vallibonavenatrix; C, strict consensus tree of 406 MPTs retained from an implied weighting analysis using a concavity constant of k=10 (see Methods for details). Important clades are labelled. Irritator as the main focus of our study is highlighted in bold face within the clade Spinosauridae.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Eight empirical phylogenetic datasets analysed in Rota et al.

<p>Eight empirical phylogenetic datasetes analysed in Rota et al. study &quot;A simple method for data partitioning based on relative evolutionary rates&quot;. All eight datasets are given in the phylip format with the alignment followed by partitions defined to correspond to the best partitioning strategy as described in the study. These are the datasets: Arctina, Calisto, Coenonymphina, Choreutidae, Geometridae, Morpho, Noctuidae, and Pieridae. The datasets have been modified from their published versions to reduce the amount of missing data. The details are in the paper.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Dataset: Phylogenetic matrix: A new myromecophilic spider from the Chihuahuan desert

<p>Dataset: Phylogenetic matrix: A new myromecophilic spider from the Chihuahuan desert</p> <p>combined.nex.txt: MrBayes Nexus data matrix&nbsp;of DNA and morphology combined.</p> <p>morphology.tnt.txt: TNT data matrix of the morphology partition.</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Brut dataset for Drivers of taxonomic, functional and phylogenetic diversities in dominant ground-dwelling arthropods of coastal heathlands

<p>Brut dataset used in&nbsp;Drivers of taxonomic, functional and phylogenetic diversities&nbsp;in dominant ground-dwelling arthropods of coastal heathlands.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

FIGURE 2 in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics

FIGURE 2. Majority-rule cladogram of nine most parsimonious trees (length: 2038, CI: 0.2498, RI: 0.5337, RC: 0.1333, HI: 0.7502) from analysis of our new dataset of 40 taxa and 368 osteological characters. All trees show optimization of an Aptornis defossor+Psophia obscura sister group. Synapomorphies are detailed in table 2. Extinct taxa are denoted with daggers. Majority-rule percentages are annotated above branches, followed by bootstrap support values greater than 50% in parentheses. Branch length ranges are below branches.

opencc-by-4.0May 2019View details →
zenodo40/100

FIGURE 5 in A new morphological dataset reveals a novel relationship for the adzebills of New Zealand (Aptornis) and provides a foundation for total evidence neoavian phylogenetics

FIGURE 5. Synapomorphies for the pelvis of Aptornis defossor (AMNH 7300, A) and Psophia obscura (AMNH 2671, B). The pelvises are shown in ventral view. Scale bars are different for each specimen and are shown below each specimen. Labels correspond to synapomorphies, with character numbers followed by character states in parentheses. Abbreviations: cio, crista iliaca obliqua; ili, preacetabular ilium; ish, postacetabular ischium; pil, postacetabular ilium; syn, synsacrum

opencc-by-4.0May 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record