Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
55
datasets available to search
ShareScore release 0.9.0
Dataset results
55 results for “ancestral reconstruction”
Bayesian Methods for Ancestral State Reconstruction in Morphosyntax
<p>Supplementary files to accompany journal submission.</p> <p>Files are:</p> <p> </p> <p>tree.pdf - pdf consensus tree, for illustration</p> <p>data.txt - coding file</p> <p>TREE_Set.t - nexus format sample of trees.</p> <p>sources.pdf - source materials used for languages</p>
Reconstruction of full-length LINE-1 progenitors from ancestral genomes (Supplementary Data)
<p><strong>Web Supplementary Files</strong></p> <ul> <li>Web Supplementary File 1 - FASTA files containing full-length reconstruction input sequences<strong>: full_length_reconstruction_input_sequence_fastas.zip</strong></li> <li>Web Supplementary File 2 - FASTA files containing Muscle alignments of the full-length reconstruction input sequences.<strong> full_length_reconstruction_input_sequence_alns.zip</strong></li> <li>Web Supplementary File 3 - FASTA file of full-length reconstructed sequences:<strong> full_length_reconstructions.fa</strong></li> <li>Web Supplementary File 4 - Table of full-length reconstruction statistics: <strong>full_length_reconstruction_stats.csv</strong></li> <li>Web Supplementary File 5 - FASTA files containing ORF reconstruction input sequences:<strong> orf_fastas.zip</strong></li> <li>Web Supplementary File 6 - FASTA files containing Macse alignments of the ORF reconstruction input sequences:<strong> ORF_reconstruction_input_sequence_alns.zip</strong></li> <li>Web Supplementary File 7 - Table of ORF reconstruction statistics: <strong>ORF_reconstructions.fa</strong></li> <li>Web Supplementary File 8 - Table of ORF reconstruction statistics: <strong>ORF_reconstruction_stats.csv</strong></li> <li>Web Supplementary File 9 - Table of Composite Sequences: <strong>bestfl_selection_fixed_CS_seqs.csv</strong></li> <li>Web Supplementary File 10 - Database of gold standards: <strong>L1_goldstandards.csv</strong></li> </ul> <p><strong>Data Underlying Figures</strong></p> <ul> <li>RepeatMasker scans of hg38 and ancestral genomes:<strong> </strong><strong>anc_gen_RM_out_files.zip</strong></li> <li><strong>Figure 4</strong> <ul> <li>4A <ul> <li>Source alignment of 54 composite sequences: <strong>220121_dropped12+L1ME3A_muscle.nt.afa</strong></li> <li>Tree produced using the alignment and FastTree: <strong>220121_dropped12+L1ME3A.tree</strong></li> </ul> </li> <li>4B <ul> <li>Source alignment of 67 Dfam L1 subfamily 3’ end models: <strong>200123_dfam_3ends.fa.muscle.aln</strong></li> <li>Tree produced using the alignment: <strong>200123_dfam_3ends.fa.muscle.aln.tree</strong></li> </ul> </li> </ul> </li> <li><strong>Figure 5</strong> <ul> <li>KZFP-TE enrichment p-values (from Barazandeh <em>et al</em> 2018):<strong> TE_KZFP_enrichment_pvals.xlsx</strong></li> <li>KZFP-TE top 500 peak overlap (from Barazandeh <em>et al</em> 2018): <strong>top500_peak_overlap.xlsx</strong></li> </ul> </li> <li><strong>Figure 6</strong> <ul> <li>RepeatMasker .out file for the Composite Sequence custom library queried against hg38: <strong>CS_RM_hg38.fa.out.gz</strong></li> </ul> </li> <li><strong>Figure S2</strong> <ul> <li>RepeatMasker scan .out file of hg38 (CG corrected Kimura Divergence values are in last column): <strong>hg38+KimDiv_RM.out</strong></li> <li>RepeatMasker scan .out file of the Progressive Cactus eutherian ancestral genome (CG corrected Kimura Divergence values are in last column): <strong>Progressive_Cactus_Euth+KimDiv_RM.out</strong></li> <li>RepeatMasker scan .out file of the Ancestors 1.1 eutherian ancestral genome (CG corrected Kimura Divergence values are in last column): <strong>Ancestors_Euth+KimDiv_RM.out</strong></li> </ul> </li> <li><strong>Figure S5</strong> <ul> <li>RepeatMasker scan .out files for Progressive Cactus simian and primate reconstructed ancestral genomes: <strong>progCactus_RM_outfiles.zip</strong></li> <li>S5A <ul> <li>FASTA files containing Cactus genome-derived reconstructed sequences equivalent to the L1MA2, L1MA4, and L1MD1-3 best full-length sequences: <strong>progCactus_reconstruction_bestFL_equivalents.zip</strong></li> </ul> </li> <li>S5B <ul> <li>FASTA files containing Muscle alignments of Cactus genome-derived full-length reconstruction input sequences: <strong>progCactus_reconstruction_input_sequence_alns.zip</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S6</strong> <ul> <li>S6A <ul> <li>Results of Conserved Domain scans of Cactus genome-derived full-length reconstructed sequences: <strong>CD_search_results_short_nms.txt</strong></li> </ul> </li> <li>S6B-D <ul> <li>Character posterior probabilities of “best” full-length reconstructed sequences: <strong>best_fl_post_probs.zip</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S7</strong> <ul> <li>S7B-C <ul> <li>Results of Conserved Domain scans of translated initial full-length reconstructed sequences: <strong>initial_recons_all_3frametrans_CD-search.txt</strong></li> <li>Results of Conserved Domain scans of translated reconstructed ORFs: <strong>recons_ORF1-2_all_3frametrans_CD-search.csv</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S15</strong> <ul> <li>S15A <ul> <li>Source alignment of 67 composite sequences: <strong>bestfl_selection_fixed_CS_seqs_muscle.nt.afa</strong></li> <li>Tree produced using the alignment: <strong>bestfl_selection_fixed_CS_seqs_muscle.nt.afa.tree</strong></li> </ul> </li> <li>S15B-E <ul> <li>Source Muscle alignments for phylogenetic trees of reconstructed sequence components: <ul> <li>ORF2: <strong>ORF2_keep54_muscle.nt.afa</strong></li> <li>5’ UTR: <strong>5utr_keep54_muscle.nt.afa</strong></li> <li>ORF1: <strong>ORF1_keep54_muscle.nt.afa</strong></li> <li>3’ UTR: <strong>3utr_keep54_muscle.nt.afa</strong></li> </ul> </li> <li>Trees produced using above alignments: <ul> <li>ORF2: <strong>ORF2_keep54_muscle.nt.afa.tree</strong></li> <li>5’ UTR: <strong>5utr_keep54_muscle.nt.afa.tree</strong></li> <li>ORF1: <strong>ORF1_keep54_muscle.nt.afa.tree</strong></li> <li>3’ UTR: <strong>3utr_keep54_muscle.nt.afa.tree</strong></li> </ul> </li> </ul> </li> </ul> </li> <li><strong>Figure S17</strong> <ul> <li>Unfiltered BLAST results of Composite Sequences queried against hg38: <strong>CS_hg38_blastn.csv.zip</strong></li> <li>BED file of L1 instances annotated using BLAST pipeline: <strong>BLAST_L1_hits.bed</strong></li> </ul> </li> </ul>
Identifying climatic drivers of hybridization with a new ancestral niche reconstruction method
<p>Applications of molecular phylogenetic approaches have uncovered evidence of hybridization across numerous clades of life, yet the environmental factors responsible for driving opportunities for hybridization remain obscure. Verbal models implicating geographic range shifts that brought species together during the Pleistocene have often been invoked, but quantitative tests using paleoclimatic data are needed to validate these models. Here, we produce a phylogeny for Heuchereae, a clade of 15 genera and 83 species in Saxifragaceae, with complete sampling of recognized species, using 277 nuclear loci and nearly complete chloroplast genomes. We then employ an improved framework with a coalescent simulation approach to test and confirm previous hybridization hypotheses and identify one new intergeneric hybridization event. Focusing on the North American distribution of Heuchereae, we introduce and implement a newly developed approach to reconstruct potential past distributions for ancestral lineages across all species in the clade and across a paleoclimatic record extending from the late Pliocene. Time calibration based on both nuclear and chloroplast trees recovers a mid- to late-Pleistocene date for most inferred hybridization events, a timeframe concomitant with repeated geographic range restriction into overlapping refugia. Our results indicate an important role for past episodes of climate change, and the contrasting responses of species with differing ecological strategies, in generating novel patterns of range contact among plant communities and therefore new opportunities for hybridization. The new ancestral niche method flexibly models the shape of niche while incorporating diverse sources of uncertainty and will be an important addition to the current comparative methods toolkit.</p>
Evolutionary insights into Felidae iris color through ancestral state reconstruction
<p>There have been almost no studies with an evolutionary perspective on eye (iris) color, outside of humans and domesticated animals. Extant members of the family Felidae have a great interspecific and intraspecific diversity of eye colors, in stark contrast to their closest relatives, all of which have only brown eyes. This makes the felids a great model to investigate the evolution of eye color in natural populations. Through machine learning cluster image analysis of publicly available photographs of all felid species, as well as a number of subspecies, five felid eye colors were identified: brown, hazel/green, yellow/beige, gray, and blue. Using phylogenetic comparative methods, the presence or absence of these colors was reconstructed on a phylogeny. Additionally, through a new color analysis method, the specific shades of the ancestors' eyes were quantitatively reconstructed. The ancestral felid population was predicted to have brown-eyed individuals, as well as a novel evolution of gray-eyed individuals, the latter being a key innovation that allowed the rapid diversification of eye color seen in modern felids, including numerous gains and losses of different eye colors. It was also found that the loss of brown eyes and the gain of yellow/beige eyes is associated with an increase in the likelihood of evolving round pupils, which in turn influence the shades present in the eyes. Along with these important insights, the unique methods presented in this work are widely applicable and will facilitate future research into phylogenetic reconstruction of color beyond irises.</p>
Identifying climatic drivers of hybridization with a new ancestral niche reconstruction method
Open the record for dataset details and reuse information.
Evolutionary insights into Felidae iris color through ancestral state reconstruction
Open the record for dataset details and reuse information.
Ancestral reconstruction of sunflower karyotypes reveals non-random chromosomal evolution
<p>Mapping the chromosomal rearrangements between species can inform our understanding of genome evolution, reproductive isolation, and speciation. Here we present a novel algorithm for identifying regions of synteny in pairs of genetic maps, which is implemented in the accompanying R package, syntR. The syntR algorithm performs as well as previous methods while being systematic and repeatable and can be used to map chromosomal rearrangements in any group of species. In addition, we present a systematic survey of chromosomal rearrangements in the annual sunflowers, which is a group known for extreme karyotypic diversity. We build high-density genetic maps for two subspecies of the prairie sunflower<i>,</i> <i>Helianthus</i> <i>petiolaris</i> ssp. <i>petiolaris</i> and <i>H. petiolaris</i> ssp. <i>fallax.</i> Using <i>syntR</i>, and we identify blocks of synteny between these two subspecies and previously published high-density genetic maps. We reconstruct ancestral karyotypes for annual sunflowers using those synteny blocks and conservatively estimate that there have been 7.9 chromosomal rearrangements per million years – a high rate of chromosomal evolution. Although the rate of inversion is even higher than the rate of translocation in this group, we further find that every extant karyotype is distinguished by between 1 and 3 translocations involving only 8 of the 17 chromosomes. This non-random exchange suggests that specific chromosomes are prone to translocation and may thus contribute disproportionately to widespread hybrid sterility in sunflowers. These data deepen our understanding of chromosome evolution and confirm that <i>Helianthus</i> has an exceptional rate of chromosomal rearrangement that may facilitate similarly rapid diversification.</p>
Dataset for the ancestral sequence reconstruction of NRC3
<p>Please refer to the material_and_method.pdf to check how we generated each file.</p><p> </p><p><strong>01_NRCH_cds_23-03-03.fasta</strong></p><p>FASTA file containing the nucleotide sequences of 2341 NRC helper sequences extracted from 124 Solanaceae NLRome dataset [1] and 20 NRC helper sequences from Adachi et al. 2023 [2].</p><p><br><strong>02_NRCH_cds_23-03-03.min2400max2800.fasta</strong></p><p>FASTA file containing the nucleotide sequences of 1753 NRC helper sequences filtered to keep only sequences between 2400 and 2800 bp.</p><p><br><strong>03_NRCH_cds_23-03-03.min2400max2800.uniq.fasta</strong></p><p>FASTA file containing the nucleotide sequences of 1116 NRC helper sequences filtered to keep only sequences between 2400 and 2800 bp and to remove duplicates.</p><p><br><strong>04_NRCH_cds_23-03-03.min2400max2800.uniq.NBARC_aa.fasta</strong></p><p>FASTA file containing the amino acid sequences of 1116 NB-ARC domains of NRC helpers after filtering to keep only sequences between 2400 and 2800 bp and to remove duplicates.</p><p><br><strong>05_NRCH_cds_23-03-03.min2400max2800.uniq.NBARC_aa.aln.fasta</strong></p><p>FASTA file containing the amino acid sequences of 1116 NB-ARC domains of NRC helpers after filtering to keep only sequences between 2400 and 2800 bp and to remove duplicates and after alignment with MAFFT.</p><p><br><strong>06_NRCH_cds_23-03-03.min2400max2800.uniq.NBARC_aa.aln.fasta.treefile</strong></p><p>Newick file containing the phylogenetic tree of the 1116 NB-ARC domains of NRC helpers reconstructed with FastTree.</p><p><br><strong>07_NRC123X_cds_23-03-03.min2400max2800.uniq.aa.fasta</strong></p><p>FASTA file containing the 324 full-length amino acid sequences of the NRC1/2/3/X clades.</p><p><br><strong>08_NRC123X_cds_23-03-03.min2400max2800.uniq.aa.aln.fasta</strong></p><p>FASTA file containing the 324 full-length amino acid sequences of the NRC1/2/3/X clades after alignment with MAFFT.</p><p><br><strong>09_NRC123X_cds_23-03-03.min2400max2800.uniq.nt.aln.fasta</strong></p><p>FASTA file containing the 324 full-length nucleotide sequences of the NRC1/2/3/X clades threaded onto the protein alignment with MAFFT.</p><p><br><strong>10_NRC123X_cds_23-03-03.min2400max2800.uniq.nt.aln.fasta.treefile</strong></p><p>Newick file containing the phylogenetic tree of the 324 full-length nucleotide sequences of the NRC1/2/3/X clades reconstructed with IQ-TREE.</p><p> </p><p><strong>11_FastML_NRC123X.zip</strong></p><p>Zip file containing the FastML results for the ancestral sequence reconstruction of the NRC1/2/3/X clades.</p><p> </p><p>1. Sugihara, Y., Toghani, A., Kamoun, S., & Kourelis, J. (2023). NLRome dataset from 124 genomes of plants in the Solanaceae family. <i>Zenodo</i>. https://doi.org/10.5281/zenodo.10354350</p><p>2. Adachi, H., Sakai, T., Harant, A., Pai, H., Honda, K., Toghani, A., Claeys, J., Duggan, C., Bozkurt, T. O., Wu, C., & Kamoun, S. (2023). An atypical NLR protein modulates the NRC immune receptor network in Nicotiana benthamiana. <i>PLOS Genetics</i>, 19(1), e1010500. https://doi.org/10.1371/journal.pgen.1010500</p><p> </p>
Dataset for "Ancient whale rhodopsin reconstructs dim-light vision over a major evolutionary transition: Implications for ancestral diving behaviour"
<p>Dataset files include:</p> <p>- Alignment of rhodopsin (Rh1) sequences formatted for PAML</p> <p>- Corresponding species tree in Newick format for PAML</p> <p>- Ancestral Rh1 amino acid sequences (for Cetacean and Whippomorpha nodes) estimated with PAML (random sites, clade, and amino acid models), Datamonkey, and ProtASR</p>
ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process
<p>Modern phylogenetic methods allow inference of ancestral molecular sequences given an alignment and phylogeny relating present day sequences. This provides insight into the evolutionary history of molecules, helping to understand gene function and to study biological processes such as adaptation and convergent evolution across a variety of applications. Here we propose a dynamic programming algorithm for fast joint likelihood-based reconstruction of ancestral sequences under the Poisson Indel Process (PIP). Unlike previous approaches, our method, named ARPIP, enables the reconstruction with insertions and deletions based on an explicit indel model. Consequently, inferred indel events have an explicit biological interpretation. Likelihood computation is achieved in linear time with respect to the number of sequences. Our method consists of two steps, namely finding the most probable indel points and reconstructing ancestral sequences. First, we find the most likely indel points and prune the phylogeny to reflect the insertion and deletion events per site. Second, we infer the ancestral states on the pruned subtree in a manner similar to FastML. We applied ARPIP on simulated datasets and on real data from the Betacoronavirus genus. ARPIP reconstructs both the indel events and substitutions with a high degree of accuracy. Our method fares well when compared to established state-of-the-art methods such as FastML and PAML. Moreover, the method can be extended to explore both optimal and suboptimal reconstructions, include rate heterogeneity through time and more. We believe it will expand the range of novel applications of ancestral sequence reconstruction.</p>
Ancestral Genomes: a resource for reconstructed ancestral genes and genomes across the tree of life
<p>For each ancestral gene, we assign a stable identifier, and provide additional information designed to facilitate analysis: an inferred name (based on its descendants in extant genomes), a reconstructed protein sequence, a set of inferred Gene Ontology (GO) annotations, and a “proxy gene” for each ancestral gene, defined as the least-diverged descendant of the ancestral gene in a given extant genome.</p>
FIGURE 5 Ancestral state reconstructions. A. Whorl count. B. Body length. C in Phylogeny and systematic revision of the helicarionid semislugs of eastern Queensland (Stylommatophora, Helicarionidae)
FIGURE 5 Ancestral state reconstructions. A. Whorl count. B. Body length. C. Altitude.
Data from: Ancestral reconstruction of sunflower karyotypes reveals non-random chromosomal evolution
Open the record for dataset details and reuse information.
Data and R Script from: The ancestor of sharks and rays laid eggs, but ancestral state reconstructions need empirically supported traits and transparent reporting
Open the record for dataset details and reuse information.
Data from: A fish-focused menu: An interdisciplinary reconstruction of Ancestral Tsleil-Waututh diets
Open the record for dataset details and reuse information.
ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process
Open the record for dataset details and reuse information.
Ancestral state reconstruction for regeneration and autotomy in arthopods and reptiles
<p>Some form of regeneration occurs in all lifeforms and extends from single-cell organisms to humans. The degree to which regenerative ability is distributed across different taxa, however, is harder to ascertain given the potential for phylogenetic constraint or inertia, and adaptive processes to shape this pattern. Here, we examine the phylogenetic history of regeneration in two groups where the trait has been well-studied: arthropods and reptiles. Because autotomy is often present alongside regeneration in these groups, we performed ancestral state reconstructions for both traits to more precisely assess the timing of their origins and the degree to which these traits coevolve. Using an ancestral trait reconstruction, we find that autotomy and regeneration were present at the base of the arthropod and reptile trees. We also find that when autotomy is lost it does not re-evolve easily. Lastly, we find that the distribution of regeneration is intimately connected to autotomy with the association being stronger in reptiles than in arthropods. While these patterns suggest that decoupling autotomy and regeneration at a broad phylogenetic scale may be difficult, the available data provides useful insight into their entanglement. Ultimately, our reconstructions provide important groundwork to explore how selection may have played a role during the loss of regeneration in specific lineages.</p>
Data from: Rate heterogeneity across Squamata, misleading ancestral state reconstruction and the importance of proper null model specification
The binary-state speciation and extinction (BiSSE) model has been used in many instances to identify state-dependent diversification and reconstruct ancestral states. However, recent studies have shown that the standard procedure of comparing the fit of the BiSSE model to constant-rate birth–death models often inappropriately favours the BiSSE model when diversification rates vary in a state-independent fashion. The newly developed HiSSE model enables researchers to identify state-dependent diversification rates while accounting for state-independent diversification at the same time. The HiSSE model also allows researchers to test state-dependent models against appropriate state-independent null models that have the same number of parameters as the state-dependent models being tested. We reanalyse two data sets that originally used BiSSE to reconstruct ancestral states within squamate reptiles and reached surprising conclusions regarding the evolution of toepads within Gekkota and viviparity across Squamata. We used this new method to demonstrate that there are many shifts in diversification rates across squamates. We then fit various HiSSE submodels and null models to the state and phylogenetic data and reconstructed states under these models. We found that there is no single, consistent signal for state-dependent diversification associated with toepads in gekkotans or viviparity across all squamates. Our reconstructions show limited support for the recently proposed hypotheses that toepads evolved multiple times independently in Gekkota and that transitions from viviparity to oviparity are common in Squamata. Our results highlight the importance of considering an adequate pool of models and null models when estimating diversification rate parameters and reconstructing ancestral states.
FIGURE 1. Ancestral character-state reconstructions for Characters 1–9 in Concentrated evolutionary novelties in the foot musculature of Odontophrynidae (Anura: Neobatrachia), with comments on adaptations for burrowing
FIGURE 1. Ancestral character-state reconstructions for Characters 1–9. Ambiguities in Macrogenioglottus alipioi and Odontophrynus carvalhoi in Characters 5–7 are due to polymorphism.
Figure 2 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 2. Bayesian consensus tree resulting from the LW-Rh and Wg gene alignments (871 bp). Coloured dots on the branches indicate the values of posterior probability (PP): green dots represent values between 1.00 and 0.95, yellow dots between 0.94 and 0.90, and red dots ≤ 0.89. The nodes are indicated with numbers. Values above and below the branches represent the ancestral genome size (GS; 1C-values, in picograms) at particular nodes: in blue is the value generated by the maximum likelihood (ML) [asterisks are related to confidence interval (CI) values shown in Supporting Information, Table S4]; orange is the value generated by maximum parsimony (MP); and black, given below the branches, is the value generated by Bayesian inference (BI). Genome size data (1C-values) were obtained in the present work (pink dots) or taken from the literature (grey dots).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.