Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.9.0
Dataset results
10 results for “Ancestral Genome Reconstruction”
Reconstruction of full-length LINE-1 progenitors from ancestral genomes (Supplementary Data)
<p><strong>Web Supplementary Files</strong></p> <ul> <li>Web Supplementary File 1 - FASTA files containing full-length reconstruction input sequences<strong>: full_length_reconstruction_input_sequence_fastas.zip</strong></li> <li>Web Supplementary File 2 - FASTA files containing Muscle alignments of the full-length reconstruction input sequences.<strong> full_length_reconstruction_input_sequence_alns.zip</strong></li> <li>Web Supplementary File 3 - FASTA file of full-length reconstructed sequences:<strong> full_length_reconstructions.fa</strong></li> <li>Web Supplementary File 4 - Table of full-length reconstruction statistics: <strong>full_length_reconstruction_stats.csv</strong></li> <li>Web Supplementary File 5 - FASTA files containing ORF reconstruction input sequences:<strong> orf_fastas.zip</strong></li> <li>Web Supplementary File 6 - FASTA files containing Macse alignments of the ORF reconstruction input sequences:<strong> ORF_reconstruction_input_sequence_alns.zip</strong></li> <li>Web Supplementary File 7 - Table of ORF reconstruction statistics: <strong>ORF_reconstructions.fa</strong></li> <li>Web Supplementary File 8 - Table of ORF reconstruction statistics: <strong>ORF_reconstruction_stats.csv</strong></li> <li>Web Supplementary File 9 - Table of Composite Sequences: <strong>bestfl_selection_fixed_CS_seqs.csv</strong></li> <li>Web Supplementary File 10 - Database of gold standards: <strong>L1_goldstandards.csv</strong></li> </ul> <p><strong>Data Underlying Figures</strong></p> <ul> <li>RepeatMasker scans of hg38 and ancestral genomes:<strong> </strong><strong>anc_gen_RM_out_files.zip</strong></li> <li><strong>Figure 4</strong> <ul> <li>4A <ul> <li>Source alignment of 54 composite sequences: <strong>220121_dropped12+L1ME3A_muscle.nt.afa</strong></li> <li>Tree produced using the alignment and FastTree: <strong>220121_dropped12+L1ME3A.tree</strong></li> </ul> </li> <li>4B <ul> <li>Source alignment of 67 Dfam L1 subfamily 3’ end models: <strong>200123_dfam_3ends.fa.muscle.aln</strong></li> <li>Tree produced using the alignment: <strong>200123_dfam_3ends.fa.muscle.aln.tree</strong></li> </ul> </li> </ul> </li> <li><strong>Figure 5</strong> <ul> <li>KZFP-TE enrichment p-values (from Barazandeh <em>et al</em> 2018):<strong> TE_KZFP_enrichment_pvals.xlsx</strong></li> <li>KZFP-TE top 500 peak overlap (from Barazandeh <em>et al</em> 2018): <strong>top500_peak_overlap.xlsx</strong></li> </ul> </li> <li><strong>Figure 6</strong> <ul> <li>RepeatMasker .out file for the Composite Sequence custom library queried against hg38: <strong>CS_RM_hg38.fa.out.gz</strong></li> </ul> </li> <li><strong>Figure S2</strong> <ul> <li>RepeatMasker scan .out file of hg38 (CG corrected Kimura Divergence values are in last column): <strong>hg38+KimDiv_RM.out</strong></li> <li>RepeatMasker scan .out file of the Progressive Cactus eutherian ancestral genome (CG corrected Kimura Divergence values are in last column): <strong>Progressive_Cactus_Euth+KimDiv_RM.out</strong></li> <li>RepeatMasker scan .out file of the Ancestors 1.1 eutherian ancestral genome (CG corrected Kimura Divergence values are in last column): <strong>Ancestors_Euth+KimDiv_RM.out</strong></li> </ul> </li> <li><strong>Figure S5</strong> <ul> <li>RepeatMasker scan .out files for Progressive Cactus simian and primate reconstructed ancestral genomes: <strong>progCactus_RM_outfiles.zip</strong></li> <li>S5A <ul> <li>FASTA files containing Cactus genome-derived reconstructed sequences equivalent to the L1MA2, L1MA4, and L1MD1-3 best full-length sequences: <strong>progCactus_reconstruction_bestFL_equivalents.zip</strong></li> </ul> </li> <li>S5B <ul> <li>FASTA files containing Muscle alignments of Cactus genome-derived full-length reconstruction input sequences: <strong>progCactus_reconstruction_input_sequence_alns.zip</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S6</strong> <ul> <li>S6A <ul> <li>Results of Conserved Domain scans of Cactus genome-derived full-length reconstructed sequences: <strong>CD_search_results_short_nms.txt</strong></li> </ul> </li> <li>S6B-D <ul> <li>Character posterior probabilities of “best” full-length reconstructed sequences: <strong>best_fl_post_probs.zip</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S7</strong> <ul> <li>S7B-C <ul> <li>Results of Conserved Domain scans of translated initial full-length reconstructed sequences: <strong>initial_recons_all_3frametrans_CD-search.txt</strong></li> <li>Results of Conserved Domain scans of translated reconstructed ORFs: <strong>recons_ORF1-2_all_3frametrans_CD-search.csv</strong></li> </ul> </li> </ul> </li> <li><strong>Figure S15</strong> <ul> <li>S15A <ul> <li>Source alignment of 67 composite sequences: <strong>bestfl_selection_fixed_CS_seqs_muscle.nt.afa</strong></li> <li>Tree produced using the alignment: <strong>bestfl_selection_fixed_CS_seqs_muscle.nt.afa.tree</strong></li> </ul> </li> <li>S15B-E <ul> <li>Source Muscle alignments for phylogenetic trees of reconstructed sequence components: <ul> <li>ORF2: <strong>ORF2_keep54_muscle.nt.afa</strong></li> <li>5’ UTR: <strong>5utr_keep54_muscle.nt.afa</strong></li> <li>ORF1: <strong>ORF1_keep54_muscle.nt.afa</strong></li> <li>3’ UTR: <strong>3utr_keep54_muscle.nt.afa</strong></li> </ul> </li> <li>Trees produced using above alignments: <ul> <li>ORF2: <strong>ORF2_keep54_muscle.nt.afa.tree</strong></li> <li>5’ UTR: <strong>5utr_keep54_muscle.nt.afa.tree</strong></li> <li>ORF1: <strong>ORF1_keep54_muscle.nt.afa.tree</strong></li> <li>3’ UTR: <strong>3utr_keep54_muscle.nt.afa.tree</strong></li> </ul> </li> </ul> </li> </ul> </li> <li><strong>Figure S17</strong> <ul> <li>Unfiltered BLAST results of Composite Sequences queried against hg38: <strong>CS_hg38_blastn.csv.zip</strong></li> <li>BED file of L1 instances annotated using BLAST pipeline: <strong>BLAST_L1_hits.bed</strong></li> </ul> </li> </ul>
Ancestral Genomes: a resource for reconstructed ancestral genes and genomes across the tree of life
<p>For each ancestral gene, we assign a stable identifier, and provide additional information designed to facilitate analysis: an inferred name (based on its descendants in extant genomes), a reconstructed protein sequence, a set of inferred Gene Ontology (GO) annotations, and a “proxy gene” for each ancestral gene, defined as the least-diverged descendant of the ancestral gene in a given extant genome.</p>
Figure 2 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 2. Bayesian consensus tree resulting from the LW-Rh and Wg gene alignments (871 bp). Coloured dots on the branches indicate the values of posterior probability (PP): green dots represent values between 1.00 and 0.95, yellow dots between 0.94 and 0.90, and red dots ≤ 0.89. The nodes are indicated with numbers. Values above and below the branches represent the ancestral genome size (GS; 1C-values, in picograms) at particular nodes: in blue is the value generated by the maximum likelihood (ML) [asterisks are related to confidence interval (CI) values shown in Supporting Information, Table S4]; orange is the value generated by maximum parsimony (MP); and black, given below the branches, is the value generated by Bayesian inference (BI). Genome size data (1C-values) were obtained in the present work (pink dots) or taken from the literature (grey dots).
Figure 1 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 1. Fluorescence intensity histograms obtained from three different species, with Drosophila melanogaster as internal standard, stained with propidium iodide (PI; A–C) or 4,6-diamidino-2-phenylindole (DAPI; D–F). The x-axis corresponds to the scale of fluorescence intensity, and the y-axis represents the number of nuclei with that fluorescence intensity.
Figure 3 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 3. Mean genome size (in picograms and megabase pairs) estimated for Formicidae subfamilies. The phylogenetic tree generated in the present study was redrawn, with collapsed branches corresponding to species of the same subfamily.
Fig. 1 in Reconstruction of the ancestral metazoan genome reveals an increase in genomic novelty
Fig. 1 Reconstruction of ancestral genomes. Evolutionary relationships of the major groups included in his study2. Different categories of HG are indicated in each node, from top to bottom, Ancestral HG, Novel HG, Novel Core HG, and Lost HG. Values assume sponges as the sister group to other animals, and placozoans as sister group to Planulozoa (=Cnidaria + Bilateria); alternative phylogenetic hypotheses are explored in Supplementary Data 3-8. Organism outlines from phylopic.org and the authors
Figure 3 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 3. Mean genome size (in picograms and megabase pairs) estimated for Formicidae subfamilies. The phylogenetic tree generated in the present study was redrawn, with collapsed branches corresponding to species of the same subfamily.
Figure 2 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 2. Bayesian consensus tree resulting from the LW-Rh and Wg gene alignments (871 bp). Coloured dots on the branches indicate the values of posterior probability (PP): green dots represent values between 1.00 and 0.95, yellow dots between 0.94 and 0.90, and red dots ≤ 0.89. The nodes are indicated with numbers. Values above and below the branches represent the ancestral genome size (GS; 1C-values, in picograms) at particular nodes: in blue is the value generated by the maximum likelihood (ML) [asterisks are related to confidence interval (CI) values shown in Supporting Information, Table S4]; orange is the value generated by maximum parsimony (MP); and black, given below the branches, is the value generated by Bayesian inference (BI). Genome size data (1C-values) were obtained in the present work (pink dots) or taken from the literature (grey dots).
Fig. 2 in Reconstruction of the ancestral metazoan genome reveals an increase in genomic novelty
Fig. 2 Novelty in ancestral genomes. a Proportion of Novel HG in the Ancestral HG for different holozoan ancestors. b Percentage of Core HG that are novel, and percentage of highly preserved genes among the Novel HG across different LCA. c Number of Protein Class GO hits for the fruit fly representatives of the Novel HG for the various phylogenetic nodes
Figure 1 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 1. Fluorescence intensity histograms obtained from three different species, with Drosophila melanogaster as internal standard, stained with propidium iodide (PI; A–C) or 4,6-diamidino-2-phenylindole (DAPI; D–F). The x-axis corresponds to the scale of fluorescence intensity, and the y-axis represents the number of nuclei with that fluorescence intensity.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.