Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
111
datasets available to search
ShareScore release 0.9.0
Dataset results
111 results for “genome duplications”
Reproductive complexity, whole genome duplication, and genome size data across vascular plants
Open the record for dataset details and reuse information.
Intercontinental dispersal and whole‐genome duplication contribute to loss of self‐incompatibility in a polyploid complex
Open the record for dataset details and reuse information.
The genomic architecture of the passerine MHC region: high repeat content and contrasting evolutionary histories of single copy and tandemly duplicated MHC genes
Open the record for dataset details and reuse information.
Nested whole-genome duplications coincide with diversification and high morphological disparity in Brassicaceae
Open the record for dataset details and reuse information.
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
Open the record for dataset details and reuse information.
Data from: Whole-genome duplication and host genotype affect rhizosphere microbial communities
Open the record for dataset details and reuse information.
Effects of glaciation and whole genome duplication on the distribution of the Campanula rotundifolia polyploid complex
Open the record for dataset details and reuse information.
Data from: Screening of duplicated loci reveals hidden divergence patterns in a complex salmonid genome
Open the record for dataset details and reuse information.
Data from: A meta-analysis of whole genome duplication and the effects on flowering traits in plants
Open the record for dataset details and reuse information.
Data for: Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure
Open the record for dataset details and reuse information.
Genome-structural analyses support an allotetraploid origin of the walnut family from within Myricaceae and shared genome duplications reveal substitution rate variation
Open the record for dataset details and reuse information.
Data from: Genetic drift dominates genome-wide regulatory evolution following an ancient whole genome duplication in Atlantic salmon
Open the record for dataset details and reuse information.
Nucleotide alignments of eight meiosis genes under extreme selection following whole genome duplication in Arabidopsis lyrata/A.arenosa.
Open the record for dataset details and reuse information.
Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"
<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data" [Zaccaria & Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at <a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em> which contains the entire dataset has been compressed with standard <em>zip</em>.</p>
Exploring whole-genome duplicate gene retention with complex genetic interaction analysis
<p>Whole-genome duplication<b> </b>has played a central role in genome evolution of many organisms, including the human genome. Most duplicated genes are eliminated and factors that influence the retention of persisting duplicates remain poorly understood. Here, we describe a systematic complex genetic interaction analysis with yeast paralogs derived from the whole-genome duplication event. Mapping digenic interactions for a deletion mutant of each paralog and trigenic interactions for the double mutant provides insight into their roles and a quantitative measure of their functional redundancy. Trigenic interaction analysis distinguishes two classes of paralogs, a more functionally divergent subset and another that retained more functional overlap. Gene feature analysis and modeling suggest that evolutionary trajectories of duplicated genes are dictated by combined functional and structural entanglement factors.</p>
Data from: Whole genome duplication and transposable element proliferation drive genome expansion in Corydoradinae catfishes
Genome size varies significantly across eukaryotic taxa and the largest changes are typically driven by macro-mutations such as whole genome duplications (WGDs) and proliferation of repetitive elements. These two processes may affect the evolutionary potential of lineages by increasing genetic variation and changing gene expression. Here we elucidate the evolutionary history and mechanisms underpinning genome size variation in a species rich group of Neotropical catfishes (Corydoradinae) with extreme variation in genome size - 0.6pg to 4.4 pg per haploid cell. Firstly, genome size was quantified in 65 species and mapped onto a novel fossil-calibrated phylogeny. Two evolutionary shifts in genome size were identified across the tree - the first between 43-49 Mya (95% highest posterior density (HPD) 36.2-68.1 Mya) and the second at ~19 Mya (95% HPD 15.3-30.14 Mya). Secondly, RAD sequencing was used to identify potential WGD events and quantify transposable element abundance in different lineages. Evidence of two lineage scale WGDs were identified across the phylogeny, the first event occurring between 54-66 Mya (95% HPD 42.56-99.5 Mya) and the second at 20-30 Mya (95% HPD 15.3-45 Mya) based on haplotype numbers per contig and between 35-44 Mya (95% HPD 30.29-64.51 Mya) and 20-30 Mya (95% HPD 15.3-45 Mya) based on SNP read ratios. Transposable element abundance increased considerably in parallel with genome size, with a single TE-family (TC1-IS630-Pogo) showing several increases across the Corydoradinae, with the most recent at 20-30 Mya (95% HPD 15.3-45 Mya) and an older event at 35-44 Mya (95% HPD 30.29-64.51 Mya). We identified signals congruent with two WGD duplication events, as well as an increase in TE abundance across different lineages, making the Corydoradinae an excellent model system to study the effects of WGD and TEs on genome and organismal evolution.
Data from: Comparative genomics of chemosensory protein genes reveals rapid evolution and positive selection in ant-specific duplicates
Gene duplications can have a major role in adaptation, and gene families underlying chemosensation are particularly interesting due to their essential role in chemical recognition of mates, predators and food resources. Social insects add yet another dimension to the study of chemosensory genomics, as the key components of their social life rely on chemical communication. Still, chemosensory gene families are little studied in social insects. Here we annotated chemosensory protein (CSP) genes from seven ant genomes and studied their evolution. The number of functional CSP genes ranges from 11 to 21 depending on species, and the estimated rates of gene birth and death indicate high turnover of genes. Ant CSP genes include seven conservative orthologous groups present in all the ants, and a group of genes that has expanded independently in different ant lineages. Interestingly, the expanded group of genes has a differing mode of evolution from the orthologous groups. The expanded group shows rapid evolution as indicated by a high dN/dS (nonsynonymous to synonymous changes) ratio, several sites under positive selection and many pseudogenes, whereas the genes in the seven orthologous groups evolve slowly under purifying selection and include only one pseudogene. These results show that adaptive changes have played a role in ant CSP evolution. The expanded group of ant-specific genes is phylogenetically close to a conservative orthologous group CSP7, which includes genes known to be involved in ant nestmate recognition, raising an interesting possibility that the expanded CSPs function in ant chemical communication.
Data from: A well-constrained estimate for the timing of the salmonid whole genome duplication reveals major decoupling from species diversification
Whole genome duplication (WGD) is often considered to be mechanistically associated with species diversification. Such ideas have been anecdotally attached to a WGD at the stem of the salmonid fish family, but remain untested. Here, we characterized an extensive set of gene paralogues retained from the salmonid WGD, in species covering the major lineages (subfamilies Salmoninae, Thymallinae and Coregoninae). By combining the data in calibrated relaxed molecular clock analyses, we provide the first well-constrained and direct estimate for the timing of the salmonid WGD. Our results suggest that the event occurred no later in time than 88 Ma and that 40–50 Myr passed subsequently until the subfamilies diverged. We also recovered a Thymallinae–Coregoninae sister relationship with maximal support. Comparative phylogenetic tests demonstrated that salmonid diversification patterns are closely allied in time with the continuous climatic cooling that followed the Eocene–Oligocene transition, with the highest diversification rates coinciding with recent ice ages. Further tests revealed considerably higher speciation rates in lineages that evolved anadromy—the physiological capacity to migrate between fresh and seawater—than in sister groups that retained the ancestral state of freshwater residency. Anadromy, which probably evolved in response to climatic cooling, is an established catalyst of genetic isolation, particularly during environmental perturbations (for example, glaciation cycles). We thus conclude that climate-linked ecophysiological factors, rather than WGD, were the primary drivers of salmonid diversification.
RAW DATA parthenogenetic lizard Darevskia armeniaca premeiotic genome duplication
<p>parthenogenetic lizard Darevskia armeniaca premeiotic genome duplication raw data images</p>
Evolution of binding preferences among whole-genome duplicated transcription factors
<p>Throughout evolution, new transcription factors (TFs) emerge by gene duplication, promoting growth and rewiring of transcriptional networks. How TF duplicates diverge is known for only a few studied cases. To provide a genome-scale view, we considered the 35% of budding yeast TFs, classified as whole-genome duplication (WGD)-retained paralogs. Using high-resolution profiling, we find that ~60% of paralogs evolved differential binding preferences. We show that this divergence results primarily from variations outside the DNA binding domains (DBDs), while DBD preferences remain largely conserved. Analysis of non-WGD orthologs revealed that ancestral preferences are unevenly split between duplicates, while new targets are acquired preferentially by the least conserved paralog (biased sub/neo-functionalization). Dimer-forming paralogs evolved mostly one-sided dependency, while other paralogs interacted through low-magnitude DNA-binding competition that minimized paralog interference. We discuss the implications of our findings for the evolutionary design of transcriptional networks.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.