Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

111

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

111 results for “genome duplications”

Learn how ShareScore rates datasets ↗
zenodo40/100

Phylogenomics reveals an extensive history of genome duplication in diatoms (Bacillariophyta)

<p>Abstract</p> <p>Premise of the Study</p> <p>Diatoms are one of the most species‐rich lineages of microbial eukaryotes. Similarities in clade age, species richness, and primary productivity motivate comparisons to angiosperms, whose genomes have been inordinately shaped by whole‐genome duplication (WGD). WGDs have been linked to speciation, increased rates of lineage diversification, and identified as a principal driver of angiosperm evolution. We synthesized a large but scattered body of evidence that suggests polyploidy may be common in diatoms as well.</p> <p>Methods</p> <p>We used gene counts, gene trees, and distributions of synonymous divergence to carry out a phylogenomic analysis of WGD across a diverse set of 37 diatom species.</p> <p>Key Results</p> <p>Several methods identified WGDs of varying age across diatoms. Determining the occurrence, exact number, and placement of events was greatly impacted by uncertainty in gene trees. WGDs inferred from synonymous divergence of paralogs varied depending on how redundancy in transcriptomes was assessed, gene families were assembled, and synonymous distances (Ks) were calculated. Our results highlighted a need for systematic evaluation of key methodological aspects of Ks‐based approaches to WGD inference. Gene tree reconciliations supported allopolyploidy as the predominant mode of polyploid formation, with strong evidence for ancient allopolyploid events in the thalassiosiroid and pennate diatom clades.</p> <p>Conclusions</p> <p>Our results suggest that WGD has played a major role in the evolution of diatom genomes. We outline challenges in reconstructing paleopolyploid events in diatoms that, together with these results, offer a framework for understanding the impact of genome duplication in a group that likely harbors substantial genomic diversity.</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Autopolyploidy Genome Duplication Preserves Other Ancient Genome Duplications in Atlantic Salmon (Salmo salar) Supplementary Datasets

<p>For various species, alignments were found between a protein database (produced from Zebrafish) and the sequenced genome of that species.  Using Perl scripts and the alignments, gene models were identified in the various species based on the protein sequences. </p> <ul> <li>The gene models, for the various species, can be found in the .gff3 files.  Some of the .gff3 files have had ribosomal proteins removed. </li> <li>Homeologous regions were then identified using Perl scripts and can be found in .gff3 files as well.  They have Homeologous_Regions.gff3 in their title. </li> <li>Homeologous genes in these regions were counted (named XX_XX_Homeologous_Regions.txt), and compared to all of the genes (not just homeologous genes) in these regions (named Gene_Count_Homeolgous_XX_XX_XX.txt) to find the density. </li> <li>Homeologous gene sequences were compared to each other to identify the Ps values between them using a program called SNAP (Files with _Homeologous_region_analysis_version_1.2.txt at the end). </li> <li>The analyses of these files are summarized in "Pn_Ps_Values_Vertebrate_Homeologous_Regions.ods." </li> <li>The synteny between species can be found in the files with .seg extensions (These can be opened in IGV). </li> <li>A comparison between the gene density and Ps value for each homeologous region can be found in the file, "Gene_Density_Compared_to_Ps_Values.ods."</li> </ul> <p>Included is an extended readme file and Perl scripts (.pl extension) in a compressed file (Final_Scripts.tar.gz).</p>

opencc-by-4.0Feb 2017View details →
dryad40/100

Data From: Polygenic basis and the role of genome duplication in adaptation to similar selective environments

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad40/100

Data for: The 3-dimensional genome drives the evolution of asymmetric gene duplicates via enhancer capture-divergence

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways

Background: Teak, a member of the Lamiaceae family, produces one of the most expensive hardwoods in the world. High demand coupled with deforestation have caused a decrease in natural teak forests, and future supplies will be reliant on teak plantations. Hence, selection of teak tree varieties for clonal propagation with superior growth performance is of great importance, and access to high-quality genetic and genomic resources can accelerate the selection process by identifying genes underlying desired traits. Findings: To facilitate teak research and variety improvement, we generated a highly contiguous, chromosomal-scale genome assembly using high-coverage PacBio long reads coupled with high-throughput chromatin conformation capture. Of the 18 teak chromosomes, we generated 17 near-complete pseudomolecules with one chromosome present as two chromosome arm scaffolds. Genome annotation yielded 31,168 genes encoding 46,826 gene models, of which, 39,930 and 41,155 had Pfam domain and expression evidence, respectively. We identified 14 clusters of tandem-duplicated terpene synthases (TPSs), genes central to the biosynthesis of terpenes which are involved in plant defense and pollinator attraction. Transcriptome analysis revealed 10 TPSs highly expressed in woody tissues, of which, 8 were in tandem, revealing the importance of resolving tandemly duplicated genes and the quality of the assembly and annotation. We also validated the enzymatic activity of four TPSs to demonstrate the function of key TPSs. Conclusions: In summary, this high-quality chromosomal-scale assembly and functional annotation of the teak genome will facilitate the discovery of candidate genes related to traits critical for sustainable production of teak and for anti-insecticidal natural products.

opencc-zeroDec 2018View details →
dryad36/100

Data from: Genome duplication effects on functional traits and fitness are genetic context and species dependent: studies of synthetic polyploid Fragaria

PREMISE OF THE STUDY Divergence in functional traits and adaptive responses to environmental change underlies the ecological advantage of polyploid plants in the wild. While established polyploids may benefit from combined outcomes of genome doubling, hybridization and polyploidy-enabled adaptive evolution, it remains less clear whether genome doubling alone can drive ecological divergence or whether the outcome is genetically variable.METHODS Using synthetic, colchicine-induced, autotetraploid (4x) plants derived from self-pollinated diploid (2x) seeds, and their colchicine-treated but unconverted diploid (2x.nc) full sibs from two diploid wild strawberry taxa (Fragaria vesca ssp. vesca and F. vesca ssp. bracteata), we examined the effects of genome doubling on functional traits, heat stress tolerance and fitness components across taxa and maternal families (i.e. genetic families) within taxa.KEY RESULTS Comparisons between 2x and 2x.nc plants indicated a negligible effect of colchicine treatment on functional traits. Genome doubling increased stomatal length, and decreased stomatal density, specific leaf area and leaf vein density, recapitulating patterns observed in wild polyploid Fragaria. Trichome density, heat stress tolerance and relative growth rate were not significantly affected by genome doubling. Although a reduction in clonal reproduction was observed in response to genome doubling, this effect was strongly genetic family dependent.CONCLUSIONS The results suggest that genome doubling during incipient speciation alone can generate ecological divergence and variation among genetic lineages. This response potentially allows for rapid short-term evolutionary adaptation and fuels genomic diversity and independent origins of polyploidy.

opencc-zeroSep 2020View details →
dryad36/100

The Rhododendron genome and chromosomal organization provide insight into shared whole-genome duplications across the heath family (Ericaceae)

<p>The genus <em>Rhododendron</em> (Ericaceae), which includes horticulturally important plants such as azaleas, is a highly diverse and widely distributed genus of &gt;1,000 species. Here, we report the chromosome-scale de novo assembly and genome annotation of <em>Rhododendron williamsianum</em> as a basis for continued study of this large genus. We created multiple short fragment genomic libraries, which were assembled using ALLPATHS-LG. This was followed by contiguity preserving transposase sequencing (CPT-seq) and fragScaff scaffolding of a large fragment library, which improved the assembly by decreasing the number of scaffolds and increasing scaffold length. Chromosome-scale scaffolding was performed by proximity-guided assembly (LACHESIS) using chromatin conformation capture (Hi-C) data. Chromosome-scale scaffolding was further refined and linkage groups defined by restriction-site associated DNA (RAD) sequencing of the parents and progeny of a genetic cross. The resulting linkage map confirmed the LACHESIS clustering and ordering of scaffolds onto chromosomes and rectified large-scale inversions. Assessments of the <em>R. williamsianum</em> genome assembly and gene annotation estimate them to be 89% and 79% complete, respectively. Predicted coding sequences from genome annotation were used in syntenic analyses and for generating age distributions of synonymous substitutions/site between paralgous gene pairs, which identified whole-genome duplications (WGDs) in <em>R. williamsianum</em>. We then analyzed other publicly available Ericaceae genomes for shared WGDs. Based on our spatial and temporal analyses of paralogous gene pairs, we find evidence for two shared, ancient WGDs in <em>Rhododendron</em> and <em>Vaccinium</em> (cranberry/blueberry) members that predate the Ericaceae family and, in one case, the Ericales order.</p>

opencc-zeroOct 2020View details →
dryad36/100

Pioneering polyploids: the impact of whole-genome duplication on biome shifting in New Zealand Coprosma (Rubiaceae) and Veronica (Plantaginaceae)

<p>The role of whole-genome duplication in facilitating shifts into novel biomes remains unknown. Focusing on two diverse woody plant groups in New Zealand, <i>Coprosma </i>(Rubiaceae) and <i>Veronica </i>(Plantaginaceae), we investigate how biome occupancy varies with ploidy level, and test the hypothesis that whole-genome duplication increases the rate of biome shifting.</p> <p>Ploidy levels and biome occupancy (forest, open, and alpine) were determined for indigenous species in both clades. The distribution of low ploidy (<i>Coprosma</i>: 2<i>x</i>, <i>Veronica</i>: 6<i>x</i>) vs high ploidy (<i>Coprosma</i>: 4–10<i>x</i>, <i>Veronica</i>: 12–18<i>x</i>) species across biomes was tested statistically. Estimation of the phylogenetic history of biome occupancy and whole-genome duplication was performed using time-calibrated phylogenies and the R package BioGeoBEARS. Trait-dependent dispersal models were implemented to determine support for an increased rate of biome shifting among high ploidy lineages.</p> <p>We find support for a greater than random portion of high ploidy species occupying multiple biomes. We also find strong support for high ploidy taxa showing a three to eight-fold increase in the rate of biome shifts. These results suggest that whole-genome duplication promotes ecological expansion into new biomes.</p>

opencc-zeroNov 2020View details →
dryad36/100

Data from: Gene duplication, population genomics and species-level differentiation within a tropical mountain shrub

Gene duplication leads to paralogy, which complicates the de novo assembly of genotyping-by-sequencing (GBS) data. The issue of paralogous genes is exacerbated in plants, because they are particularly prone to gene duplication events. Paralogs are normally filtered from GBS data before undertaking population genomics or phylogenetic analyses. However, gene duplication plays an important role in the functional diversification of genes and it can also lead to the formation of postzygotic barriers. Using populations and closely related species of a tropical mountain shrub, we examine: (1) the genomic differentiation produced by putative orthologs, and (2) the distribution of recent gene duplication among lineages and geography. We find high differentiation among populations from isolated mountain peaks and species-level differentiation within what is morphologically described as a single species. The inferred distribution of paralogs among populations is congruent with taxonomy and shows that GBS could be used to examine recent gene duplication as a source of genomic differentiation of non-model species.

opencc-zeroDec 2013View details →
dryad36/100

Novel mitochondrial genome rearrangements including duplications and extensive heteroplasmy could underlie temperature adaptations in Antarctic notothenioid fishes

<p>Mitochondrial genomes are known for their compact size and conserved gene order, however, recent studies employing long-read sequencing technologies have revealed the presence of atypical mitogenomes in some species. In this study, we assembled and annotated the mitogenomes of five Antarctic notothenioids, including four icefishes (Champsocephalus gunnari, C. esox, Chaenocephalus aceratus, and Pseudochaenichthys georgianus) and the cold-specialized Trematomus borchgrevinki. Antarctic notothenioids are known to harbor some rearrangements in their mt genomes, however the extensive duplications in icefishes observed in our study have never been reported before. In the icefishes, we observed duplications of the protein coding gene ND6, two transfer RNAs, and the control region with different copy number variants present within the same individuals and with some ND6 duplications appearing to follow the canonical Duplication-Degeneration-Complementation (DDC) model in C. esox and C. gunnari. In addition, using long-read sequencing and k-mer analysis, we were able to detect extensive heteroplasmy in C. aceratus and C. esox. We also observed a large inversion in the mitogenome of T. borchgrevinki, along with the presence of tandem repeats in its control region. This study is the first in using long-read sequencing to assemble and identify structural variants and heteroplasmy in notothenioid mitogenomes and signifies the importance of long-reads in resolving complex mitochondrial architectures. Identification of such wide-ranging structural variants in the mitogenomes of these fishes could provide insight into the genetic basis of the atypical icefish mitochondrial physiology and more generally may provide insights about their potential role in cold adaptation.</p>

opencc-zeroDec 2023View details →
dryad36/100

Data from: Species tree estimation and the impact of gene loss following whole-genome duplication

<p>Whole-genome duplication (WGD) has been demonstrated to occur broadly and repeatedly in the evolutionary history of eukaryotes, and is recognized as a prominent evolutionary force, especially in plants. Immediately following WGD, most genes are present in two copies as paralogs. Due to this redundancy, one copy of a paralog pair commonly undergoes pseudogenization and is eventually lost. When speciation occurs shortly after WGD, however, differential loss of paralogs may lead to spurious phylogenetic inference resulting from the inclusion of pseudoorthologs – paralogous genes mistakenly identify as orthologs because they are present in single copes within each sampled species. The influence and impact of including pseudoorthologs versus true orthologs as result of gene extinction (or incomplete laboratory sampling) in a phylogenetic context is only recently starting to gain empirical attention. Moreover, few of these studies have yet to investigate this phenomenon in an explicit coalescent framework. Here, using mathematical models, numerous simulated data sets, and two newly assembled empirical data sets, we assess the effect of pseudoorthologs on species tree estimation under varying levels of incomplete lineage sorting (ILS) and different patterns of gene loss following WGD. When gene loss occurs in the terminal branches of the species tree, the alignment-based (BPP) and gene-tree-based (ASTRAL, MP-EST, and STAR) coalescent methods are adversely affected as the level of ILS increases. This can be greatly improved by sampling a sufficiently large number of genes. Under the same circumstances, however, concatenation methods consistently estimate incorrect species trees as the number of sampled genes increases. Furthermore, pseudoorthologs can mislead species tree inference if gene loss occurs in the internal branches of the species tree, where both coalescent and concatenation methods are prone to produce inconsistent results. However, pseudoorthologs are problematic when filtering only for single-copy genes in phylogenomic data sets. Pruning orthologs or even randomly selecting a copy from multi-copy genes can avoid most of those pseudoorthologs. These results underscore the importance of understanding the influence of pseudoorthologs in the phylogenomics era.</p>

opencc-zeroJun 2022View details →
zenodo36/100

Microscopy images - "An elevated rate of whole-genome duplications associated with carcinogen exposure in Black cancer patients"

<p>This repository contains microscopy data from the manuscript "An elevated rate of whole-genome duplications associated with carcinogen exposure in Black cancer patients" by Leanne M. Brown, Ryan A. Hagenson, Tilen Koklič, Iztok Urbančič, Janez Strancar, and Jason M. Sheltzer (preprint: 10.1101/2023.11.10.23298349; accepted for publication in Nature Communications).</p> <p>&nbsp;</p> <p>Each zip contains the set of images from individual multi-channel multi-position time-lapse experiment with different combinations of cells exposed to one nanomaterial. Files are named as: IMGxxxx_[ExperimentCode]_ROIxx_tile[TileNumber]_[Channel]_[CellType]_t[Timepoint].tif, where each of the varying elements in [..] denotes the following:</p> <ul> <li>[ExperimentCode]: tells which material the cells were exposed to - see decoding table in the file "material-codes.xlsx"</li> <li>[TileNumber]: two xy-tiles per condition</li> <li>[Channel]: cytoplasm of epi cells (LA4), membrane (MEM), cytoplasm of imu cells and tubulin (MHSTUB), nanomaterial (NANO)</li> <li>[CellType]: mono-culture of lung epithelial cells (epi), their coculture with macrophages (epiimu)</li> <li>[Timepoint]: consecutive number of the frame in the time-lapse</li> </ul> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Jul 2024View details →
dryad36/100

Investigating the effects of whole genome duplication on phenotypic plasticity: Implications for the invasion success of Giant Goldenrod (Solidago gigantea)

<p>Polyploidy commonly occurs in invasive species and phenotypic plasticity (PP, the ability to alter one's phenotype in different environments), is predicted to be enhanced in polyploids and contribute to their invasive success. However, empirical support that increased PP is frequent in polyploids and/or confers invasive success is limited. Here, we investigated if polyploids are more pre-adapted to become invasive than diploids via the scaling of trait values and PP with ploidy-level, and if post-introduction selection has led to a divergence in trait values and PP responses between native- and non-native cytotypes. We grew diploid, tetraploid (from both native North American and non-native European ranges), and hexaploid <em>Solidago gigantea</em> in pots outside with low, medium, and high soil nitrogen and phosphorus (NP) amendments, and measured traits related to growth, asexual reproduction, physiology, and insects/pathogen resistance. We found little evidence to suggest that polyploidy and post-selection shaped mean trait and PP responses. To examine invasion dynamics, we compared diploids to tetraploids (as their introduction into Europe was more likely), and found that tetraploids had greater pathogen resistance, photosynthetic capacities, and water-use efficiencies and generally performed better under NP enrichments. Furthermore, tetraploids invested more into roots than shoots in low NP and into shoots than roots in high NP and this resource strategy is beneficial under variable NP conditions. Lastly, native-tetraploids exhibited greater plasticity in biomass accumulation, clonal-ramet production and water-use efficiency. Cumulatively, tetraploid <em>S. gigantea</em> possesses traits that might have pre-disposed and enabled them to become successful invaders. Our findings highlight that trait expression and invasive species dynamics are nuance while also providing insight into the invasion success and cyto-geographic patterning of <em>S. gigantea </em>that can be broadly applied to other invasive species with polyploid complexes.</p>

opencc-zeroAug 2023View details →
dryad36/100

Kinetochore and ionomic adaptation to whole genome duplication

<p>Whole genome duplication (WGD) brings challenges to key processes like meiosis but nevertheless is associated with diversification in all kingdoms. How is WGD tolerated, and what processes commonly evolve to stabilize the new polyploid lineage? Here we study this in <em>Cochlearia</em> spp., which have experienced multiple rounds of WGD in the last 300,000 years. We first generate a chromosome-scale genome and sequence 113 individuals from 33 diploid, tetraploid, hexaploid, and outgroup populations. We detect the clearest post-WGD selection signatures in functionally interacting kinetochore components and ion transporters. We structurally model these derived selected alleles, associating them with known WGD-relevant functional variation, and compare these results to independent recent post-WGD selection in <em>Arabidopsis</em> <em>arenosa</em> and <em>Cardamine</em> <em>amara</em>. Some of the same biological processes evolve in all three WGDs, but specific genes recruited are flexible. This points to a polygenic basis for modifying systems that control the kinetochore, meiotic crossover number, DNA repair, ion homeostasis, and cell cycle. Given that DNA management (especially repair) is the most salient category with the strongest selection signal, we speculate that the generation rate of structural genomic variants may be altered by WGD in young polyploids, contributing to their occasionally spectacular adaptability observed across kingdoms.</p>

opencc-zeroOct 2023View details →
dryad36/100

Data from: Genome duplication effects on functional traits and fitness are genetic context and species dependent: studies of synthetic polyploid Fragaria

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad36/100

Pioneering polyploids: the impact of whole-genome duplication on biome shifting in New Zealand Coprosma (Rubiaceae) and Veronica (Plantaginaceae)

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad36/100

Investigating the effects of whole genome duplication on phenotypic plasticity: Implications for the invasion success of Giant Goldenrod (Solidago gigantea)

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad36/100

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad36/100

Novel mitochondrial genome rearrangements including duplications and extensive heteroplasmy could underlie temperature adaptations in Antarctic notothenioid fishes

Open the record for dataset details and reuse information.

publicDec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record