Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

104

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

104 results for “gene duplications”

Learn how ShareScore rates datasets ↗
zenodo44/100

Supplemental data for: Increased mutation and gene conversion within human segmental duplications

<p>Data used for figure generation and analysis in: <strong>Increased mutation and gene conversion within human segmental duplications</strong></p> <ol> <li>new-assemblies.zip contains all the new assemblies added in this work beyond the HPRC assemblies (Clint PTR, CHM1, HG00514, NA12878, HG03125). All other assemblies used in this analysis are available&nbsp;through the HPRC:&nbsp;<a href="https://github.com/human-pangenomics/HPP_Year1_Assemblies/blob/main/assembly_index/Year1_assemblies_v2_genbank.index">assembly_index/Year1_assemblies_v2_genbank.index</a>.</li> <li>all-sample.vcf is a vcf file with all the variant calls used in this analysis.&nbsp;</li> <li>alignments.zip contains all the syntenic&nbsp;alignments used for analysis.&nbsp;</li> <li>data.zip contains annotation data and other information used in analysis and figure making.&nbsp;</li> <li>Online tables 1-4 (Online-tables.xlsx)</li> </ol> <p>Code used in figure making and analysis is&nbsp;on <a href="https://github.com/mrvollger/sd-divergence-and-igc-figures">GitHub</a>.</p> <p>Snakemake pipelines used in the analysis are also on GitHub:</p> <ul> <li>Assembly alignment and IGC calling: https://github.com/mrvollger/asm-to-reference-alignment</li> <li>Variant calling from assembly alignments: https://github.com/mrvollger/sd-divergence</li> <li>Analysis of the triplet content of SNVs: https://github.com/mrvollger/mutyper_workflow</li> </ul>

opencc-by-4.0Feb 2023View details →
dryad40/100

Data for: The 3-dimensional genome drives the evolution of asymmetric gene duplicates via enhancer capture-divergence

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways

Background: Teak, a member of the Lamiaceae family, produces one of the most expensive hardwoods in the world. High demand coupled with deforestation have caused a decrease in natural teak forests, and future supplies will be reliant on teak plantations. Hence, selection of teak tree varieties for clonal propagation with superior growth performance is of great importance, and access to high-quality genetic and genomic resources can accelerate the selection process by identifying genes underlying desired traits. Findings: To facilitate teak research and variety improvement, we generated a highly contiguous, chromosomal-scale genome assembly using high-coverage PacBio long reads coupled with high-throughput chromatin conformation capture. Of the 18 teak chromosomes, we generated 17 near-complete pseudomolecules with one chromosome present as two chromosome arm scaffolds. Genome annotation yielded 31,168 genes encoding 46,826 gene models, of which, 39,930 and 41,155 had Pfam domain and expression evidence, respectively. We identified 14 clusters of tandem-duplicated terpene synthases (TPSs), genes central to the biosynthesis of terpenes which are involved in plant defense and pollinator attraction. Transcriptome analysis revealed 10 TPSs highly expressed in woody tissues, of which, 8 were in tandem, revealing the importance of resolving tandemly duplicated genes and the quality of the assembly and annotation. We also validated the enzymatic activity of four TPSs to demonstrate the function of key TPSs. Conclusions: In summary, this high-quality chromosomal-scale assembly and functional annotation of the teak genome will facilitate the discovery of candidate genes related to traits critical for sustainable production of teak and for anti-insecticidal natural products.

opencc-zeroDec 2018View details →
dryad36/100

Data from: Gene duplication, population genomics and species-level differentiation within a tropical mountain shrub

Gene duplication leads to paralogy, which complicates the de novo assembly of genotyping-by-sequencing (GBS) data. The issue of paralogous genes is exacerbated in plants, because they are particularly prone to gene duplication events. Paralogs are normally filtered from GBS data before undertaking population genomics or phylogenetic analyses. However, gene duplication plays an important role in the functional diversification of genes and it can also lead to the formation of postzygotic barriers. Using populations and closely related species of a tropical mountain shrub, we examine: (1) the genomic differentiation produced by putative orthologs, and (2) the distribution of recent gene duplication among lineages and geography. We find high differentiation among populations from isolated mountain peaks and species-level differentiation within what is morphologically described as a single species. The inferred distribution of paralogs among populations is congruent with taxonomy and shows that GBS could be used to examine recent gene duplication as a source of genomic differentiation of non-model species.

opencc-zeroDec 2013View details →
zenodo36/100

Supporting data for Bryłka et al., 2023 Gene duplication, shifting selection, and functional diversification of silicon transporter proteins in marine and freshwater diatoms.

<p>These are supporting files used in the analysis in Bryłka et al., 2023&nbsp;&quot;<em>Gene duplication, shifting selection, and functional diversification of silicon transporter proteins in marine and freshwater diatoms&quot;.&nbsp;</em> All data newly generated from the Alverson lab.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
dryad36/100

Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster

<p>Tremendous plant metabolic diversity arises from phylogenetically-restricted specialized metabolic pathways. Specialized metabolites are synthesized in dedicated cells or tissues, with pathway genes sometimes colocalizing in biosynthetic gene clusters (BGCs). However, the mechanisms by which spatial expression patterns arise and the role of BGCs in pathway evolution remain underappreciated. In this study, we investigated the mechanisms driving acylsugar evolution in the Solanaceae. Previously thought to be restricted to glandular trichomes, acyl sugars were recently discovered in cultivated tomato roots. We demonstrated that acyl sugars in cultivated tomato roots and trichomes have different sugar cores, identified root-enriched paralogs of trichome acyl sugar pathway genes, and characterized a key paralog required for root acyl sugar biosynthesis, <em>SlASAT1-LIKE</em> (<em>SlASAT1-L</em>), which is nested within a previously-reported trichome acyl sugar BGC. Finally, we provided evidence that <em>ASAT1-L</em> arose through duplication of its paralog, <em>ASAT1</em>, and was trichome-expressed before acquiring root-specific expression in the <em>Solanum</em> genus. Our results illuminate the genomic context and molecular mechanisms underpinning metabolic diversity in plants.</p>

opencc-zeroApr 2024View details →
zenodo36/100

The adaptive value of recombination in resolving intralocus sexual conflict by gene duplication.

<p>Datasets and code for the study titled "The adaptive value of recombination in resolving intralocus sexual conflict by gene duplication".</p>

opencc-by-4.0Apr 2024View details →
dryad36/100

Data from: Species tree estimation and the impact of gene loss following whole-genome duplication

<p>Whole-genome duplication (WGD) has been demonstrated to occur broadly and repeatedly in the evolutionary history of eukaryotes, and is recognized as a prominent evolutionary force, especially in plants. Immediately following WGD, most genes are present in two copies as paralogs. Due to this redundancy, one copy of a paralog pair commonly undergoes pseudogenization and is eventually lost. When speciation occurs shortly after WGD, however, differential loss of paralogs may lead to spurious phylogenetic inference resulting from the inclusion of pseudoorthologs – paralogous genes mistakenly identify as orthologs because they are present in single copes within each sampled species. The influence and impact of including pseudoorthologs versus true orthologs as result of gene extinction (or incomplete laboratory sampling) in a phylogenetic context is only recently starting to gain empirical attention. Moreover, few of these studies have yet to investigate this phenomenon in an explicit coalescent framework. Here, using mathematical models, numerous simulated data sets, and two newly assembled empirical data sets, we assess the effect of pseudoorthologs on species tree estimation under varying levels of incomplete lineage sorting (ILS) and different patterns of gene loss following WGD. When gene loss occurs in the terminal branches of the species tree, the alignment-based (BPP) and gene-tree-based (ASTRAL, MP-EST, and STAR) coalescent methods are adversely affected as the level of ILS increases. This can be greatly improved by sampling a sufficiently large number of genes. Under the same circumstances, however, concatenation methods consistently estimate incorrect species trees as the number of sampled genes increases. Furthermore, pseudoorthologs can mislead species tree inference if gene loss occurs in the internal branches of the species tree, where both coalescent and concatenation methods are prone to produce inconsistent results. However, pseudoorthologs are problematic when filtering only for single-copy genes in phylogenomic data sets. Pruning orthologs or even randomly selecting a copy from multi-copy genes can avoid most of those pseudoorthologs. These results underscore the importance of understanding the influence of pseudoorthologs in the phylogenomics era.</p>

opencc-zeroJun 2022View details →
dryad36/100

DNA methylation signatures of duplicate gene evolution in angiosperms

<p><span>Gene duplication is a source of evolutionary novelty. DNA methylation may play a role in the evolution of duplicate genes through its association with gene expression. While this relationship is examined to varying extents in a few individual species, the generalizability of these results at either a broad phylogenetic scale with species of differing duplication histories or across a population remains unknown. We apply a comparative epigenomics approach to 43 angiosperm species across the phylogeny and a population of 928 <em>Arabidopsis</em> <em>thaliana</em> accessions, examining the association of DNA methylation with paralog evolution. Genic DNA methylation is differentially associated with duplication type, the age of duplication, sequence evolution, and gene expression. Whole genome duplicates are typically enriched for CG-only gene-body methylated or unmethylated genes, while single-gene duplications are typically enriched for non-CG methylated or unmethylated genes. Non-CG methylation, in particular, was characteristic of more recent single-gene duplicates. Core angiosperm gene families are differentiated into those which preferentially retain paralogs and 'duplication-resistant' families, which convergently revert to singletons following duplication. Duplication-resistant families which still have paralogous copies are, uncharacteristically for core angiosperm genes, enriched for non-CG methylation. Non-CG methylated paralogs have higher rates of sequence evolution, higher frequency of presence-absence variation, and more limited expression. This suggests that silencing by non-CG methylation may be important to maintaining dosage following duplication and be a precursor to fractionation. Our results indicate that genic methylation marks differing evolutionary trajectories and fates between paralogous genes and have a role in maintaining dosage following duplication.</span></p>

opencc-zeroMar 2023View details →
ClinicalTrials.gov36/100

AAV9 U7snRNA Gene Therapy to Treat Boys With DMD Exon 2 Duplications.

ClinicalTrials.gov study NCT04240314. IPD Sharing: Not stated. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

DupLoss-2: Improved phylogenomic species tree inference under gene duplication and loss

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad36/100

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Data from: A chromosomal-scale genome assembly of Tectona grandis reveals the importance of tandem gene duplication and enables discovery of genes in natural product biosynthetic pathways

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad36/100

Data for: Tomato root specialized metabolites evolved through gene duplication and regulatory divergence within a biosynthetic gene cluster

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Data from: Gene duplication and gene expression changes play a role in the evolution of candidate pollen feeding genes in Heliconius butterflies

Open the record for dataset details and reuse information.

publicJul 2017View details →
dryad36/100

Gene duplication captures morph-specific promoter usage in the evolution of aphid wing dimorphisms

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad36/100

Data from: Revisiting ancient whole-genome duplications in the seed and flowering plants through the lens of dosage-sensitive genes

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad36/100

Data from: Gene duplication, population genomics and species-level differentiation within a tropical mountain shrub

Open the record for dataset details and reuse information.

publicSep 2014View details →
dryad36/100

DNA methylation signatures of duplicate gene evolution in angiosperms

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Data from: Species tree estimation and the impact of gene loss following whole-genome duplication

Open the record for dataset details and reuse information.

publicJun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record