Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

104

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

104 results for “gene duplications”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Invasive invertebrates associated with highly duplicated gene content

Open the record for dataset details and reuse information.

publicMar 2019View details →
dryad28/100

Data from: Subfunctionalization of peroxisome proliferator response elements accounts for retention of duplicated fabp1 genes in zebrafish

Open the record for dataset details and reuse information.

publicJul 2016View details →
geo24/100

Population Structure and Comparative Genome Hybridization of European flor yeast reveal a unique group of Saccharomyces cerevisiae strains with few gene duplications in their genome

GEO Series GSE55925. Saccharomyces cerevisiae; Schizosaccharomyces pombe; Saccharomyces cerevisiae x Saccharomyces kudriavzevii. 25 samples. Type: Genome variation profiling by array.

openGEO-OpenJun 2014View details →
geo24/100

Maedi-visna virus (MVV) preferrentially integrates into genes and generates 6-bp duplications

GEO Series GSE87786. Ovis aries. 2 samples. Type: Other.

openGEO-OpenJan 2017View details →
geo24/100

PICKLE RELATED 2 is a neofunctionalised gene duplicate under positive Darwinian selection with antagonistic effects to the ancestral PICKLE gene on the seed transcriptome

GEO Series GSE236025. Arabidopsis thaliana. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2023View details →
geo24/100

Gene duplication of type-B ARR transcription factors systematically extends transcriptional regulatory structures in Arabidopsis

GEO Series GSE62597. Arabidopsis thaliana. 20 samples. Type: Expression profiling by array.

openGEO-OpenFeb 2015View details →
geo24/100

Cumulative Impact of Polychlorinated Biphenyl and Large Chromosomal Duplications on DNA Methylation, Chromatin, and Expression of Autism Candidate Genes.

GEO Series GSE81541. Homo sapiens. 64 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenDec 2016View details →
geo24/100

Duplication of autism-related gene Chd8 leads to behavioral hyperactivity and neurodevelopmental defects in mice

GEO Series GSE263334. Mus musculus. 40 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenApr 2025View details →
geo24/100

RNA-seq analysis of alternative splicing events in duplicated genes of Arabidopsis thaliana indicates considerable qualitative and quantitative divergence

GEO Series GSE57579. Arabidopsis thaliana. 3 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2014View details →
geo24/100

Gene expression variation during grape genome duplication

GEO Series GSE119442. Vitis vinifera. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2020View details →
geo24/100

The naked endosperm genes encode duplicate ID domain transcription factors required for maize endosperm differentiation

GEO Series GSE61057. Zea mays. 12 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2015View details →
geo24/100

Mutational and Transcriptional Landscape of Spontaneous Gene Duplications and Deletions in Caenorhabditis elegans

GEO Series GSE112821. Caenorhabditis elegans. 64 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenApr 2018View details →
geo24/100

PML is a novel constitutive component of chromatin domains of enriched repetitive elements and duplicated gene clusters in cancer cells

GEO Series GSE261929. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2024View details →
geo24/100

Duplication of autism-related gene Chd8 leads to behavioral hyperactivity and neurodevelopmental defects in mice [cortex]

GEO Series GSE288732. Mus musculus. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2025View details →
geo24/100

Array-CGH and Next-generation sequencing of duplication CNVs reveals that most are tandem and some disrupt genes at breakpoints

GEO Series GSE62657. Homo sapiens. 161 samples. Type: Genome variation profiling by array.

openGEO-OpenOct 2014View details →
geo24/100

Investigating determinants of aneuploidy toxicity using gene duplication in Saccharomyces cerevisiae

GEO Series GSE263221. Saccharomyces cerevisiae. 83 samples. Type: Other.

openGEO-OpenMay 2024View details →
geo24/100

A Y-linked duplication of anti-Mullerian hormone is the sex determination gene in threespine stickleback

GEO Series GSE296766. Gasterosteus aculeatus. 36 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2025View details →
geo24/100

Cdx ParaHox genes acquired distinct developmental roles after gene duplication in vertebrate evolution

GEO Series GSE71006. Xenopus tropicalis. 18 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2015View details →
dryad24/100

Data from: Analysis of phylogenomic datasets reveals conflict, concordance, and gene duplications with examples from animals and plants

Background: The use of transcriptomic and genomic datasets for phylogenetic reconstruction has become increasingly common as researchers attempt to resolve recalcitrant nodes with increasing amounts of data. The large size and complexity of these datasets introduce significant phylogenetic noise and conflict into subsequent analyses. The sources of conflict may include hybridization, incomplete lineage sorting, or horizontal gene transfer, and may vary across the phylogeny. For phylogenetic analysis, this noise and conflict has been accommodated in one of several ways: by binning gene regions into subsets to isolate consistent phylogenetic signal; by using gene-tree methods for reconstruction, where conflict is presumed to be explained by incomplete lineage sorting (ILS); or through concatenation, where noise is presumed to be the dominant source of conflict. The results provided herein emphasize that analysis of individual homologous gene regions can greatly improve our understanding of the underlying conflict within these datasets. Results: Here we examined two published transcriptomic datasets, the angiosperm group Caryophyllales and the aculeate Hymenoptera, for the presence of conflict, concordance, and gene duplications in individual homologs across the phylogeny. We found significant conflict throughout the phylogeny in both datasets and in particular along the backbone. While some nodes in each phylogeny showed patterns of conflict similar to what might be expected with ILS alone, the backbone nodes also exhibited low levels of phylogenetic signal. In addition, certain nodes, especially in the Caryophyllales, had highly elevated levels of strongly supported conflict that cannot be explained by ILS alone. Conclusion: This study demonstrates that phylogenetic signal is highly variable in phylogenomic data sampled across related species and poses challenges when conducting species tree analyses on large genomic and transcriptomic datasets. Further insight into the conflict and processes underlying these complex datasets is necessary to improve and develop adequate models for sequence analysis and downstream applications. To aid this effort, we developed the open source software phyparts (https://bitbucket.org/blackrim/phyparts), which calculates unique, conflicting, and concordant bipartitions, maps gene duplications, and outputs summary statistics such as internode certainy (ICA) scores and node-specific counts of gene duplications.

opencc-zeroDec 2014View details →
dryad24/100

Data from: Molecular evolution accompanying functional divergence of duplicated genes along the plant starch biosynthesis pathway

Background: Starch is the main source of carbon storage in the Archaeplastida. The Starch Biosynthesis Pathway (SBP) emerged from cytosolic glycogen metabolism shortly after plastid endosymbiosis and was redirected to the plastid stroma during the green lineage divergence. The SBP is a complex network of genes, most of which are members of large multigene families. While some gene duplications occurred in the Archaeplastida ancestor, most were generated during the SBP redirection process, and the remaining few paralogs were generated through compartmentalization or tissue specialization during the evolution of the land plants. In the present study, we tested models of duplicated gene evolution in order to understand the evolutionary forces that have led to the development of SBP in angiosperms. We combined phylogenetic analyses and tests on the rates of evolution along branches emerging from major duplication events in six gene families encoding SBP enzymes. Results: We found evidence of positive selection along branches following cytosolic or plastidial specialization in two starch phosphorylases and identified numerous residues that exhibited changes in volume, polarity or charge. Starch synthases, branching and debranching enzymes functional specializations were also accompanied by accelerated evolution. However, none of the sites targeted by selection corresponded to known functional domains, catalytic or regulatory. Interestingly, among the 13 duplications tested, 7 exhibited evidence of positive selection in both branches emerging from the duplication, 2 in only one branch, and 4 in none of the branches. Conclusions: The majority of duplications were followed by accelerated evolution targeting specific residues along both branches. This pattern was consistent with the optimization of the two sub-functions originally fulfilled by the ancestral gene before duplication. Our results thereby provide strong support to the so-called "Escape from Adaptive Conflict" (EAC) model. Because none of the residues targeted by selection occurred in characterized functional domains, we propose that enzyme specialization has occurred through subtle changes in affinity, activity or interaction with other enzymes in complex formation, while the basic function defined by the catalytic domain has been maintained.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record