Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
104
datasets available to search
ShareScore release 0.9.0
Dataset results
104 results for “gene duplications”
Data from: Invasive invertebrates associated with highly duplicated gene content
Open the record for dataset details and reuse information.
Data from: Subfunctionalization of peroxisome proliferator response elements accounts for retention of duplicated fabp1 genes in zebrafish
Open the record for dataset details and reuse information.
Population Structure and Comparative Genome Hybridization of European flor yeast reveal a unique group of Saccharomyces cerevisiae strains with few gene duplications in their genome
GEO Series GSE55925. Saccharomyces cerevisiae; Schizosaccharomyces pombe; Saccharomyces cerevisiae x Saccharomyces kudriavzevii. 25 samples. Type: Genome variation profiling by array.
Maedi-visna virus (MVV) preferrentially integrates into genes and generates 6-bp duplications
GEO Series GSE87786. Ovis aries. 2 samples. Type: Other.
PICKLE RELATED 2 is a neofunctionalised gene duplicate under positive Darwinian selection with antagonistic effects to the ancestral PICKLE gene on the seed transcriptome
GEO Series GSE236025. Arabidopsis thaliana. 8 samples. Type: Expression profiling by high throughput sequencing.
Gene duplication of type-B ARR transcription factors systematically extends transcriptional regulatory structures in Arabidopsis
GEO Series GSE62597. Arabidopsis thaliana. 20 samples. Type: Expression profiling by array.
Cumulative Impact of Polychlorinated Biphenyl and Large Chromosomal Duplications on DNA Methylation, Chromatin, and Expression of Autism Candidate Genes.
GEO Series GSE81541. Homo sapiens. 64 samples. Type: Methylation profiling by high throughput sequencing.
Duplication of autism-related gene Chd8 leads to behavioral hyperactivity and neurodevelopmental defects in mice
GEO Series GSE263334. Mus musculus. 40 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
RNA-seq analysis of alternative splicing events in duplicated genes of Arabidopsis thaliana indicates considerable qualitative and quantitative divergence
GEO Series GSE57579. Arabidopsis thaliana. 3 samples. Type: Expression profiling by high throughput sequencing.
Gene expression variation during grape genome duplication
GEO Series GSE119442. Vitis vinifera. 6 samples. Type: Expression profiling by high throughput sequencing.
The naked endosperm genes encode duplicate ID domain transcription factors required for maize endosperm differentiation
GEO Series GSE61057. Zea mays. 12 samples. Type: Expression profiling by high throughput sequencing.
Mutational and Transcriptional Landscape of Spontaneous Gene Duplications and Deletions in Caenorhabditis elegans
GEO Series GSE112821. Caenorhabditis elegans. 64 samples. Type: Genome variation profiling by genome tiling array.
PML is a novel constitutive component of chromatin domains of enriched repetitive elements and duplicated gene clusters in cancer cells
GEO Series GSE261929. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Duplication of autism-related gene Chd8 leads to behavioral hyperactivity and neurodevelopmental defects in mice [cortex]
GEO Series GSE288732. Mus musculus. 10 samples. Type: Expression profiling by high throughput sequencing.
Array-CGH and Next-generation sequencing of duplication CNVs reveals that most are tandem and some disrupt genes at breakpoints
GEO Series GSE62657. Homo sapiens. 161 samples. Type: Genome variation profiling by array.
Investigating determinants of aneuploidy toxicity using gene duplication in Saccharomyces cerevisiae
GEO Series GSE263221. Saccharomyces cerevisiae. 83 samples. Type: Other.
A Y-linked duplication of anti-Mullerian hormone is the sex determination gene in threespine stickleback
GEO Series GSE296766. Gasterosteus aculeatus. 36 samples. Type: Expression profiling by high throughput sequencing.
Cdx ParaHox genes acquired distinct developmental roles after gene duplication in vertebrate evolution
GEO Series GSE71006. Xenopus tropicalis. 18 samples. Type: Expression profiling by high throughput sequencing.
Data from: Analysis of phylogenomic datasets reveals conflict, concordance, and gene duplications with examples from animals and plants
Background: The use of transcriptomic and genomic datasets for phylogenetic reconstruction has become increasingly common as researchers attempt to resolve recalcitrant nodes with increasing amounts of data. The large size and complexity of these datasets introduce significant phylogenetic noise and conflict into subsequent analyses. The sources of conflict may include hybridization, incomplete lineage sorting, or horizontal gene transfer, and may vary across the phylogeny. For phylogenetic analysis, this noise and conflict has been accommodated in one of several ways: by binning gene regions into subsets to isolate consistent phylogenetic signal; by using gene-tree methods for reconstruction, where conflict is presumed to be explained by incomplete lineage sorting (ILS); or through concatenation, where noise is presumed to be the dominant source of conflict. The results provided herein emphasize that analysis of individual homologous gene regions can greatly improve our understanding of the underlying conflict within these datasets. Results: Here we examined two published transcriptomic datasets, the angiosperm group Caryophyllales and the aculeate Hymenoptera, for the presence of conflict, concordance, and gene duplications in individual homologs across the phylogeny. We found significant conflict throughout the phylogeny in both datasets and in particular along the backbone. While some nodes in each phylogeny showed patterns of conflict similar to what might be expected with ILS alone, the backbone nodes also exhibited low levels of phylogenetic signal. In addition, certain nodes, especially in the Caryophyllales, had highly elevated levels of strongly supported conflict that cannot be explained by ILS alone. Conclusion: This study demonstrates that phylogenetic signal is highly variable in phylogenomic data sampled across related species and poses challenges when conducting species tree analyses on large genomic and transcriptomic datasets. Further insight into the conflict and processes underlying these complex datasets is necessary to improve and develop adequate models for sequence analysis and downstream applications. To aid this effort, we developed the open source software phyparts (https://bitbucket.org/blackrim/phyparts), which calculates unique, conflicting, and concordant bipartitions, maps gene duplications, and outputs summary statistics such as internode certainy (ICA) scores and node-specific counts of gene duplications.
Data from: Molecular evolution accompanying functional divergence of duplicated genes along the plant starch biosynthesis pathway
Background: Starch is the main source of carbon storage in the Archaeplastida. The Starch Biosynthesis Pathway (SBP) emerged from cytosolic glycogen metabolism shortly after plastid endosymbiosis and was redirected to the plastid stroma during the green lineage divergence. The SBP is a complex network of genes, most of which are members of large multigene families. While some gene duplications occurred in the Archaeplastida ancestor, most were generated during the SBP redirection process, and the remaining few paralogs were generated through compartmentalization or tissue specialization during the evolution of the land plants. In the present study, we tested models of duplicated gene evolution in order to understand the evolutionary forces that have led to the development of SBP in angiosperms. We combined phylogenetic analyses and tests on the rates of evolution along branches emerging from major duplication events in six gene families encoding SBP enzymes. Results: We found evidence of positive selection along branches following cytosolic or plastidial specialization in two starch phosphorylases and identified numerous residues that exhibited changes in volume, polarity or charge. Starch synthases, branching and debranching enzymes functional specializations were also accompanied by accelerated evolution. However, none of the sites targeted by selection corresponded to known functional domains, catalytic or regulatory. Interestingly, among the 13 duplications tested, 7 exhibited evidence of positive selection in both branches emerging from the duplication, 2 in only one branch, and 4 in none of the branches. Conclusions: The majority of duplications were followed by accelerated evolution targeting specific residues along both branches. This pattern was consistent with the optimization of the two sub-functions originally fulfilled by the ancestral gene before duplication. Our results thereby provide strong support to the so-called "Escape from Adaptive Conflict" (EAC) model. Because none of the residues targeted by selection occurred in characterized functional domains, we propose that enzyme specialization has occurred through subtle changes in affinity, activity or interaction with other enzymes in complex formation, while the basic function defined by the catalytic domain has been maintained.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.