Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

281

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

281 results for “copy number variation”

Learn how ShareScore rates datasets ↗
dryad40/100

Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture

<p>Anterior cruciate ligament (ACL) rupture is an important condition of the human knee. Second ruptures are common and societal costs are substantial. Canine cranial cruciate ligament (CCL) rupture closely models the human disease. CCL rupture is common in the Labrador Retriever (5.79% prevalence), ~100-fold more prevalent than in humans. Labrador Retriever CCL rupture is a polygenic complex disease, based on genome-wide association study (GWAS) of single nucleotide polymorphism (SNP) markers. Dissection of genetic variation in complex traits can be enhanced by studying structural variation, including copy number variants (CNVs). Dogs are an ideal model for CNV research because of reduced genetic variability within breeds and extensive phenotypic diversity across breeds. We studied the genetic etiology of CCL rupture by association analysis of CNV regions (CNVRs) using 110 case and 164 control Labrador Retrievers. CNVs were called from SNPs using three different programs (PennCNV, CNVPartition, and QuantiSNP). After quality control, CNV calls were combined to create CNVRs using ParseCNV and an association analysis was performed. We found no strong effect CNVRs but found 46 small effect (max(T) permutation P&lt;0.05) CCL rupture associated CNVRs in 22 autosomes; 25 were deletions and 21 were duplications. Of the 46 CCL rupture associated CNVRs, we identified 39 unique regions. Thirty four were identified by a single calling algorithm, 3 were identified by two calling algorithms, and 2 were identified by all three algorithms. For 42 of the associated CNVRs, frequency in the population was &lt;10% while 4 occurred at a frequency in the population ranging from 10-25%. Average CNVR length was 198,872bp and CNVRs covered 0.11 to 0.15% of the genome. All CNVRs were associated with case status. CNVRs did not overlap previous canine CCL rupture risk loci identified by GWAS. Associated CNVRs contained 152 annotated genes; 12 CNVRs did not have genes mapped to CanFam3.1. Using pathway analysis, a cluster of 19 homeobox domain transcript regulator genes was associated with CCL rupture (P=6.6E-13). This gene cluster influences cranial-caudal body pattern formation during embryonic limb development. Clustered genes were found in 3 CNVRs on chromosome 14 (HoxA), 28 (NKX6-2), and 36 (HoxD). When analysis was limited to deletion CNVRs, the association was strengthened (P=8.7E-16). This study suggests a component of the polygenic risk of CCL rupture in Labrador Retrievers is associated with small effect CNVs and may include aspects of stifle morphology regulated by homeobox domain transcript regulator genes.</p>

opencc-zeroDec 2020View details →
zenodo40/100

CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data

<p>To date, ancient genome analyses have been largely confined to the study of single nucleotide polymorphisms (SNPs). Copy number variants (CNVs) are a major contributor of disease and of evolutionary adaptation, but identifying CNVs in ancient shotgun-sequenced genomes is hampered by (i) most published genomes being &lt;1x&nbsp;coverage, (ii) ancient DNA fragments being typically &lt;80 bps. These characteristics preclude state-of-the-art CNV detection software to be effectively applied to ancient genomes. Here we present CONGA, an algorithm tailored for genotyping deletion and duplication events in genomes with low depths of coverage. Simulations and down-sampling experiments show that CONGA can genotype deletions &gt;1 kbps with F-scores &gt;0.75 at &gt;=1x, and distinguish between heterozygous and homozygous states. Using CONGA, we analyse deletion events at 10,018 loci in 56 ancient human genomes spanning the last 50,000 years, with coverages 0.4x-26x. We show that inter-individual genetic diversity measured using deletions and SNPs are highly correlated, as in modern-day genomes, confirming that deletion frequencies broadly reflect demographic history. We also identify signatures of strong purifying selection on deletions in ancient-genomes, such as an excess of singletons compared to those in SNPs. CONGA paves the way for systematic studies of drift, mutation load, and adaptation in ancient and modern-day gene pools through the lens of CNVs.</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

Dataset for: Ace and ace-like genes of invasive redlegged earth mite: Copy number variation, target-site mutations, and their associations with organophosphate insensitivity

<p class="MsoNormal">This repository contains the scripts and data required to replicate the analyses in Thia et al.'s, "Evolution of an acetylcholinesterase<em> </em>gene complex and its contribution toward organophosphate insensitivity in an invasive mite pest", submitted to <em>Pest Management Science</em>.</p> <p class="MsoNormal">In this work, Thia et al. use a combination of experimental selection and pool-seq genomic analyses to understand the genetic mechanisms underpinning organophosphate insensitivity in the redlegged earth mite, <em>Halotydeus destructor</em>. There is a special emphasis on disentangling the roles of copy number variation and target-site mutations in the acetylcholinesterase genes, <em>ace,</em> and radiated <em>ace</em>-like genes<span>.</span></p> <p class="MsoNormal">There are three major analyses: (1) an F<sub>ST</sub> genome scan to identify outlier loci between alive (insensitive) and dead (sensitive) mites; (2) an analysis of <em>ace </em>copy number variation between alive and dead mites; and (3) an analysis of candidate target-site mutations in the <em>ace</em> gene.</p>

opencc-zeroMay 2023View details →
zenodo40/100

Copy number variations and their effect on the plasma proteome | Associations

<p>Results and supplementary data from a GWAS investigation the relationship of CNVs and blood protein measurements.</p>

opencc-by-4.0Sep 2022View details →
dryad40/100

Dataset for: Ace and ace-like genes of invasive redlegged earth mite: Copy number variation, target-site mutations, and their associations with organophosphate insensitivity

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad40/100

Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture

Open the record for dataset details and reuse information.

publicDec 2020View details →
zenodo36/100

Adaptation by copy number variation increases insecticide resistance in the fall armyworm (vcf files)

<p>Here, I deposite&nbsp;vcf files used for a paper, entitled &#39;Adaptation by copy number variation increases insecticide resistance in the fall armyworm&#39;.</p> <p>genotype.vcf and&nbsp;SNP.filtered.vcf.gz have the information of CNVs and SNPs, respectively.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
dryad36/100

Data from: Extreme copy number variation at a tRNA ligase gene affecting phenology and fitness in yellow monkeyflowers

Copy number variation (CNV) is a major part of the genetic diversity segregating within populations, but remains poorly understood relative to single nucleotide variation. Here, we report on a tRNA ligase gene (RLG1a) exhibiting unprecedented, and fitness-relevant, CNV within an annual population of the yellow monkeyflower Mimulus guttatus. Variation at RLG1a was associated with multiple traits in pooled population resequencing (PoolSeq) scans of phenotypic and phenological cohorts. Five of 35 (14%) of resequenced inbred lines carried three-copy variants of RLG1a (trip+), and trip+ lines exhibited elevated RLG1a expression. trip+ carriers, in addition to being over-represented in late-flowering and large-flowered PoolSeq populations, flowered later under stressful conditions in a greenhouse experiment (P &lt; 0.05). In early-flowering wild cohorts, we discovered an additional rare variant (high+) that carries 250-300 copies of RLG1a totaling ~5.7Mb, equivalent to 20-40% of a chromosome. Mendelian segregation of diagnostic alleles and qPCR-based copy counts In the progeny of a high+ carrier, indicate that high+ is a single tandem array unlinked from the single copy RLG1a locus in the reference genome. In the wild, high+ carriers had highest fitness in two dry and/or hot years (2015 and 2017; both P &lt; 0.01), while single copy individuals were twice as fecund as either CNV genotype in a lush year (2016: p &lt; 0.005). Our results demonstrate fluctuating selection on CNVs affecting phenological traits in a wild population, suggest that plant tRNA ligases mediate stress-responsive life-history traits, and introduce a novel system for investigating the molecular mechanisms of gene amplification.

opencc-zeroDec 2017View details →
dryad36/100

Somatic copy number and structural variation in RPE-1 cells with induced chromosomal instability

<p><span><span><span><span><span><span><span><span><span><span><span>The chromosome breakage-fusion-bridge (BFB) cycle is a mutational process that produces gene amplification and genome instability. Signatures of BFB cycles can be observed in cancer genomes alongside chromothripsis, another catastrophic mutational phenomenon. Here, we explain this association by elucidating a mutational cascade, downstream of <a>th</a></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>e single cell division error of chromosome bridge formation, that rapidly generates extreme genomic complexity.  We show that actomyosin forces are required for initial bridge breakage and mutagenesis, following which chromothripsis accumulates with aberrant interphase replication of bridge DNA.  This is then followed by an unexpected burst of DNA replication in the next mitosis, generating extensive DNA damage.  During this second cell division, broken bridge chromosomes frequently mis-segregate and form micronuclei, promoting additional chromothripsis. We <a>fu</a></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>rther show that this mutational cascade generates the continuing evolution and sub-clonal heterogeneity characteristic of many human cancers.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroFeb 2020View details →
dryad36/100

Data from: The fire ant social supergene is characterized by extensive gene and transposable element copy number variation

In the fire ant Solenopsis invicta, a supergene composed of ~600 genes and having two variants, SB and Sb, regulates colony social form. In single queen colonies all individuals carry only the SB allele, while in multiple queen colonies, some individuals carry the Sb allele. In this study we characterized genes with copy number variation between SB and Sb-carrying individuals. We showed extensive acquisition of gene duplicates in Sb genome, with some likely involved in polygyne-related phenotypes. We found 260 genes with differences in copy number between SB and Sb, of which 239 are in greater copy number in Sb. We observed TE accumulation on Sb, likely due to the accumulation of repetitive elements on the non-recombining chromosome. We found a weak correlation between TE copy number and differential expression, suggesting some TEs may still be proliferating in Sb, however many of the duplicated TEs were already silenced. Among the 115 non-TE genes with higher copy in Sb, enzymes responsible for cuticular hydrocarbon synthesis were highly represented. These include a desaturase and an elongase; both potentially responsible for differential queen odor and likely beneficial for polygyne ants. These genes seem to have translocated into the supergene from other chromosomes and proliferated by multiple duplication events. While the presence of transposable elements (TEs) in supergenes is well documented, little is known about duplication of non-TE genes and their possible adaptive role. Overall, our results suggest that gene duplications may be an important factor leading to monogyne and polygyne ant societies.

opencc-zeroJan 2020View details →
dryad36/100

Data from: Putative climate adaptation in American pikas (Ochotona princeps) is associated with copy number variation across environmental gradients

<p>Improved understanding of the genetic basis of adaptation to climate change is necessary for maintaining global biodiversity moving forward. Studies to date have largely focused on sequence variation, yet there is growing evidence that suggests that changes in genome structure may be an even more significant source of adaptive potential. The American pika (<em>Ochotona princeps</em>) is an alpine specialist that shows some evidence of adaptation to climate along elevational gradients, but previous work has been limited to single nucleotide polymorphism (SNP)-based analyses within a fraction of the species range. Here, we investigated the role of copy number variation underlying patterns of local adaptation in the American pika using genome-wide data previously collected across the entire species range. We identified 37-193 putative copy number variants (CNVs) associated with environmental variation (temperature, precipitation, solar radiation) within each of the six major American pika lineages, with patterns of divergence largely following elevational and latitudinal gradients. Genes associated (<em>n</em>=158) with independent annotations across lineages, variables, and/or CNVs had functions related to mitochondrial structure/function, immune response, hypoxia, olfaction, and DNA repair, some of which have been previously linked to putative high elevation and/or climate adaptation that may serve as important targets in future studies.</p>

opencc-zeroDec 2023View details →
zenodo36/100

The population genetics of adaptation through copy-number variation in a fungal plant pathogen

<p>Supplementary Tables S1-S8 for the manuscript &quot;The population genetics of adaptation through copy-number variation in a fungal plant pathogen&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Quantitative PCR from human genomic DNA: the determination of gene copy numbers for congenital adrenal hyperplasia and RCCX copy number variation

<p>The dataset is related a study in which we aimed to simultaneously assess the performance of 7 quantitative polymerase chain reaction (qPCR) assays for the gene copy number (GCN) determination of the genetic elements of RCCX copy number variation (CNV). A single laboratory method validations of duplex qPCR assays with hydrolysis probes on <em>CYP21A1P</em> and <em>CYP21A2</em> genes, which are responsible for congenital adrenal hyperplasia, were performed using 46 human genomic DNA samples. We also performed the verifications on 5 qPCR assays for the genetic elements of RCCX CNV such as <em>C4A</em> gene, <em>C4B</em>, gene, RCCX CNV breakpoint, HERV-K(C4) CNV deletion and insertion alleles. The dataset contains the data of genomic DNA samples, the raw quantification cycle values of all qPCR experiments, the peak heights and dosage quotient of multiplex ligation-dependent probe amplification (MLPA) experiments, and the detailed GCN results based on qPCR and MLPA. All other analyses are available in our publication under the same title.</p>

opencc-by-4.0Dec 2021View details →
dryad36/100

Accumulation of gene copy number variations during the early phase of free-spawning abalone speciation

<p><span>The g</span><span>enetic basis of speciation in free-spawning marine invertebrates is poorly understood. Although gene copy number variations (GCNVs) as well as nucleotide variations possibly trigger the speciation of these organisms, empirical evidence for such </span><span>a hypothesis</span><span> is limited. In this study, we searched for genomic signatures of GCNVs that may contribute to the speciation of Western Pacific abalone species. Whole-genome sequencing data suggested the existence of significant amounts of GCNVs in closely related abalones, <em>Haliotis discus</em> and <em>H. madaka</em>, in the early phase of speciation. In addition, the degree of interspecies genetic differentiation in the genes where GCNVs were estimated was higher than </span><span>that </span><span>in other genes, suggesting that nucleotide divergence also accumulate</span><span>s in the genes with GCNVs.</span><span> GCNVs in some genes were also detected in other related abalone species, suggesting that these GCNVs are derived from both ancestral and <em>de novo</em> mutations.</span> <span>Our findings </span><span>suggest that GCNVs have been accumulated in </span><span>the early phase</span><span> of free-spawning abalone speciation.</span></p>

opencc-zeroMay 2024View details →
zenodo36/100

Copy number variation introduced by a massive mobile element facilitates global thermal adaptation in a fungal wheat pathogen - Supplementary Data files

<p>Supplementary Data file 3-4 included in the manuscript Copy number variation introduced by a massive mobile element facilitates global thermal adaptation in a fungal wheat pathogen.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
dryad36/100

Data from: Evolutionary variation in gene conversion at the avian MHC is explained by fluctuating selection, gene copy numbers, and life history

<p>The Major Histocompatibility Complex (MHC) multigene family encodes key pathogen-recognition molecules of the vertebrate adaptive immune system. Hyper-polymorphism of MHC genes is <em>de novo</em> generated by point mutations, but new haplotypes may also arise by re-shuffling of existing variation through intra- and inter-locus gene conversion. Although the occurrence of gene conversion at the MHC has been known for decades, we still have limited understanding of its functional importance. Here, I took advantage of extensive genetic resources (~9000 sequences) to investigate a broad scale macroevolutionary patterns in gene conversion processes at the MHC across nearly 200 avian species. Gene conversion was found to constitute a universal mechanism in birds, as 83% of species showed footprints of gene conversion at either MHC class and 25% of all allelic variants were attributed to gene conversion. Gene conversion processes were stronger at MHC-II than MHC-I, but inter-specific variation at both MHC classes was explained by similar evolutionary scenarios, reflecting fluctuating selection towards different optima and drift. Gene conversion showed uneven phylogenetic distribution across birds and was driven by gene copy number variation, supporting significant role of inter-locus gene conversion processes in the evolution of the avian MHC. Finally, MHC gene conversion was stronger in species with fast life histories (high fecundity) and in long-distance migrants, likely reflecting variation in population sizes and host-pathogen coevolutionary dynamics. The results provide a robust comparative framework for understanding macroevolutionary variation in gene conversion at the avian MHC and reinforce important contribution of this mechanism to functional MHC diversity.</p>

opencc-zeroJun 2024View details →
zenodo36/100

Copy number variation heterogeneity & hierarchical cancer classifications

<p>This is the CNV data from progenetix on the NCIt morphology tree used in the study Copy number variation heterogeneity reveals biological inconsistency in hierarchical cancer classifications. The columns are separated by tabs, and the values indicate the max CNV value of the biosample on the corresponding bin (1MB) of the genome.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Pangenome graph analysis reveals extensive effector copy-number variation in spinach downy mildew

<p>Data produced for the comparison of six&nbsp;<em>Peronospora effusa</em> isolates. For each isolate, we provide the genome assemblies, gene and repeat annotation, effector clustering, and gene variation. Additionally, we provide the repeat library that was used to annotate the transposable elements for each isolate and the pangenome graph.</p> <p>DOI: https://doi.org/10.1101/2024.05.30.596583&nbsp;</p>

opencc-by-4.0Aug 2024View details →
dryad36/100

Amaranthus palmeri EPSPS copy number and glyphosate resistance variation

<p>Gene copy number variation (CNV) has been increasingly associated with organismal responses to environmental stress, but we know little about the quantitative relation between CNV and phenotypic variation. In this study we quantify the relation between variation in <i>EPSPS</i> (5-enolpyruvylshikimate-3-phosphate synthase) copy number using digital drop PCR and variation in phenotypic glyphosate resistance in 22 populations of <i>Amaranthus palmeri</i> (Palmer Amaranth), a range-expanding agricultural weed. Overall, we detected a significant positive relation between population mean copy number and resistance. The majority of populations exhibited high glyphosate resistance yet maintained low-resistance individuals, resulting in bimodality in many populations. We also investigated threshold models for the relation between copy number and resistance, and found evidence for a threshold of ~15 <i>EPSPS</i> copies: there was a steep increase in resistance below the threshold, followed by a much shallower increase. Across 924 individuals, as copy number increases the range of variation in resistance decreases, yielding an increasing frequency of high phenotypic resistance individuals. Among populations <span>we detected a decline in variation (s.d.) as mean phenotypic resistance increased from moderate to high, consistent with the prediction that as phenotypic resistance increases in populations, stabilizing selection decreases variation in the trait. Our study demonstrates that populations of <i>A. palmeri</i> can harbour wide variation in <i>EPSPS</i> copy number and phenotypic glyphosate resistance, reflecting the history of, and template for future, resistance evolution.</span></p>

opencc-zeroSep 2021View details →
zenodo36/100

Copy number variation introduced by a massive mobile element underpins global thermal adaptation in a fungal wheat pathogen - Supplementary Tables..

<p>Supplementary Tables S1 - S11 of the manuscript&nbsp;<strong>Copy number variation introduced by a massive mobile element underpins global thermal adaptation in a fungal wheat pathogen.&nbsp;</strong></p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record