Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
55
datasets available to search
ShareScore release 0.9.0
Dataset results
55 results for “genomic structural variation”
Large structural variations in the haplotype-resolved African cassava genome
<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p> </p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>
Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""
<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. rufipogon</em> DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. glaberrima</em> IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. barthii</em> W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. nivara</em> W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p>
Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture
<p>Anterior cruciate ligament (ACL) rupture is an important condition of the human knee. Second ruptures are common and societal costs are substantial. Canine cranial cruciate ligament (CCL) rupture closely models the human disease. CCL rupture is common in the Labrador Retriever (5.79% prevalence), ~100-fold more prevalent than in humans. Labrador Retriever CCL rupture is a polygenic complex disease, based on genome-wide association study (GWAS) of single nucleotide polymorphism (SNP) markers. Dissection of genetic variation in complex traits can be enhanced by studying structural variation, including copy number variants (CNVs). Dogs are an ideal model for CNV research because of reduced genetic variability within breeds and extensive phenotypic diversity across breeds. We studied the genetic etiology of CCL rupture by association analysis of CNV regions (CNVRs) using 110 case and 164 control Labrador Retrievers. CNVs were called from SNPs using three different programs (PennCNV, CNVPartition, and QuantiSNP). After quality control, CNV calls were combined to create CNVRs using ParseCNV and an association analysis was performed. We found no strong effect CNVRs but found 46 small effect (max(T) permutation P<0.05) CCL rupture associated CNVRs in 22 autosomes; 25 were deletions and 21 were duplications. Of the 46 CCL rupture associated CNVRs, we identified 39 unique regions. Thirty four were identified by a single calling algorithm, 3 were identified by two calling algorithms, and 2 were identified by all three algorithms. For 42 of the associated CNVRs, frequency in the population was <10% while 4 occurred at a frequency in the population ranging from 10-25%. Average CNVR length was 198,872bp and CNVRs covered 0.11 to 0.15% of the genome. All CNVRs were associated with case status. CNVRs did not overlap previous canine CCL rupture risk loci identified by GWAS. Associated CNVRs contained 152 annotated genes; 12 CNVRs did not have genes mapped to CanFam3.1. Using pathway analysis, a cluster of 19 homeobox domain transcript regulator genes was associated with CCL rupture (P=6.6E-13). This gene cluster influences cranial-caudal body pattern formation during embryonic limb development. Clustered genes were found in 3 CNVRs on chromosome 14 (HoxA), 28 (NKX6-2), and 36 (HoxD). When analysis was limited to deletion CNVRs, the association was strengthened (P=8.7E-16). This study suggests a component of the polygenic risk of CCL rupture in Labrador Retrievers is associated with small effect CNVs and may include aspects of stifle morphology regulated by homeobox domain transcript regulator genes.</p>
The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)
<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p> import numpy as np<br> dense = np.load("filename.npz")<br> dense["chr1"]<br> <br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM. </p> <p><strong>Supplemental File 5</strong> Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer <= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>
Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads"
<p>Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads".</p> <p>The archive contains files that are necessary to reproduce the cell line benchmarks from the paper, including:</p> <ul> <li>Scripts and command lines</li> <li>Original VCF outpurs of all tools used in benchmarking</li> <li>Minda evaluations and truthset VCF files</li> <li>Full Severus outputs + visualizations</li> <li>truvari calls</li> </ul>
Structural genomic variation in the inbred Scandinavian wolf population contributes to the realized genetic load but is positively affected by immigration
Open the record for dataset details and reuse information.
Combined analysis of transposable elements and structural variation in maize genomes reveals genome contraction outpaces expansion
Open the record for dataset details and reuse information.
Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture
Open the record for dataset details and reuse information.
Detection and analysis of complex structural variation in human genomes across populations and in brains of donors with psychiatric disorders
Open the record for dataset details and reuse information.
Wide spectrum and high frequency of genomic structural variation, including transposable elements, in large double stranded DNA viruses
Our knowledge of the diversity and frequency of genomic structural variation segregating in populations of large double stranded (ds) DNA viruses is limited. Here we sequenced the genome of a baculovirus (AcMNPV) purified from beet armyworm (Spodoptera exigua) larvae at depths >195,000X using both short-read (Illumina) and long-read (PacBio) technologies. Using a pipeline relying on hierarchical clustering of structural variants (SVs) detected in individual short- and long-reads by six variant callers, we identified a total of 1,141 SVs in AcMNPV, including 464 deletions, 443 inversions, 160 duplications and 74 insertions. These variants are considered robust and unlikely to result from technical artifacts because they were independently detected in at least three long reads as well as at least three short reads. SVs are distributed along the entire AcMNPV genome and may involve large genomic regions (30,496 bp on average). We show that no less than 39.9% of genomes carry at least one SV in AcMNPV populations, that the vast majority of SVs (75%) segregate at very low frequency (<0.01%) and that very few SVs persist after 10 replication cycles, consistent with a negative impact of most SVs on AcMNPV fitness. Using short-read sequencing datasets, we then show that populations of two iridoviruses and one herpesvirus are also full of SVs, as they contain between 426 and 1102 SVs carried by 52.4 to 80.1% of genomes. Finally, AcMNPV long reads allowed us to identify 1,757 transposable elements (TEs) insertions, 895 of which are truncated and occur at one extremity of the reads. This further supports the role of baculoviruses as possible vectors of horizontal transfer of TEs. Altogether, we found that SVs, which evolve mostly under rapid dynamics of gain and loss in viral populations, represent an important feature in the biology of large dsDNA viruses.
Data from: Pan-genome analysis highlights the role of structural variation in the evolution and environmental adaptation of Asian honeybees
<p>The <em>Asian honeybee</em>, <em>Apis cerana</em>, is an ecologically and economically important pollinator. Mapping its genetic variation is key to understanding population-level health, histories, and potential capacities to respond to environmental changes. However, most efforts to date were focused on single nucleotide polymorphisms (SNPs) based on a single reference genome, thereby ignoring larger-scale genomic variation. We employed long-read sequencing technologies to generate a chromosome-scale reference genome for the ancestral group of<em> A. cerana</em>. Integrating this with 525 resequencing datasets, we constructed the first pan-genome of <em>A. cerana</em>, encompassing almost the entire gene content. We found that 31.32% of genes in the pan-genome were variably present across populations, providing a broad gene pool for environmental adaptation. We identified and characterized structural variations (SVs) and found that they were not closely linked with SNP distributions, however, the formation of SVs was closely associated with transposable elements. Furthermore, phylogenetic analysis using SVs revealed a novel <em>A. cerana</em> ecological group not recoverable from the SNP data. Performing environmental association analysis identified a total of 44 SVs likely to be associated with environmental adaptation. Verification and analysis of one of these, a 330 bp deletion in the Atpalpha gene, indicated that this SV may promote the cold adaptation of <em>A. cerana</em> by altering gene expression. Taken together, our study demonstrates the feasibility and utility of applying pan-genome approaches to map and explore genetic feature variations of honeybee populations, and in particular to examine the role of SVs in the evolution and environmental adaptation of <em>A. cerana</em>.</p>
Whole-genome sequencing reveals contribution of rare and common variation to structural kidney and urinary tract malformations
<p>Supplementary tables detailing analysis of whole-genome sequencing data from 992 patients with congenital anomalies of the kidneys and urinary tract (CAKUT). </p>
Resolution of structural variation in diverse mouse genomes reveals chromatin remodeling due to transposable elements
<p>Structural variant calls, RepeatMasker annotations, and genome assemblies of diverse mouse genomes. </p>
Wide spectrum and high frequency of genomic structural variation, including transposable elements, in large double stranded DNA viruses
Open the record for dataset details and reuse information.
Data from: Pan-genome analysis highlights the role of structural variation in the evolution and environmental adaptation of Asian honeybees
Open the record for dataset details and reuse information.
Structural genomic variation and behavioral interactions underpin a balanced sexual mimicry polymorphism
Open the record for dataset details and reuse information.
Evolutionary genomics of structural variation in Asian rice (Oryza sativa) domestication
<p>DATA from Kou et al. 2020 Evolutionary Genomics of Structural Variation in Asian Rice (<em>Oryza sativa</em>) Domestication, <em>Molecular Biology and Evolution</em>, Volume 37, Issue 12, December 2020, Pages 3507–3524, <a href="https://doi.org/10.1093/molbev/msaa185">https://doi.org/10.1093/molbev/msaa185</a> </p> <p>Gene and TE annotation of Nipponbare (<em>Oryza sativa </em>ssp.<em> japonica</em> ) genome V5.0 updated using SMRT long reads</p> <p>Raw SV calls jointly detected in Asian rice (<em>Oryza sativa</em>) and its wild progenitor (<em>O. rufipogon</em>)</p> <p> </p>
Duck pan-genome reveals two transposon-derived structural variations caused bodyweight enlarging and white plumage phenotype formation during evolution
<p><span>Structural variations (SVs) are a major source of domestication and improvement traits. We present the first duck pan-genome constructed using five genome assemblies capturing ~40.98 Mb new sequences. This pan-genome together with high-depth sequencing data (>46.5X) identified 101,041 SVs, of which substantial proportions were derived from transposable element (TE) activity. Many TE-derived SVs anchored in a gene body or regulatory region are linked to domestication and improvement. By combining quantitative genetics with molecular experiments, we dissect how TE-derived SVs change gene expression of <em>IGF2BP1</em> and generate novel transcripts of <em>MITF</em>, shaping body weight and plumage color. In the <em>IGF2BP1</em> locus, the TE-derived SV explains the largest effect on body weight among avian species (27.61% of phenotypic variation). Our findings highlight the </span><span>importance of using a pan-genome as a reference in genomics studies</span><span> and explore the roles of TE-derived SVs in trait formation and in livestock breeding.</span></p>
Dataset from: Distinct patterns of genetic variation at low-recombining genomic regions represent haplotype structure
<p>Genetic variation of the entire genome represents population structure, yet individual loci can show distinct patterns. Such deviations identified through genome scans have often been attributed to effects of selection instead of randomness. This interpretation assumes that long enough genomic intervals average out randomness in underlying genealogies, which represent local genetic ancestries. However, an alternative explanation to distinct patterns has not been fully addressed: too few genealogies to average out the effect of randomness. Specifically, distinct patterns of genetic variation may be due to reduced local recombination rate, which<br>reduces the number of genealogies in a genomic window. Here, we associate distinct patterns of local genetic variation with reduced recombination rates in a songbird, the Eurasian blackcap (<em>Sylvia atricapilla</em>), using genome sequences and recombination maps. We find that distinct patterns of local genetic variation reflect haplotype structure at low-recombining regions either shared in most populations or found only in a few populations. At the former species-wide low-recombining regions, genetic variation depicts conspicuous haplotypes segregating in multiple populations. At the latter population-specific low-recombining regions, genetic variation represents variance among cryptic haplotypes within the low-recombining populations. With simulations, we confirm that these distinct patterns of haplotype structure evolve due<br>to reduced recombination rate, on which the effects of selection can be overlaid. Our results highlight that distinct patterns of genetic variation can emerge through evolution of reduced local recombination rate. Recombination landscape as an evolvable trait therefore plays an important role determining the heterogeneous distribution of genetic variation along the genome.</p>
The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications
<div> <div> <p><a href="https://onlinelibrary.wiley.com/doi/10.1111/age.13181">https://onlinelibrary.wiley.com/doi/10.1111/age.13181</a></p> <h1>The de novo assembly of a European wild boar genome revealed unique patterns of chromosomal structural variations and segmental duplications</h1> <div> </div> <div> <div> <div> <div><a href="https://onlinelibrary.wiley.com/authored-by/Chen/Jianhai">Jianhai Chen</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Zhong/Jie">Jie Zhong</a>, <a href="https://onlinelibrary.wiley.com/authored-by/He/Xuefei">Xuefei He</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Li/Xiaoyu">Xiaoyu Li</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Ni/Pan">Pan Ni</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Safner/Toni">Toni Safner</a>, <a href="https://onlinelibrary.wiley.com/authored-by/%C5%A0prem/Nikica">Nikica Šprem</a>, <a href="https://onlinelibrary.wiley.com/authored-by/Han/Jianlin">Jianlin Han</a></div> </div> </div> </div> <p>The rapid progress of sequencing technology has greatly facilitated the de novo genome assembly of pig breeds. However, the assembly of the wild boar genome is still lacking, hampering our understanding of chromosomal and genomic evolution during domestication from wild boars into domestic pigs. Here, we sequenced and de novo assembled a European wild boar genome (ASM2165605v1) using the long-range information provided by 10× Linked-Reads sequencing. We achieved a high-quality assembly with contig N50 of 26.09 Mb. Additionally, 1.64% of the contigs (222) with lengths from 107.65 kb to 75.36 Mb covered 90.3% of the total genome size of ASM2165605v1 (~2.5 Gb). Mapping analysis revealed that the contigs can fill 24.73% (93/376) of the gaps present in the orthologous regions of the updated pig reference genome (Sscrofa11.1). We further improved the contigs into chromosome level with a reference-assistant scaffolding method. Using the ‘assembly-to-assembly’ approach, we identified intra-chromosomal large structural variations (SVs, length >1 kb) between ASM2165605v1 and Sscrofa11.1 assemblies. Interestingly, we found that the number of SV events on the X chromosome deviated significantly from the linear models fitting autosomes (<em>R</em><sup>2</sup> > 0.64, <em>p</em> < 0.001). Specifically, deletions and insertions were deficient on the X chromosome by 66.14 and 58.41% respectively, whereas duplications and inversions were excessive on the X chromosome by 71.96 and 107.61% respectively. We further used the large segmental duplications (SDs, >1 kb) events as a proxy to understand the large-scale inter-chromosomal evolution, by resolving parental-derived relationships for SD pairs. We revealed a significant excess of SD movements from the X chromosome to autosomes (<em>p</em> < 0.001), consistent with the expectation of meiotic sex chromosome inactivation. Enrichment analyses indicated that the genes within derived SD copies on autosomes were significantly related to biological processes involving nervous system, lipid biosynthesis and sperm motility (<em>p</em> < 0.01). Together, our analyses of the de novo assembly of ASM2165605v1 provides insight into the SVs between European wild boar and domestic pig, in addition to the ongoing process of meiotic sex chromosome inactivation in driving inter-chromosomal interaction between the sex chromosome and autosomes.</p> </div> </div> <div>The work has been pulished here: https://onlinelibrary.wiley.com/doi/full/10.1111/age.13181</div> <div> </div> <div>The current dataset include the genome annotation files.</div> <div> </div> <div>For the whole-genomic assembly, please check NCBI: </div> <div>https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_021656055.1/</div> <div> <table> <tbody> <tr> <th> </th> <th>GenBank</th> </tr> </tbody> <tbody> <tr> <td>Genome size</td> <td>2.5 Gb</td> </tr> <tr> <td>Total ungapped length</td> <td>2.4 Gb</td> </tr> <tr> <td>Number of scaffolds</td> <td>12,642</td> </tr> <tr> <td>Scaffold N50</td> <td>28.3 Mb</td> </tr> <tr> <td>Scaffold L50</td> <td>25</td> </tr> <tr> <td>Number of contigs</td> <td>41,323</td> </tr> <tr> <td>Contig N50</td> <td>157.9 kb</td> </tr> <tr> <td>Contig L50</td> <td>4,562</td> </tr> <tr> <td>GC percent</td> <td>42</td> </tr> <tr> <td>Genome coverage</td> <td>56.0x</td> </tr> <tr> <td>Assembly level</td> <td>Scaffold</td> </tr> </tbody> </table> <p> </p> <h2>Assembly methods</h2> <div>Sequencing technology 10xgenomics Assembly method Supernova v. 2.1.1 <p> </p> <p>part_** are genome fasta for the GCA_021656055.1</p> <p>You could use the following to combine and uncompress.</p> </div> </div> <div> <div> <div><code><span>cat</span> part_* > archive_combined.zip </code></div> </div> <div> <div> </div> <div><code>unzip archive_combined.zip</code></div> </div> </div> <div> </div> <div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.