Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
129
datasets available to search
ShareScore release 0.9.0
Dataset results
129 results for “Structural variant”
Structural variant discovery and genotyping in next-generation sequencing data
<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels >50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project
SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.
CADD-SV - A framework to score the effects of structural variants in health and disease
<p>Required annotation data-set to run the CADD-SV framework; a method to retrieve and integrate a wide set of annotations to predict the effects of SVs. Pre-scored variants as well as additional information on used features.<br> A webserver for online scoring as well as data downloads is available at: https://cadd-sv.bihealth.org/<br> Source code for CADD-SV is available at GitHub: https://github.com/kircherlab/CADD-SV</p>
Phenopackets for case reports of structural variants
<p>A collection of 188 published deleterious structural variants based on 182 cases published in 146 clinical case reports that describe individuals with Mendelian diseases.</p>
Dataset and structure database for an ML model to predict diffusivity in ZIF variants
<p>This dataset accompanies the publication titled "Data Mining for Predicting Gas Diffusivity in Zeolitic-imidazolate Frameworks (ZIFs)" (DOI: <a href="https://doi.org/10.1039/D2TA02624D">https://doi.org/10.1039/D2TA02624D</a>)</p> <p><a href="https://zenodo.org/api/files/b80f6d07-3bf4-484c-97ac-5d579fb0cc27/ESI_2_dataset.xlsx?versionId=dc4525d0-1c5c-478c-9bef-a156587ad69b">ESI_2_dataset.xlsx</a>: Descriptors for all ZIFs of the publication and simulations output, in the form of diffusivities of gas molecules (He up to iso-butane), in all ZIFs.</p> <p>ZIF_database.zip: ZIP file containing all ZIFs prepared by the authors (as discussed in the publication), through various units replacements, in the SOD topology, in .pdb format.</p>
i2QTL HipSci Structural Variant and Short Tandem Repeat Genotypes
<p>Here we provide structural variant and short tandem repeat variant calls from 204 HipSci donors as described in the manuscript Jakubosky et al. "Discovery and quality analysis of a comprehensive set of structural variants and short tandem repeats". </p> <p><a href="https://www.biorxiv.org/content/10.1101/713198v2">https://www.biorxiv.org/content/10.1101/713198v2</a></p> <p>Jakubosky D, Smith EN, D’Antonio M, Bonder MJ, Young Greenwald WW, Matsui H, D’Antonio-Chronowska A, (Hipsci), Stegle O, Montgomery SB, DeBoever C, Frazer KA. Discovery and Quality Analysis of a Comprehensive Set of Structural Variants and Short Tandem Repeats. bioRxiv. January 2019:713198. doi:10.1101/713198.</p> <p> </p>
Native MS dataset for: "Insights into the pathogenesis of primary hyperoxaluria type I from the structural dynamics of alanine:glyoxylate aminotransferase variants"
<p>Native mass spectrometry dataset used in: <strong>Insights into the pathogenesis of primary hyperoxaluria type I from the structural dynamics of alanine:glyoxylate aminotransferase variants.</strong> Pavla Vankova, Juan Luis Pacheco-Garcia, Dmitry S. Loginov, Atanasio Gómez-Mulas, Alan Kádek, José Manuel Martín-Garcia, Eduardo Salido, Petr Man and Angel L. Pey. FEBS Letters (2024)</p> <p><strong>Description:</strong></p> <p>Native mass spectrometry (MS) analysis verifying the oligomeric state of alanine:glyoxylate aminotransferase (AGT) protein and its P11L and I340M (LM) polymorphism and LM G170R mutation variants in primary hyperoxaluria type I.</p> <p><strong>Sample processing:</strong></p> <p>AGT protein as well as its LM and LM G130R mutants were buffer exchanged into 150 mM aqueous ammonium acetate solution (pH 7.5, MS-grade, Sigma-Aldrich) through six cycles of tenfold dilution and re-concentration using centrifugal concentrators Vivaspin 500 (30 kDa cut-off, <em>Sartorius</em>). Desalted proteins were introduced into a Synapt G2Si mass spectrometer (Waters) via static nanoelectrospray ionization from in-house prepared gold-coated borosilicate glass capillaries Kwik-Fil 1B120F-4 (<em>World Precision Instruments</em>). Protein concentration in samples was determined by 280 nm absorbance measurements using DeNovix DS-11 spectrophotometer. Samples were diluted in ammonium acetate and electrosprayed at 1 and 2 µM concentration. The mass spectrometer was carefully tuned for best signal quality and intensity, while keeping ion activation and unfolding minimal. Namely, electrospray voltage was kept at 1.3 kV, source desolvation temperature 80°C, sampling cone 80 V and 10 V collision voltage with 6 ml/min flow of argon in the trap region for thermalization of ions. Quadrupole was operated in a broadband transmission mode up to 8000 m/z while the spectra were acquired in mass range 500 – 20000 m/z. Spectra were externally mass recalibrated using known masses of caesium iodide clusters.</p> <p><strong>Data processing:</strong></p> <p>Raw mass spectra were averaged over 75 scans and further processed in Waters MassLynx 4.1. The averaged spectra were exported for ZENODO deposition as plain in plain m/z vs intensity .txt files as well uploaded as part of the .raw file format of the whole analysis (including initial metadata) with scan descriptions and parameter changes described in a stand-alone .txt descriptor file.</p>
Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions
<p>Using AlphaFold2 and HADDOCK, we have generated a predicted structure for the SARS-CoV-2 B.1.1.529 variant's Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in the receptor binding domain of B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>
Supplementary material for: "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"
<p>Supplementary material for the publication "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"</p> <p>The realated preprint can be found at Research Square (<a href="https://doi.org/10.21203/rs.3.rs-861830/v1">https://doi.org/10.21203/rs.3.rs-861830/v1</a>)</p> <p>Supplementary file 1: Supplementary results, tables and figures.</p> <p>Supplementary file 2: MultiQC report.</p> <p>Supplementary file 3: Observer concordance of the visual filtering step.</p> <p>Supplementary file 4: Snakemake workflow and scripts.</p> <p> </p>
Local adaptation and archaic introgression shape global diversity at human structural variant loci
<p>Supporting data associated with the manuscript "Local adaptation and archaic introgression shape global diversity at human structural variant loci". These include:</p> <ul> <li>structural variant genotypes (Paragraph; <a href="https://github.com/Illumina/paragraph">https://github.com/Illumina/paragraph</a>)</li> <li>eQTL mapping results (fastqtl permutation pass; see <a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a> for column descriptions)</li> <li>eQTL fine-mapping results (CAVIAR; see <a href="http://genetics.cs.ucla.edu/caviar/index.html">http://genetics.cs.ucla.edu/caviar/index.html</a>)</li> <li>structural variant selection scan results (Ohana; <a href="https://github.com/jade-cheng/ohana">https://github.com/jade-cheng/ohana</a>)</li> </ul> <p>Description of files in this directory:</p> <p><strong>Structural variant genotypes</strong></p> <p><code>SVs_paragraphFormat.vcf.gz</code> - merged long-read structural variant calls</p> <p><code>SVs_1KGP_pgGTs.vcf.gz</code> - genotypes for 1000 Genomes samples in VCF format</p> <p><strong>eQTL mapping results</strong></p> <p><code>fastqtl_out.txt</code> - results from fastQTL permutation pass; see <a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a> for column descriptions</p> <p><code>caviar_out.txt</code> - results from fine-mapping SNPs and SVs at significant SV eQTL loci with CAVIAR. Description of columns:</p> <ul> <li>query_sv: SV that was a significant eQTL and underwent fine-mapping</li> <li>gene_id: gene exhibiting an expression association with the query_sv</li> <li>var_id: variant (SNV or SV) that was tested for expression association with the above gene in the fine-mapping analysis</li> <li>var_in_credible_causal_set: Boolean variable denoting whether the above variant is in the 95% credible causal set</li> <li>prob_in_pcausal_set: the amount that this variant contributes to 95% credible causal set</li> <li>causal_post_prob: the posterior probability that the variant is causal in the expression association</li> </ul> <p><strong>Structural variant selection scan results</strong></p> <p><code>chr21_pruned_50_Q.matrix</code> - admixture proportion matrix (generated by Ohana; <a href="https://github.com/jade-cheng/ohana">https://github.com/jade-cheng/ohana</a>)</p> <p><code>chr21_pruned_50_F.matrix</code> - matrix of inferred ancestral allele frequencies (generated by Ohana)</p> <p><code>chr21_pruned_50_C.matrix</code> - matrix of ancestry component covariances (generated by Ohana) Entries of the matrix can be modified to produce "selection hypothesis" matrices where allele frequencies are allowed to vary in one ancestry component (<a href="https://github.com/jade-cheng/ohana/wiki/Population-or-ancestry-specific-selection-scan">https://github.com/jade-cheng/ohana/wiki/Population-or-ancestry-specific-selection-scan</a>).</p> <p><code>selscan_50_k8_p*.txt.gz</code> - raw output of Ohana selscan (see <a href="https://github.com/jade-cheng/ohana">https://github.com/jade-cheng/ohana</a>)</p> <p><code>selscan_res.txt.gz</code> - Ohana selection scan results. These results have been filtered to exclude SVs that have low genotyping rates (<50% of samples), violate Hardy-Weinberg equilibrium expectations (excess of heterozygotes) in more than half of populations, or have extreme global log likelihood estimate (LLE) values. Description of columns:</p> <ul> <li>ID: SV ID</li> <li>#CHROM: SV chromosome</li> <li>POS: SV start position</li> <li>SVLEN: SV length (negative for deletions)</li> <li>step: number of steps needed to interpolate between genome-wide and selection hypothesis models</li> <li>lle_ratio: likelihood ratio statistic (LRS) of the genome-wide vs. selection hypothesis model</li> <li>global-lle: log likelihood of the genome-wide model</li> <li>local-lle: log likelihood of the selection hypothesis model</li> <li>f-pop0: inferred allele frequency in ancestry component 0</li> <li>f-pop1: inferred allele frequency in ancestry component 1</li> <li>f-pop2: inferred allele frequency in ancestry component 2</li> <li>f-pop3: inferred allele frequency in ancestry component 3</li> <li>f-pop4: inferred allele frequency in ancestry component 4</li> <li>f-pop5: inferred allele frequency in ancestry component 5</li> <li>f-pop6: inferred allele frequency in ancestry component 6</li> <li>f-pop7: inferred allele frequency in ancestry component 7</li> <li>ancestry_component: ancestry component tested by the selection hypothesis model. Note that we have added 1 to the ancestry component numbers to match the terminology used in paper (which orders the components from 1-8 rather than 0-7 for interpretability)</li> <li>snp_perc: SV's percentile in the LRS distribution for frequency-matched SNPs</li> <li>p_nominal: nominal p-value calculated from the likelihood ratio</li> <li>p_adj: adjusted p-value calculated from the likelihood ratio</li> </ul> <p> </p>
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 5. Варианты преΑсказанной Αоменной структуры скавенΑжер-рецепторов гемоцитов моΛΛюсков Planorbarius corneus. Сокращения (зΑесь и ΑаΛее): SR — богатый цистеином Αомен скавенΑжер-рецептора, Filament — Αомен промежуточного фиΛамента, TSP1 — повторы тромбоспонΑина типа 1, KR — крингΛ-Αомен, LDLa — Αомен рецептора Λипопротеинов низкой пΛотности кΛасса А Fig. 5. Variants of the predicted domain structure of scavenger receptors from hemocytes of Planorbarius corneus molluscs. Abbreviations (here and in what follows): SR — scavenger receptor Cys-rich domain, Filament — intermediate filament protein, TSP1 — thrombospondin type 1 repeats, KR — kringle domain, LDLa — low-density lipoprotein receptor domain class A
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain
Structural variants of calreticulin mutants associated with essential thrombocythemia
<p>Video S1. CALRwt molecular simulation with Ca2+ ions binding to the structure. This movie shows the dynamics of Ca2+ ions over the 40 ns and how they rapidly bind to the CALRwt structure. Most of the ions bind to the C-terminal part of CALRwt, where the majority of the negative residues are located. The binding of Ca2+ ions creates an interaction between residues that can induce specific folding. In this example, residues E407 and E416 interact with a calcium ion and fold, as shown at the end of the movie.</p> <p>Video S2. CALRwt molecular simulation with calcium ions binding to the structure. This movie shows the dynamics of CALRwt and Ca2+ ions over 400 ns and how ions are rapidly bound to the CALRwt structure. Most of the ions bind to the C-terminal region of CALRwt, where the majority of the negative residues are located. The binding of calcium ions creates an interaction between residues that can induce specific folding (a specific example of fold is shown on Video S1). On this full movie, we are able to see the unfolding of the N-terminal region of the helix and the stabilization of the C- terminal region thanks to calcium binding.</p> <p>Video S3. CALRwt molecular simulation with Na+ ions. This movie shows the dynamics of CALRwt with Na+ ions. The absence of binding from calcium ions leaves the structure free of any external constraint, especially for the C-terminal region. The latter appears to be flexible but this region is in fact quite stable at the local level and has its own dynamics, unstructured slightly the C-terminal of the helix.</p> <p>Video S4. CALRm class A molecular simulation. This movie shows the dynamics of CALRm class A over 400ns. The C-terminal region of the structure is quite flexible at first and interacts with the N-terminal region after several ns. This interaction locally constrains the structure around residues 400-404, stabilizing an unstructured structure for this region.</p> <p>Video S5. CALRm class B molecular simulation. This movie shows the dynamics of CALRm class B over 400ns. The first and last helices move away from their initial position with great flexibility of the coiled regions between the helices. The first helix interacts closely with the second helix and forms a specific T-shaped fold which stabilize the whole structure.</p> <p>Video S6. CALRm class C molecular simulation. This movie shows the dynamics of CALRm class C over 400ns. Extremities are highly flexible but the helix remains stable. This high flexibility allow the C- terminal region to have some interaction with the helix at some frames</p> <p>Video S7. CALRm class D molecular simulation. This movie shows the dynamics of CALRm class D over 400ns. The C-terminal region appears to be flexible, but the helix remains stable the whole simulation.</p> <p>Video S8. CALRm class E molecular simulation. This movie shows the dynamics of CALRm class E over 400ns. This movie shows the dynamics of CALRm class E and calcium ions over 400 ns and how ions are rapidly bound to the CALRwt structure. Most of the ions bind to the C-terminal region of CALRm class E, where the majority of the negative residues are located. The N-terminal region of the helix is being unstructured and the C-terminal region is stabilized with calcium ions. This simulation of CALRm class E is very similar to the simulation of CALRwt.</p> <p>Video S9. Dimeric form of CALRm class A molecular simulation. This movie shows the dynamics of two CALRm class A monomers forming a dimer through disulphide bonds. Both chains seem to repulse each other due to electrostatic charges, but disulphide bonds maintain the dimeric form, otherwise both chains would have been separated.</p> <p>Video S10. Dimeric form of CALRm class A with broken disulphide bonds molecular simulation. This movie shows the dynamics of two CALRm class A monomers and their attempt to form a dimer with broken disulfide bonds. As each monomer repels each other, the dimeric form cannot be stable without any disulfide bonds. This is demonstrated by the separation of each chain from each other.</p> <p>Video S11. Dimeric form of CALRm class B molecular simulation. This movie shows the dynamics of two CALRm class B monomers forming a dimer through disulphide bonds. Both chains are interacting together to form a specific shape, similar to the simulation of the monomer of class B. This interaction implies that even without any disulphide bonds, the dimeric form of class B could be stable, contrary to class A.</p> <p>Video S12. Dimeric form of CALRm class B molecular simulation. It is an interesting replicate of the same system than Video S11. Helices are also well maintained.</p> <p>Video S13. Dimeric form of CALRm class C molecular simulation. This movie shows the dynamics of two CALRm class C monomers forming a dimer through disulphide bonds. Both chains seem to repulse each other due to electrostatic charges, but disulphide bonds maintain the dimeric form, otherwise both chains would have been separated.</p> <p>Video S14. Dimeric form of CALRm class E molecular simulation. This movie shows the dynamics of two CALRm class E monomers and their attempt to form a dimer without any disulphide bonds. Chains repel and are moving away from each other after several ns, indicating the inability for class E to form a dimer.</p> <p>Video S15. Dimeric form of CALRwt molecular simulation. This movie shows the dynamics of two CALRwt monomers and their attempt to form a dimer. Some interactions occur between the N- terminus of each chain, but this is not sufficient and the chains move away from each other. Subsequently, the dynamics of each chain resembles the dynamics of the CALRwt monomer simulated with sodium ions (Video S2). CALRwt is not able to be stable as dimer.</p> <p>Video S16. Dimeric form of CALRm class D molecular simulation. This movie shows the dynamics of two CALRm class D monomers and their attempt to form a dimer without any disulphide bonds. Both chains are separated very quickly which indicate that they can not be stable as dimer.</p>
Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants
Open the record for dataset details and reuse information.
MD simulations from "#GotGlycans: Role of N343 Glycosylation on the SARS-CoV-2 S RBD Structure and Co-Receptor Binding Across Variants of Concern
<p>This folder contains all the MD simulations (saved in frames of 1 ns in PDB format) analysed and discussed in the paper titled "#GotGlycans: Role of N343 Glycosylation on the SARS-CoV-2 S RBD Structure and Co-Receptor Binding Across Variants of Concern" DOI https://doi.org/10.1101/2023.12.05.570076. The naming reflects the specific variant and the presence ('g' or 'gly') or absence ('ng' or 'nogly') of glycosylation at N343 and N331 sites in the SARS-CoV-2 S RBD. Gaussian accelerated MD simulations are indicated with 'gamd', all others represent conventional (deteriministic) sampling. For all details please refer to the original manuscript.</p>
Calling structural variants with confidence from short-read data in wild bird populations
<p>Comprehensive characterisation of structural variation in natural populations has only become feasible in the last decade. To investigate the population genomic nature of structural variation (SV), reproducible and high-confidence SV callsets are first required. We created a population-scale reference of the genome-wide landscape of structural variation across 33 Nordic house sparrows (<em>Passer domesticus</em>) individuals. To produce a consensus callset across all samples using short-read data, we compare heuristic-based quality filtering and visual curation (Samplot/PlotCritic and Samplot-ML) approaches. We demonstrate that curation of SVs is important for reducing putative false positives and that the time invested in this step outweighs the potential costs of analysing short-read discovered SV datasets that include many potential false positives. We find that even a lenient manual curation strategy (e.g. applied by a single curator) can reduce the proportion of putative false positives by up to 80%, thus enriching the proportion of high-confidence variants. Crucially, in applying a lenient manual curation strategy with a single curator, nearly all (>99%) variants rejected as putative false positives were also classified as such by a more stringent curation strategy using three additional curators. Furthermore, variants rejected by manual curation failed to reflect the expected population structure from SNPs, whereas variants passing curation did. Combining heuristic-based quality-filtering with rapid manual curation of structural variants in short-read data can therefore become a time- and cost-effective first step for functional and population genomic studies requiring high-confidence SV callsets.</p>
A phased genome of the highly heterozygous 'Texas' almond uncovers patterns of allele-specific expression linked to heterozygous structural variants
<h2># Genomic datasets associated to the publication: </h2> <h3># Gene-ID conversion with previous genome version</h3> <p>Texasv3_vs_Texasv2_GeneID.txt -- gene ID conversion between Texasv3 and Texasv2 (https://www.rosaceae.org/analysis/295)</p> <p>pdulcis26_to_F1_liftoff_polished.gff3 -- Texasv2 gene annotation liftoff on Phase-1 assembly (Phase-1 coordinates)</p> <h3># Phase-1</h3> <p>Texas_F1_K80_chr.fasta -- genome asssembly, phase-1 <br>Texas_F1_gene_models.gff3 -- phase-1 gene annotation (de novo annotation)<br>Functional_annotation_TexasF1.csv -- phase-1 gene functions <br>Texas_F1_ref_SV.vcf --- Structural variations relative to phase-0 (This file uses Phase-1 as reference)</p> <h3># Phase-0</h3> <p>Texas_F0_K80_chr.fasta -- genome asssembly, phase-0 <br>Texas_F0_gene_models.gff3 -- phase-0 gene annotation (liftoff from Phase-1)<br>Functional_annotation_TexasF0.csv -- phase-0 gene functions <br>Texas_F0_ref_SV.vcf --- Structural variations relative to phase-1 (This file uses Phase-0 as reference)</p> <h3># Transposable element annotation</h3> <p>Texas_F0_HiConf_TE_v3.gff3 --- TE annotation in Phase-0<br>Texas_F1_HiConf_TE_v3.gff3 --- TE annotation in Phase-1<br>Texasv3_TElib.fa --- TE library of TexasV3 (non-redundant repeat consensuses taking the account the two genome phases)</p> <h3># Gene sequences in fasta</h3> <p>Transcript, CDS and protein sequences in fasta format for phase-0 (F0) and phase-1 (F1) </p>
Structural variants in the barley gene pool: precision and sensitivity to detect them using short-read sequencing and their association with gene expression and phenotypic variation
<p>SNV of 23 parental barley inbreds of the double round robin population (DRR) (<a href="https://doi.org/10.1111/pbi.13746">https://doi.org/10.1111/pbi.13746</a>) used in the publication "Structural variants in the barley gene pool: precision and sensitivity to detect them using short-read sequencing and their association with gene expression and phenotypic variation". SV, INDELs, and additional data are available via figshare (https://doi.org/10.6084/m9.figshare.16802473).</p>
The role of structural variants in pest adaptation and genome evolution of the Colorado potato beetle, Leptinotarsa decemlineata (Say)
<p>Structural variation has been associated with genetic diversity and adaptation in diverse taxa. Despite these observations, it is not yet clear what their relative importance is for microevolution, especially with respect to known drivers of diversity, e.g., nucleotide substitutions, in rapidly adapting species. Here we examine the significance of structural variants (SVs) in pesticide resistance evolution of the agricultural super-pest, the Colorado potato beetle,<em> Leptinotarsa decemlineata</em>. By employing a parent offspring trio sequencing procedure, we develop highly contiguous reference genomes to characterize structural variation within this species. These updated assemblies represent >100-fold improvement of contiguity and include derived pest and ancestral non-pest individuals. We identify >200,000 SVs, which appear to be non-randomly distributed across the genome as they co-occur with transposable elements and genes. SVs intersect exons for a large proportion of gene annotations (~20%) and are associated with insecticide resistance, development, and transcription, most notably cytochrome P450 (CYP) genes. To understand the role that SVs might play in adaptation we measure allele frequencies of SVs for an additional 57 individuals, using whole genome resequencing data, representing pest and non-pest populations of North America. Incorporating multiple independent tests of significance using SNP data, we identify 14<strong> </strong>positively selected genes that include SVs and SNPs of elevated frequency within the sampled pest lineages. Among these, four are associated with insecticide resistance. One of these genes, glycosyltransferase-13, is a duplicated gene enclosed within a structural variant that resides inside the <em>CYP4g15</em> genic region. Both gene products have been observed to be co-induced during insecticide exposure. These results demonstrate the significance of structural variations as a genomic feature to describe species history, genetic diversity, and adaptation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.