Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

244

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

244 results for “genomic variants”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection

Root and butt rot caused by members of the Heterobasidion annosum species complex is the most economically important disease of conifer trees in boreal forests. Wood decay in the infected trees dramatically decreases their value and causes considerable losses to forest owners. Trees vary in their susceptibility to Heterobasidion infection, but the genetic determinants underlying the variation in the susceptibility are not well-understood. We performed the identification of Norway spruce genes associated with the resistance to Heterobasidion parviporum infection using genome-wide exon-capture approach. Sixty-four clonal Norway spruce lines were phenotyped, and their responses to H. parviporum inoculation were determined by lesion length measurements. Afterwards, the spruce lines were genotyped by targeted resequencing and identification of genetic variants (SNPs). Genome-wide association analysis identified 10 SNPs located within 8 genes as significantly associated with the larger necrotic lesions in response to H. parviporum inoculation. The genetic variants identified in our analysis are potential marker candidates for future screening programs aiming at the differentiation of disease-susceptible and resistant trees.

opencc-zeroDec 2017View details →
dryad32/100

Data from: The role of structural genomic variants in population differentiation and ecotype formation in Timema cristinae walking sticks

Theory predicts that structural genomic variants such as inversions can promote adaptive diversification and speciation. Despite increasing empirical evidence that adaptive divergence can be triggered by one or a few large inversions, the degree to which widespread genomic regions under divergent selection are associated with structural variants remains unclear. Here we test for an association between structural variants and genomic regions that underlie parallel host-plant associated ecotype formation in Timema cristinae stick insects. Using mate-pair re-sequencing of 20 new whole genomes we find that modest-sized structural variants such as inversions, deletions, and duplications are widespread across the genome, being retained as standing variation within and among populations. Using 160 previously published, standard-orientation whole genome sequences we find little to no evidence that the DNA sequences within inversions exhibit accentuated differentiation between ecotypes. In contrast, a formerly described large region of reduced recombination that harbors genes controlling color-pattern exhibits evidence for accentuated differentiation between ecotypes, which is consistent with differences in the frequency of color-pattern morphs between host-associated ecotypes. Our results suggest that some types of structural variants (e.g., large inversions) are more likely to underlie adaptive divergence than others, and that structural variants are not required for subtle yet genome-wide genetic differentiation with gene flow.

opencc-zeroDec 2019View details →
zenodo32/100

Dataset for "NanoVar: a Comprehensive Workflow for Structural Variant Detection to uncover the Genome's Hidden Patterns"

<h2><strong>Output Files for Long-Read Structural Variant and Repeat Analysis in Colorectal Cancer Samples (HRR698464, HRR698460, C586, C588)</strong></h2> <h3>Description:</h3> <p>This Zenodo dataset includes comprehensive output files generated during the application of a long-read sequencing analysis protocol for structural variant (SV) detection and repeat element characterization in colorectal cancer samples. The dataset is organized into two main directories:</p> <p><strong>1. HRR698464_MSI-H_Tumor</strong><br>This directory contains all primary output files generated from the analysis pipeline applied to the MSI-H tumor sample HRR698464 (also referred to as patient C586.T). Each subdirectory corresponds to a specific stage in the protocol:</p> <ul> <li>NanoPlot_output<br>Output from Stage 1 &ndash; Quality assessment of raw reads using NanoPlot.</li> <li>SAMtools_output<br>BAM file processing outputs from Stage 2 &ndash; Alignment of long reads to the reference genome using SAMtools.</li> <li>NanoVar_output<br>Output from Stage 3 &ndash; Structural variant calling using NanoVar.</li> <li>VCF_filtering_output<br>Output from Stage 4 &ndash; Filtering of structural variants using SURVIVOR and BCFtools; includes the filtered VCF files.</li> <li>NanoINSight_output<br>Output from Stage 5 &ndash; Characterization of repeat elements using NanoINSight.</li> <li>VEP_output<br>Output from Stage 6 &ndash; Annotation of structural variants using Ensembl Variant Effect Predictor (VEP).</li> </ul> <p>&nbsp;</p> <p><strong>2. Additional_output_files</strong><br>This directory contains supplementary output files used for comparison and visualization in Figures 4&ndash;7 of the associated publication. These include:</p> <ul> <li>HRR698460.NanoPlot.report.html<br>NanoPlot quality summary of a lower-quality tumor sample (HRR698460), used in Figure 4 for comparison with HRR698464.</li> <li>C586.N.nanovar.pass.vcf<br>NanoVar VCF output for the matched normal sample of patient C586, used to filter somatic calls in Stage 4.</li> <li>C586.N.nanovar.pass.report.html<br>NanoVar summary report of the normal sample of C586; used in Figure 5a.</li> <li>C588.N.nanovar.pass.vcf<br>NanoVar VCF output of the MSS normal sample (C588) for comparison with the MSI-H patient (C586).</li> <li>C588.N.nanovar.pass.report.html<br>NanoVar summary report of the MSS normal sample; used in Figure 5b.</li> <li>C588.T.nanovar.pass.vcf<br>NanoVar VCF output of the MSS tumor sample (C588); used in comparative analyses with the MSI-H sample.</li> <li>C588.T.nanovar.pass.report.html<br>NanoVar summary report of the MSS tumor sample; used in Figure 5b.</li> <li>MSS.tumor.unique.vcf<br>VCF file of somatic SVs in the MSS sample, generated by comparing matched tumor and normal pairs.</li> <li>MSS.tumor.unique.RepeatMasker.tbl<br>RepeatMasker output annotating somatic insertions in the MSS tumor sample; used in Figure 6.</li> <li>MSS.tumor.unique.vep.html<br>Ensembl VEP annotation report of somatic SVs in the MSS patient; used in Figures 7a and 7b.</li> <li>This dataset supports reproducibility and transparency of the protocol and offers a valuable resource for researchers interested in long-read-based SV detection, repeat annotation, and comparative cancer genomics.</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".

<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>

opencc-by-4.0Nov 2024View details →
dryad32/100

Data from: Population genomic evidence of selection on structural variants in a natural hybrid zone

<p><span>Structural variants (SVs) can promote speciation by directly causing reproductive isolation or by suppressing recombination across large genomic regions. Whereas examples of each mechanism have been documented, systematic tests of the role of SVs in speciation are lacking. Here, we take advantage of long-read (Oxford nanopore) whole-genome sequencing and a hybrid zone between two </span><em>Lycaeides</em> butterfly taxa (<em>L. melissa</em> and Jackson Hole <em>Lycaeides</em>) to comprehensively evaluate genome-wide patterns of introgression for SVs and relate these patterns to hypotheses about speciation. We found &gt;100,000 SVs segregating within or between the two hybridizing species. SVs and SNPs exhibited similar levels of genetic differentiation between species, with the exception of inversions, which were more differentiated. We detected credible variation in patterns of introgression among SV loci in the hybrid zone, with 562 of 1419 ancestry-informative SVs exhibiting genomic clines that deviated from null expectations based on genome-average ancestry. Overall, hybrids exhibited a directional shift towards Jackson Hole <em>Lycaeides</em> ancestry at SV loci, consistent with the hypothesis that these loci experienced more selection on average than SNP loci. Surprisingly, we found that deletions, rather than inversions, showed the highest skew towards excess ancestry from Jackson Hole <em>Lycaeides</em>. Excess Jackson Hole <em>Lycaeides</em> ancestry in hybrids was also especially pronounced for Z-linked SVs and inversions containing many genes. In conclusion, our results show that SVs are ubiquitous and suggest that SVs in general, but especially deletions, might disproportionately affect hybrid fitness and thus contribute to reproductive isolation.</p>

opencc-zeroApr 2022View details →
dryad32/100

Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa

<p>A comprehensive understanding of the degree to which genomic variation is maintained by selection versus drift and gene flow is lacking in many important species such as <em>Cannabis</em> <em>sativa </em>(<em>C. sativa</em>), one of the oldest known crops to be cultivated by humans worldwide. We generated whole genome resequencing data across diverse samples of feralized (escaped domesticated lineages) and domesticated lineages of <em>C. sativa</em>. We performed analyses to examine population structure, and genome wide scans for FST, balancing selection, and positive selection. Our analyses identified evidence for sub-population structure and further support the Asian origin hypothesis of this species. Feral plants sourced from the U.S. exhibited broad regions on chromosomes 4 and 10 with high <span>𝐹̅</span>ST which may indicate chromosomal inversions maintained at high frequency in this sub-population. Both our balancing and positive selection analyses identified loci that may reflect differential selection for traits favored by natural selection and artificial selection in feral versus domesticated sub-populations. In the U.S. feral sub-population, we found six loci related to stress response under balancing selection and one gene involved in disease resistance under positive selection, suggesting local adaptation to new climates and biotic interactions. In the marijuana sub-population, we identified the gene <em>SMALLER TRICHOMES</em> <em>WITH VARIABLE BRANCHES 2 </em>to be under positive selection which suggests artificial selection for increased tetrahydrocannabinol yield. Overall the data generated, and results obtained from our study help to form a better understanding of the evolutionary history in <em>C. sativa</em>.</p>

opencc-zeroAug 2022View details →
zenodo32/100

SNiffles structural variant vcf SHRSP genome

<p>Variant cell format&nbsp;file generated by Sniffles2/SURVIVOR analysis</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Variant metadata for 1K genomes reference panel

<p>1K genomes reference panel - variant metadata</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Sanger sequencing of target and off-target genomic regions for gene-edited iPSC clones with SETBP1 genetic variants

<p>This data set includes chromatograms generated using sanger sequencing of targeted regions of genomic DNA from clonal iPSC lines. The iPSC lines include clones generated using CRISPR/Cas9 homology directed repair to introduce genetic variants into <em>SETBP1,</em> and their wild-type controls. Additional files have been included in the data set to link chromatogram (ab1) files to specific iPSC clones for genomic regions across the variant in <em>SETBP1 (</em>SETBP1 clones genetic variant sanger sequencing.xslx)<em> </em>and top<em> </em>off-target sites (SETBP1 clones off-target sanger sequencing.xlsx).&nbsp;</p>

opencc-by-4.0Sep 2024View details →
dryad32/100

Data from: Whole genome sequencing and rare variant analysis in essential tremor families

Essential tremor (ET) is one of the most common movement disorders. The etiology of ET remains largely unexplained. Whole genome sequencing (WGS) is likely to be of value in understanding a large proportion of ET with Mendelian and complex disease inheritance patterns. In ET families with Mendelian inheritance patterns, WGS may lead to gene identification where WES analysis failed to identify the causative single nucleotide variant (SNV) or indel due to incomplete coverage of the entire coding region of the genome, in addition to accurate detection of larger structural variants (SVs) and copy number variants (CNVs). Alternatively, in ET families with complex disease inheritance patterns with gene x gene and gene x environment interactions enrichment of functional rare coding and non-coding variants may explain the heritability of ET. We performed WGS in eight ET families (n=40 individuals) enrolled in the Family Study of Essential Tremor. The analysis included filtering WGS data based on allele frequency in population databases, rare SNV and indel classification and association testing using the Mixed-Model Kernel Based Adaptive Cluster (MM-KBAC) test. A separate analysis of rare SV and CNVs segregating within ET families was also performed. Prioritization of candidate genes identified within families was performed using phenolyzer. WGS analysis identified candidate genes for ET in 5/8 (62.5%) of the families analyzed. WES analysis in a subset of these families in our previously published study failed to identify candidate genes. In one family, we identified a deleterious and damaging variant (c.1367G&gt;A, p.(Arg456Gln)) in the candidate gene, CACNA1G, which encodes the pore forming subunit of T-type Ca(2+) channels, CaV3.1, and is expressed in various motor pathways and has been previously implicated in neuronal autorhythmicity and ET. Other candidate genes identified include SLIT3 which encodes an axon guidance molecule and in three families, phenolyzer prioritized genes that are associated with hereditary neuropathies (family A, KARS, family B, KIF5A and family F, NTRK1). Functional studies of CACNA1G and SLIT3 suggest a role for these genes in ET disease pathogenesis.

opencc-zeroAug 2019View details →
dryad32/100

Genomic structural variants constrain and facilitate adaptation in natural populations of Theobroma cacao, the Chocolate Tree

<p>Genomic structural variants (SVs) can play important roles in adaptation and speciation. Yet, the overall fitness effects of SVs are poorly understood, partly because accurate population-level identification of SVs requires multiple high-quality genome assemblies. Here, we use 31 chromosome-scale, haplotype-resolved genome assemblies of Theobroma cacao – an outcrossing, long-lived tree species that is the source of chocolate – to investigate the fitness consequences of SVs in natural populations. Among the 31 accessions, we find over 160 thousand SVs, which together cover eight times more of the genome than SNPs and short indels (125 Mb vs. 15 Mb). Our results indicate that a vast majority of these SVs are deleterious: they segregate at low frequencies and are depleted from functional regions of the genome. We show that SVs influence gene expression, which likely impairs gene function and contributes to the detrimental effects of SVs. We also provide empirical support for a theoretical prediction that SVs, particularly inversions, increase genetic load through the accumulation of deleterious nucleotide variants as a result of suppressed recombination.<br> Despite the overall detrimental effects, we identify individual SVs bearing signatures of local adaptation, several of which are associated with genes differentially expressed between populations. Genes involved in pathogen resistance are strongly enriched among these candidates, highlighting the contribution of SVs on this important local adaptation trait. Beyond revealing new empirical evidence for the evolutionary importance of SVs, these 31 de novo assemblies provide a valuable resource for genetic and breeding studies in T. cacao. </p>

opencc-zeroJul 2021View details →
zenodo32/100

Gaps and complex structurally variant loci in phased genome assemblies

<p>Supplementary data including code for an science article &#39;Gaps and complex structurally variant loci in phased genome assemblies&#39;.</p>

opencc-by-4.0Jun 2022View details →
dryad32/100

Phylogenomic data of Litsea complex based on genome-wide single-nucleotide variants (SNVs) and plastomes

<p>In this study, we focus on the Litsea complex (Lauraceae), a key lineage with dominant species of evergreen broadleaved forests (EBLFs) in East Asia to gain insights into how evergreen versus deciduous trait shifted, providing insights into the origin and historical dynamics of EBLFs in East Asia under Cenozoic climate change. We reconstructed a robust phylogeny of the Litsea complex using  genome-wide single-nucleotide variants (SNVs) and plastomes. Finally, the dataset of phylogenomic matrices was generated, including five genome-wide SNVs dataset and plastomes matrix used for phylogenetic analysis.</p>

opencc-zeroFeb 2023View details →
zenodo32/100

Single-cell somatic copy number variants in brain using different amplification methods and reference genomes

<p>Variable and constant sized bins for GRCh38 and T2T-Chm13 were generated using the buildGenome scripts provided with Ginkgo (<a href="https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts">https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts</a>).</p>

opencc-by-nd-4.0Aug 2023View details →
zenodo32/100

Comprehensive genomic analysis identifies a diverse landscape of sideroblastic and non-sideroblastic iron related anemias with novel and pathogenic variants in an iron deficient endemic setting

<p>The history for the 5 CMA cases are as&nbsp;follows:-</p> <p>&nbsp;</p> <p>1.&nbsp;<strong>HMA-2&nbsp;</strong>:- 0.6 month male child presented with poor feeding and noted to have mild-moderate pallor. CBC showed Hemoglobin 9.0gm/dl, MCV- 58fl, MCH 17.3, MCHC 29, Iron 113 ug/dl, Ferritin 297 ug/l, TSAT-45%, TIBC-297 ug/dl, Hepcidin- 20.3ng/ml. No transfusions received and antenatal and neonatal history uneventful. HPLC (both parents), alpha sequencing normal. Targeted NGS shows a heterozygous VUS in GLRX5 gene along with a VUS in BMP6 gene. WES- Does not add anything further significant.</p> <p>&nbsp;</p> <p>2.&nbsp;<strong>HMA-3&nbsp;</strong>:- 0.8 month male child with transfusion dependent microcytic hypochromic anemia and neurological developmental delay. Expired at 1 year of age due to secondary infection (acinetobacter and toxoplasmosis). Hb- 6.5, MCV-61.5, MCH-18.3, Iron 45 ug/dl, Ferritin 650. Targeted panel and WES - negative.</p> <p>&nbsp;</p> <p>3.&nbsp;<strong>HMA-4:</strong>- 5 year male child with severe microcytic hypochromic anemia and received just one transfusion outside. Has microcephaly and SNHL. Hb 5, MCV- 56, MCH- 18, MCHC- 23, Iron- 75, Ferritin - 330, TIBC- 290, TSAT- 41%. Targeted NGS and WES show a Homozygous VUS in BMP6 gene (? association?)&nbsp;</p> <p>&nbsp;</p> <p>4.&nbsp;<strong>HMA-6:-</strong>&nbsp;2 year old male child with moderate microcytic hypochromic anemia and focal seizures. Hb -8, MCV- 58, MCH- 16, MCHC-27, Iron- 94, Ferritin -380, Transferrin levels- 10 (very low). Targeted NGS and WES show a single heterozygous VUS mutation in TF gene and another intrinsic rare variant but reported in pop. Database (could it make it compound her for atransferrinemia?)&nbsp;</p> <p>&nbsp;</p> <p>5.&nbsp;<strong>HMA-7:-</strong>&nbsp;7 year male child with short stature, consanguinity in family, moderate macrocytic anemia with HB-7, MCV-113, Iron-55, Ferritin -209, TSAT-54%. Also Bone marrow shows coarse vacuolations in erythroid cells (12%) and 19% RS. Suspicion PMS- But MLPA for mitochondrial genome normal, no deletions. Targeted panel NGS and WES reveal a heterozygous VUS in LARS2 gene (? nature).</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

A pangenome graph reference of 30 chicken genomes allows genotyping of large and complex structural variants

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov32/100

Validation of Optical Genome Mapping for the Identification of Constitutional Genomic Variants in a Postnatal Cohort

ClinicalTrials.gov study NCT05295277. IPD Sharing: NO. Countries: 1. Publications: 7.

closedIPD-NOFeb 2026View details →
dryad32/100

Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad32/100

Data from: Population genomic evidence of selection on structural variants in a natural hybrid zone

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad32/100

Data from: Geographic distribution and adaptive significance of genomic structural variants: an anthropological genetics perspective

Open the record for dataset details and reuse information.

publicDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record