Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
129
datasets available to search
ShareScore release 0.9.0
Dataset results
129 results for “Structural variant”
Genome-wide comparison reveals large structural variants in the cassava landraces Authors
<p><span>Structural variants (SVs) are critical for plant genomic diversity and phenotypic variation. This study investigates a large, 9.7 Mbp highly repetitive segment on chromosome 12 of <em><span>TMEB117</span></em>, a region not previously characterized in cassava. We aim to explore its presence and variability across multiple cassava landraces, providing insights into its genomic significance and potential implications.</span></p>
POLQ mediates replication-stress induced structural variant formation throughout common fragile sites during mitosis
<p>Larger processed svCapture data files related to the linked publication. Two types of files are provided:</p> <p>1) Data packages, ending in .mdi.package.zip = processed data files that can be loaded by the svCapture Shiny app, or otherwise extracted to explore the contents.</p> <p>2) App bookmarks, ending in .mdi = saved state of the app with stored configurations from which figures were plotted.</p>
Structural variants of 20 wild isolates of Caenorhabditis elegans
<p>20 VCFs of 20 <em>Caenorhabditis elegans</em> wild isolates against VC2010 reference assembly.</p> <p>Reads come Pacific Biosciences sequencing.</p> <p>Alignment was done using ngmlr.</p> <p>Variant calling was done using Sniffles.</p> <p> </p>
Multimodal learning of noncoding variant effects using genome sequence and chromatin structure
<p>ncVarPred-1D3D:</p> <p>The data used for testing the inconsistency among genome sequence, epigenetic profile, and later, to show its relation to 3D chromatin structure can be found in sanity_check_data.tar.gz.</p> <p>Some trained model for noncoding mutation effect prediction (mapping genome sequence to epigenetic profile) can be found in CNN_MLP, CNN_GCN, CNN_RNN_MLP, CNN_RNN_GCN.tar.gz.</p> <p>The trained model for pathogenic variants prediction can be found in fewshot_pathogenic_model.tar.gz. </p> <p>The training data can be found in training_data.tar.gz.</p> <p>Some noncoding variant effects prediction results, e.g. eQTL and pathogenic variants, can be replicated using the data shared in ncVar_data.tar.gz.</p>
structural variants data
<p>The origin data are used to replicate our work. The topic is "<strong>Accuracy assessment of structural variation calling methods for breast cancer gene panel in the absence of gold standard</strong>". </p>
Multimodal learning of noncoding variant effects using genome sequence and chromatin structure
<p>ncVarPred-1D3D: pretrained models of Sei (PMID: 35817977) + our 3D structure embedding models are shared. The models are trained and validated using DeepSEA (PMID: 26301843) selected 200 bp regions (we extended to 4K bp neighboring) to predict the epigenetic profile containing 21907 epigenetic events Sei processed.</p> <p>The pretrained DeepSEA (PMID: 26301843) and reproduced DanQ (PMID: 27084946) can be found in SOTA.tar.gz.</p>
Dataset related to "CFH and CFHR structural variants in atypical Hemolytic Uremic Syndrome: Prevalence, genomic characterization and impact on outcome "
<p>The upload consists of 2 excel files, including genetic and clinical data, and 6 power point files with western blot images.</p> <p>SMRT sequencing data are deposited in the EBI European Nucletide Archive; Accession number: PRJEB44176. </p>
Calling structural variants with confidence from short-read data in wild bird populations
Open the record for dataset details and reuse information.
The role of structural variants in pest adaptation and genome evolution of the Colorado potato beetle, Leptinotarsa decemlineata (Say)
Open the record for dataset details and reuse information.
Structural variants underlie parallel adaptation following global invasion
Open the record for dataset details and reuse information.
Data from: Structural and epistatic regulatory variants cause hallmark white spotting in cattle
Open the record for dataset details and reuse information.
Data from: Genome divergence between European anchovy ecotypes fuelled by structural variants originating from trans-equatorial admixture
Open the record for dataset details and reuse information.
PacBio Simulation CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 2 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
PacBio Simulation CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 1 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
10XG CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 1 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
Illumina CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 2 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
Illumina CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 1 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
Characterization of large-scale structural variants using Linked-Reads
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
Identity-by-descent detection across 487,409 British samples reveals fine scale population structure and ultra-rare variant associations: data related to publication
<p>Data related to the following publication:</p> <p>"Identity-by-descent detection across 487,409 British samples reveals fine scale population structure and ultra-rare variant associations"</p>
Data from: The role of structural genomic variants in population differentiation and ecotype formation in Timema cristinae walking sticks
Theory predicts that structural genomic variants such as inversions can promote adaptive diversification and speciation. Despite increasing empirical evidence that adaptive divergence can be triggered by one or a few large inversions, the degree to which widespread genomic regions under divergent selection are associated with structural variants remains unclear. Here we test for an association between structural variants and genomic regions that underlie parallel host-plant associated ecotype formation in Timema cristinae stick insects. Using mate-pair re-sequencing of 20 new whole genomes we find that modest-sized structural variants such as inversions, deletions, and duplications are widespread across the genome, being retained as standing variation within and among populations. Using 160 previously published, standard-orientation whole genome sequences we find little to no evidence that the DNA sequences within inversions exhibit accentuated differentiation between ecotypes. In contrast, a formerly described large region of reduced recombination that harbors genes controlling color-pattern exhibits evidence for accentuated differentiation between ecotypes, which is consistent with differences in the frequency of color-pattern morphs between host-associated ecotypes. Our results suggest that some types of structural variants (e.g., large inversions) are more likely to underlie adaptive divergence than others, and that structural variants are not required for subtle yet genome-wide genetic differentiation with gene flow.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.