Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
409
datasets available to search
ShareScore release 0.7.1
Dataset results
409 results for “structural variation”
Regional and local variation in chemical, structural, and physical leaf traits for tree species in the northeastern United States, 2016-2023.
This dataset is a compilation of leaf trait measurements for 25 different Northern American tree species in the northeastern United States collected between 2016 and 2023 by the Terrestrial Ecosystems Analysis Lab at the University of New Hampshire. Currently, this dataset contains measurements for 2,006 samples across 18 chemical, physical, and structural traits. Measured traits include stable isotopes for carbon (C) and nitrogen (N), chlorophyll estimates, leaf and petiole dimensions, and leaf and petiole water content. Traits have been measured at plots spanning a wide range of latitude, longitude, elevation, and forest types. A simple table containing these plot descriptions has been included. Additional leaf physiological and optical traits have been measured concurrently on many of these samples and have been or will be published separately. This is a continuous dataset that will be updated on an as needed basis.
The genetic basis of structural colour variation in mimetic Heliconius butterflies
<p>Raw USAXS data from discal region of <em>Heliconius </em>butterflies (<em>H. erato </em>and<em> H. melpomene</em>). The data comes from wings of individuals of two intercross families, one from each species and was used to estimate scale structure variation and a QTL analysis.</p>
Microsatellite genotypes for «Genetic diversity and spatial genetic structure support the specialist‑generalist variation hypothesis in two sympatric woodpecker species»
<p>Species are often arranged along a continuum from “specialists” to “generalists”. Specialists typically use fewer resources, occur in more patchily distributed habitats and have overall smaller population sizes than generalists. Accordingly, the specialist-generalist variation hypothesis (SGVH) proposes that populations of habitat specialists have lower genetic diversity and are genetically more differentiated due to reduced gene flow compared to populations of generalists. Here, expectations of the SGVH were tested by examining genetic diversity, spatial genetic structure and contemporary gene flow in two sympatric woodpecker species differing in habitat specialization. Compared to the generalist great spotted woodpecker (<em>Dendrocopos major</em>), lower genetic diversity was found in the specialist middle spotted woodpecker (<em>Dendrocoptes medius</em>). Evidence for recent bottlenecks was revealed in some populations of the middle spotted woodpecker, but in none of the great spotted woodpecker. Substantial spatial genetic structure and a significant correlation between genetic and geographic distances were found in the middle spotted woodpecker, but only weak spatial genetic structure and no significant correlation between genetic and geographic distances in the great spotted woodpecker. Finally, estimated levels of contemporary gene flow did not differ between the two species. Results are consistent with all but one expectations of the SGVH. This study adds to the relatively few investigations addressing the SGVH in terrestrial vertebrates.</p>
Database for GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.
<p>GWAS SVatalog is a novel visualization tool and database for structural variants (SV) found in a predominantly European population of 101 individuals with Cystic Fibrosis (CF). Aside from the CF-causing variants on chromosome 7 and the LD block in which they lie, the remainder of the genome is comparable to a the 1000 Genomes healthy European population. This data is a collection of SV calls and their linkage disequilibrium (LD) statistics with GWAS-significant SNPs reported in the GWAS Catalog.</p> <p> </p> <p>The goal of this project is to provide a resource to aid fine mapping of GWAS loci using SVs. GWAS loci are generally identified by SNPs which account for an incomplete proportion of genetic variation and phenotypic heritability. Their relevance to the phenotype might be limited, tagging other polymorphisms, such as SVs, that could be the cause of the association signal. To leverage this data to its full potential, visit the <a href="https://svatalog.research.sickkids.ca/" target="_blank" rel="noopener">GWAS SVatalog</a> web tool. Here, interactive visualizations can illustrate SVs identified in high LD with GWAS-significant SNPs, suggesting putative causal variation that could guide additional functional investigation.</p> <p> </p> <p>For more information on how to use GWAS SVatalog, visit the<a href="https://gwas-svatalog-docs.readthedocs.io/en/latest/index.html" target="_blank" rel="noopener noreferrer"> documentation</a>.</p> <p> </p> <p>This project was accomplished in collaboration with the <a href="https://lab.research.sickkids.ca/strug/" target="_blank" rel="noopener">Strug Lab</a> at <a href="https://www.sickkids.ca/en/" target="_blank" rel="noopener">The Hospital for Sick Children (SickKids)</a>, <a href="https://www.tcag.ca/" target="_blank" rel="noopener">The Center for Applied Genomics (TCAG)</a>, and <a href="https://www.utoronto.ca/" target="_blank" rel="noopener">University of Toronto</a>.</p>
Large structural variations in the haplotype-resolved African cassava genome
<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p> </p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>
Fig. 5 in Seasonal and longitudinal variation in fish assemblage structure along an unregulated stretch of the Middle Uruguay River
Fig. 5. Detrended Correspondence Analysis (DCA) applied to ordinate samples according to variations in fish composition and abundance along the Uruguay River. Rectangles depict groups confirmed by a Multiple Response Permutation Procedure (Tab. 2). Sites: S1 = upstream; S6 = downstream. Seasons: Au= Autumn; Sp= Spring; Su= Summer and Wi= Winter.
Fig. 2 in Seasonal and longitudinal variation in fish assemblage structure along an unregulated stretch of the Middle Uruguay River
Fig. 2. Variation (mean ±standard deviation) in species richness and biomass (CPUEb/100m2) along the river channel (A and C) and among seasons (B and D), in the Middle Uruguay River. Sites: S1 = upstream; S6 = downstream. Different letters indicate statistical difference (p <0.05).
Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""
<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. rufipogon</em> DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. glaberrima</em> IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. barthii</em> W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. nivara</em> W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p>
Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture
<p>Anterior cruciate ligament (ACL) rupture is an important condition of the human knee. Second ruptures are common and societal costs are substantial. Canine cranial cruciate ligament (CCL) rupture closely models the human disease. CCL rupture is common in the Labrador Retriever (5.79% prevalence), ~100-fold more prevalent than in humans. Labrador Retriever CCL rupture is a polygenic complex disease, based on genome-wide association study (GWAS) of single nucleotide polymorphism (SNP) markers. Dissection of genetic variation in complex traits can be enhanced by studying structural variation, including copy number variants (CNVs). Dogs are an ideal model for CNV research because of reduced genetic variability within breeds and extensive phenotypic diversity across breeds. We studied the genetic etiology of CCL rupture by association analysis of CNV regions (CNVRs) using 110 case and 164 control Labrador Retrievers. CNVs were called from SNPs using three different programs (PennCNV, CNVPartition, and QuantiSNP). After quality control, CNV calls were combined to create CNVRs using ParseCNV and an association analysis was performed. We found no strong effect CNVRs but found 46 small effect (max(T) permutation P<0.05) CCL rupture associated CNVRs in 22 autosomes; 25 were deletions and 21 were duplications. Of the 46 CCL rupture associated CNVRs, we identified 39 unique regions. Thirty four were identified by a single calling algorithm, 3 were identified by two calling algorithms, and 2 were identified by all three algorithms. For 42 of the associated CNVRs, frequency in the population was <10% while 4 occurred at a frequency in the population ranging from 10-25%. Average CNVR length was 198,872bp and CNVRs covered 0.11 to 0.15% of the genome. All CNVRs were associated with case status. CNVRs did not overlap previous canine CCL rupture risk loci identified by GWAS. Associated CNVRs contained 152 annotated genes; 12 CNVRs did not have genes mapped to CanFam3.1. Using pathway analysis, a cluster of 19 homeobox domain transcript regulator genes was associated with CCL rupture (P=6.6E-13). This gene cluster influences cranial-caudal body pattern formation during embryonic limb development. Clustered genes were found in 3 CNVRs on chromosome 14 (HoxA), 28 (NKX6-2), and 36 (HoxD). When analysis was limited to deletion CNVRs, the association was strengthened (P=8.7E-16). This study suggests a component of the polygenic risk of CCL rupture in Labrador Retrievers is associated with small effect CNVs and may include aspects of stifle morphology regulated by homeobox domain transcript regulator genes.</p>
Data for "Variation in Upper Plate Crustal and Lithospheric Mantle Structure in the Greater and Lesser Antilles from Ambient Noise Tomography"
<p>This is the phase velocity information and the shear wave model for the g-cubed paper:</p> <p>"Variation in Upper Plate Crustal and Lithospheric Mantle Structure in the Greater and Lesser Antilles from Ambient Noise Tomography"</p>
Data for "Influence of variation in grain boundary parameters on the evolution of atomic structure and properties of [111] tilt grain boundaries in aluminum"
<p>This repository contains the raw data of experimental STEM images and of the simulations for the paper "Influence of variation in grain boundary parameters on the evolution of atomic structure and properties of [111] tilt boundaries in aluminum".</p>
The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)
<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p> import numpy as np<br> dense = np.load("filename.npz")<br> dense["chr1"]<br> <br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM. </p> <p><strong>Supplemental File 5</strong> Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer <= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>
Long-read sequencing reveals extensive gut phageome structural variations driven by genetic exchange with bacterial hosts
<p><span>Genetic variations are instrumental for unraveling phage evolution and deciphering their functional implications. Here we explore the underlying fine-scale genetic variations in the gut phageome, especially structural variations (SVs). By employing virome-enriched long-read metagenomics sequencing across 91 individuals, we identified a total of 14,438 non-redundant phage SVs, and revealed their prevalence within the human gut phageome. These SVs are mainly enriched in genes involved in recombination, DNA methylation, and antibiotic resistance. Strikingly, a substantial fraction of phage SV sequences share close homology with bacterial fragments, with most SVs enriched for horizontal gene transfer (HGT) mechanism. Further investigations showed that these SV sequences were genetic exchanged between specific phage-bacteria pairs, particularly between phages and their respective bacterial hosts. Temperate phages exhibits a higher frequency of genetic exchange with bacterial chromosomes then virulent phages. Collectively, our findings provide novel insights into the genetic landscape of the human gut phageome.</span></p>
Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads"
<p>Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads".</p> <p>The archive contains files that are necessary to reproduce the cell line benchmarks from the paper, including:</p> <ul> <li>Scripts and command lines</li> <li>Original VCF outpurs of all tools used in benchmarking</li> <li>Minda evaluations and truthset VCF files</li> <li>Full Severus outputs + visualizations</li> <li>truvari calls</li> </ul>
Tensile2d: 2D quasistatic non-linear structural mechanics solutions, under geometrical variations
<p>This dataset contains 2D quasistatic non-linear structural mechanics solutions, under geometrical variations. </p> <p>A Description is provided in <a href="https://arxiv.org/pdf/2305.12871.pdf">the MMGP paper</a> Sections 4.1 and A.2.</p> <p>The file format is PLAID, see <a href="https://plaid-lib.readthedocs.io/ ">the plaid documentation</a>.</p> <p>The variablity in the samples are 6 input scalars and the geometry (mesh). Outputs of interest are 4 scalars and 6 fields.</p> <p>Seven nested training sets of sizes 8 to 500 are provided, with complete input-output data. A testing set of size 200, as well as two out-of-distribution sample, are provided, for which outputs are not provided. </p> <p> </p> <p>Tips to access the data:</p> <p>After decompressing the downloaded file:</p> <p>from plaid.containers.dataset import Dataset<br>from plaid.problem_definition import ProblemDefinition</p> <p>dataset = Dataset()<br>problem = ProblemDefinition()</p> <p>problem._load_from_dir_(os.path.join(/path/to/data,'problem_definition'))<br>dataset._load_from_dir_(os.path.join(/path/to/data,'dataset'), verbose = True)</p> <p>print("problem =", problem)<br>print("dataset =", dataset)</p> <p>sample = dataset[0]<br>print("sample =", sample)</p> <p>for fn in sample.get_field_names():<br> print(f"{fn} =", sample.get_field(fn))<br>for sn in sample.get_scalar_names():<br> print(f"{sn} =", sample.get_scalar(sn))</p> <p>print("nodes =", sample.get_nodes())<br>print("elements =", sample.get_elements())<br>print("nodal_tags =", sample.get_nodal_tags())</p> <p> </p>
Fig. 3 in Altitudinal Variation In Population Density, Body Size And Morphometric Structure In C A R A B U S O D O R At U S S H I L, 1996 (C O L E O P T E R A: Carabidae)
Fig. 3. Illustration of measurements: 1-2 – elytra length (hereafter "A", 3-4 – elytra width ("B"), 5-6 – pronotum length "C"), 7-8 – pronotum width (D), 9-10 – head length (E), 11–12 – distance between the eyes (signed as "head width" or "F" in the figures).
Fig. 7 in Altitudinal Variation In Population Density, Body Size And Morphometric Structure In C A R A B U S O D O R At U S S H I L, 1996 (C O L E O P T E R A: Carabidae)
Fig. 7. Descriptive statistics of elytra length means in C. odoratus at the plots on different altitudes.
Figure S2 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census
Figure S2. – Diel variation in observation frequency of individual species. Gobiusculus flavescens and S. melops were both more frequent at daytime, while A. anguilla, M. scorpius, S. trutta, G. morhua, C. harengus and T.bubalis were, all more frequently encountered at night.
Figure S1 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census
Figure S1. – Diel variations in assemblage structure was significant (X2 P <0.05). Demersal fishes are most abundant at day, while benthic fishes are dominant at night. Abundance of pelagic fishes increase at night.
Figure 2. – Monthly average species richness observed during diurnal counts from June 2013 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census
Figure 2. – Monthly average species richness observed during diurnal counts from June 2013 to August 2014. Error bars represent ±SD. Number of counts per month are, indicated at column bases.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.