Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

55

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

55 results for “genomic structural variation”

Learn how ShareScore rates datasets ↗
dryad32/100

Genome-structural analyses support an allotetraploid origin of the walnut family from within Myricaceae and shared genome duplications reveal substitution rate variation

<p><span>In lineages of allopolyploid origin, entire parental subgenomes may coexist, with two or more sets of homoeologous chromosomes that differ in gene content and syntenic structure. Presence or absence of genes, and microsynteny along chromosomal blocks, can be used to differentiate subgenomes and can be coded as phylogenetic data. We assembled chromosome-level genomes of representative species across an ancient allopolyploid lineage, the walnut family (Juglandaceae)</span><span>, with <em>Myrica</em> and other Fagales as outgroups, and used genome-structural data to infer a phylogeny. </span><span>Microsynteny (with various collinear block sizes) and gene content analyses, using the dominant or recessive progenitor subgenomes or both, all yielded identical topologies that place <em>Engelhardia</em> (a SE Asian and Central American clade) with <em>Platycarya</em>, an </span><span>enigmatic monospecific taxon endemic in </span><span>East</span> <span>Asia</span><span>, but well-represented in the Paleocene-Eocene of North America and Europe. </span><span>Morphological studies including fossils also found the <em>Platycarya</em>/<em>Engelhardia</em> clade because of leaf architecture, floral morphology, and nut walls without lacunae, but DNA-alignment-based phylogenetics carried out here and in previous studies never detected this uniformly wind-dispersed clade, instead grouping <em>Platycarya</em> with <em>Carya</em> and <em>Juglans</em>. The novel analyses further reveal </span><span>the family's hybrid origin from extinct or unsampled progenitors nested within Myricaceae and that <em>Rhoiptelea</em> <em>chiliantha</em></span><span>, the Chinese sister species to all other Juglandaceae, </span><span>contains proportionally more genes related to DNA repair and evolved at a rate 2.6- to 3.5-times slower than the remaining species</span><span>. Our results have implications for the molecular clock hypothesis and suggest that genomic structure contains so-far undervalued phylogenetic signal</span><span>.</span></p>

opencc-zeroJun 2022View details →
dryad32/100

Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing

<p>Studies of structural variation (SV) have been challenging due to technological contraints. With the advent of third generation (long-read) sequencing technology, exploration of longer stretches of DNA not easily examined previously has been made possible. In the present study, we utilized third generation (long-read) sequencing techniques to examime SV in the <em>EGFR </em>landscape of four haplotypes derived from two human samples. We analyzed the <em>EGFR</em> gene and its landscape (+/- 500,000 base pairs) using this sequencing approach and were able to identify regions of non-coding DNA which had relatively high similarity to the most common activating <em>EGFR</em> mutation in non-small cell lung cancer. We discovered that reverse complements to the exon 19 deletion mutation which had at least 60% homology to the <em>EGFR</em> exon 19 canonical deletion and were within ± 421,000 bp of the deletion varied across the five haploid genomes examined (4 patient landscapes and hg38). Although the sample size is limited in this study, the estimated variation observed in genomic stability between the five <em>EGFR</em> haplotypes examined is novel and encourages further work to examine structural variation in larger cohorts.</p>

opencc-zeroAug 2021View details →
zenodo32/100

Data from: Chromosome-scale reference genome and RAD-based genetic map of yellow starthistle (Centaurea solstitialis) reveal putative structural variation and QTLs associated with invader traits

<p>The data directory here includes all of the data and scripts necessary to recreate the results and plots for the manuscript titled&nbsp;&quot;Chromosome-scale reference genome and RAD-based genetic map of yellow starthistle (Centaurea solstitialis) reveal putative structural variation and QTLs associated with invader traits&quot;. These data include a genetic map, QTL analysis, paleolog analysis, gene synteny analysis,&nbsp;&nbsp;and assembly validation for yellow starthistle (Centaurea solstitialis).</p>

openNov 2022View details →
dryad32/100

Data for: Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

<p>Distribution of ecotypic variation in natural populations is influenced by neutral and adaptive evolutionary forces that are challenging to disentangle without understanding of genomic architecture for phenotypic traits. This study provides a high-resolution portrait of genomic variation in Chinook salmon (<em>Oncorhynchus</em> <em>tshawytscha</em>) with emphasis on a region of major effect for ecotypic variation in migration timing. With a filtered dataset of ~13 million SNPs from low coverage whole genome resequencing of 53 populations (3,566 barcoded individuals), we contrasted patterns of genomic variation within and among major lineages and examined the extent of a selective sweep at a major effect region underlying migration timing (GREB1L/ROCK1). Allele frequency variation in GREB1L/ROCK1 was highly correlated with mean migration timing for early- and late-run populations within each of the lineages (r<sup>2</sup> between 0.58–0.95; P &lt; 0.001). However, the extent of selection within the genomic region controlling migration timing was much narrower in one lineage (interior stream-type) compared to the other two major lineages which corresponded to the breadth of phenotypic variation in migration timing observed among lineages. Evidence of a duplicated block within GREB1L/ROCK1 may be responsible for reduced recombination in this portion of the genome and contributes to phenotypic variation within and across lineages. Lastly, SNP positions across GREB1L/ROCK1 were assessed for their utility in discriminating migration timing among lineages, and we recommend multiple markers nearest the duplication to provide highest accuracy in conservation applications such as those that aim to protect early migrating Chinook salmon. These results highlight the need to investigate variation throughout the genome and the effects of structural variants on ecologically relevant phenotypic variation in natural species.</p>

opencc-zeroMar 2023View details →
ClinicalTrials.gov32/100

Human Genomic Population Structure and Phenotype-genotype Variation in ADME Genes in Four Populations

ClinicalTrials.gov study NCT02789527. IPD Sharing: UNDECIDED. Countries: 4. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Data from: The standing pool of genomic structural variation in a natural population of Mimulus guttatus

Open the record for dataset details and reuse information.

publicDec 2014View details →
dryad32/100

Data from: Oceanographic variation influences spatial genomic structure in the sea scallop, Placopecten magellanicus

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad32/100

Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad32/100

Data for: Genomic variation across Chinook salmon populations reveals effects of a duplication on migration alleles and supports fine scale structure

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad32/100

Genome-structural analyses support an allotetraploid origin of the walnut family from within Myricaceae and shared genome duplications reveal substitution rate variation

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad32/100

Data from: On the roles of landscape heterogeneity and environmental variation in determining population genomic structure in a dendritic system

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad32/100

Duck pan-genome reveals two transposon-derived structural variations caused bodyweight enlarging and white plumage phenotype formation during evolution

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad28/100

Data from: Spatiotemporally explicit demographic modelling supports a joint effect of historical barriers to dispersal and contemporary landscape composition on structuring genomic variation in a red-listed grasshopper

Inferring the processes underlying spatial patterns of genomic variation is fundamental to understand how organisms interact with landscape heterogeneity and to identify the factors determining species distributional shifts. Here, we employ genomic data (ddRADSeq) to test biologically-informed models representing historical and contemporary demographic scenarios of population connectivity for the Iberian cross-backed grasshopper Dociostaurus hispanicus, a species with a narrow distribution that currently forms highly fragmented populations. All models incorporated biological aspects of the focal taxon that could hypothetically impact its geographical patterns of genomic variation, including (a) spatial configuration of impassable barriers to dispersal defined by topographic landscapes not occupied by the species, (b) distributional shifts resulted from the interaction between the species bioclimatic envelope and Pleistocene glacial cycles, and (c) contemporary distribution of suitable habitats after extensive land clearing for agriculture. Spatiotemporally-explicit simulations under different scenarios considering these aspects and statistical evaluation of competing models within an Approximate Bayesian Computation (ABC) framework supported spatial configuration of topographic barriers to dispersal and human-driven habitat fragmentation as the main factors explaining the geographical distribution of genomic variation in the species, with no apparent impact of hypothetical distributional shifts linked to Pleistocene climatic oscillations. Collectively, this study supports that both historical (i.e., topographic barriers) and contemporary (i.e., anthropogenic habitat fragmentation) aspects of landscape composition have shaped major axes of genomic variation in the studied species and emphasizes the potential of model-based approaches to gain insights into the temporal scale at which different processes impact the demography of natural populations.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Spatiotemporally explicit demographic modelling supports a joint effect of historical barriers to dispersal and contemporary landscape composition on structuring genomic variation in a red-listed grasshopper

Open the record for dataset details and reuse information.

publicApr 2019View details →
geo24/100

Population Structure, and Selection Signatures underlying High-Altitude Adaptation Inferred from Genome-Wide Copy Number Variations in Chinese Indigenous Cattle

GEO Series GSE142218. Bos indicus; Bos grunniens; Bos taurus. 355 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenFeb 2020View details →
geo24/100

Primate genome architecture linked with formation mechanisms and functional consequences of structural variation

GEO Series GSE45741. Macaca mulatta; Pan troglodytes; Pongo pygmaeus; Pongo abelii. 30 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenAug 2013View details →
geo24/100

Comprehensive Long Span Paired-End-Tag Mapping Reveals Characteristic Patterns of Structural Variations in Epithelial Cancer Genomes

GEO Series GSE26954. Homo sapiens. 24 samples. Type: Genome variation profiling by high throughput sequencing.

openGEO-OpenApr 2011View details →
geo24/100

Ruler Arrays Reveal Haploid Genomic Structural Variation

GEO Series GSE23524. Saccharomyces cerevisiae. 2 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenSep 2010View details →
geo24/100

Structural genomic variation analysis in patients with bone marrow failure using Illumina Infinium SNP Arrays [Omni1-Quad]

GEO Series GSE48482. Homo sapiens. 55 samples. Type: SNP genotyping by SNP array.

openGEO-OpenJan 2014View details →
geo24/100

Fine-Scale Mapping and Sequencing of Structural Variation from Eight Human Genomes

GEO Series GSE10008. Homo sapiens. 38 samples. Type: Genome variation profiling by genome tiling array.

openGEO-OpenMay 2008View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record