Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
244
datasets available to search
ShareScore release 0.9.0
Dataset results
244 results for “genomic variants”
Data from: The role of structural genomic variants in population differentiation and ecotype formation in Timema cristinae walking sticks
Open the record for dataset details and reuse information.
Phylogenomic data of Litsea complex based on genome-wide single-nucleotide variants (SNVs) and plastomes
Open the record for dataset details and reuse information.
Genomic structural variants constrain and facilitate adaptation in natural populations of Theobroma cacao, the Chocolate Tree
Open the record for dataset details and reuse information.
Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa
Open the record for dataset details and reuse information.
Data from: Whole genome sequencing and rare variant analysis in essential tremor families
Open the record for dataset details and reuse information.
Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection
Open the record for dataset details and reuse information.
Data from: Population genomics reveals structure at the individual, host-tree scale and persistence of genotypic variants of the undomesticated yeast Saccharomyces paradoxus in a natural woodland
Open the record for dataset details and reuse information.
Association of common genetic variants with brain microbleeds: A genome-wide association study
<p><strong>Objective:</strong> To identify common genetic variants associated with the presence of brain microbleeds (BMB).</p> <p><strong>Methods:</strong> We performed genome-wide association studies in 11 population-based cohort studies and 3 case-control or case-only stroke cohorts. Genotypes were imputed to the Haplotype Reference Consortium or 1000 Genomes reference panel. BMB were rated on susceptibility-weighted or T2*-weighted gradient echo magnetic resonance imaging sequences, and further classified as lobar, or mixed (including strictly deep and infratentorial, possibly with lobar BMB). In a subset, we assessed the effects of <em>APOE</em> ε2 and ε4 alleles on BMB counts. We also related previously identified cerebral small vessel disease variants to BMB.</p> <p><strong>Results: </strong>BMB were detected in 3,556 of the 25,862 participants, of which 2,179 were strictly lobar and 1,293 mixed. One locus in the <em>APOE</em> region reached genome-wide significance for its association with BMB (lead SNP rs769449; OR<sub>any BMB</sub> (95% CI)=1.33 (1.21-1.45); p=2.5x10-10). <em>APOE</em> ε4 alleles were associated with strictly lobar (OR (95% CI)=1.34 (1.19- 1.50); p=1.0x10-6) but not with mixed BMB counts (OR (95% CI)=1.04 (0.86-1.25); p=0.68). <em>APOE</em> ε2 alleles did not show associations with BMB counts. Variants previously related to deep intracerebral hemorrhage and lacunar stroke, and a risk score of cerebral white matter hyperintensity variants, were associated with BMB.</p> <p><strong>Conclusions: </strong>Genetic variants in the <em>APOE</em> region are associated with the presence of BMB, most likely due to the <em>APOE</em> ε4 allele count related to a higher number of strictly lobar BMB. Genetic predisposition to small vessel disease confers risk of BMB, indicating genetic overlap with other cerebral small vessel disease markers.</p>
Data from: Candidate variants for additive and interactive effects on bioenergy traits in switchgrass (Panicum virgatum L.) identified by genome-wide association analyses
Switchgrass is a promising herbaceous energy crop, but further gains in biomass yield and quality must be achieved to enable a viable bioenergy industry. Developing DNA markers can contribute to such progress, but depiction of genetic bases should be reliable, involving not only simple additive marker effects but also interactions with genetic backgrounds, e.g., ecotypes, or synergies with other markers. We analyzed plant height, carbon content, nitrogen content, and mineral concentration in a diverse panel consisting of 512 genotypes of upland and lowland ecotype. We performed association analyses based on exome capture sequencing and tested 439,170 markers for marginal effects, but also 83,290 markers for marker-by-ecotype interactions and up to 311,445 marker pairs for pairwise interactions. Analyses of pairwise interactions focused on subsets of marker pairs preselected based on marginal marker effects, gene ontology annotation, and pairwise marker associations. Our tests identified 12 significant effects. Homology and gene expression information corroborated seven effects and indicated plausible causal pathways: flowering time and lignin synthesis for plant height; plant growth and senescence for carbon content and mineral concentration. Four pairwise interactions were detected, including three interactions preselected based on pairwise marker correlations. Furthermore, one marker-by-ecotype interaction and one pairwise interaction were confirmed in an independent switchgrass panel. Our analyses identified reliable candidate variants for important bioenergy traits in switchgrass. Moreover, they exemplified the importance of interactive effects for the depiction of genetic bases, and illustrated the usefulness of preselection of marker pairs for identifying pairwise marker interactions in association testing.
Genetic epidemiology of blood type, disease and trait variants, and genome-wide genetic diversity in over 11,000 domestic cats
<p><span>In the largest DNA-based study of domestic cat to date, 11,036 individuals (10,419 pedigreed cats from 91 breeds and breed types and 617 non-pedigreed cats) were genotyped via commercial panel testing, </span><span>elucidating the distribution and frequency of known genetic variants associated with blood type, disease and physical traits across cat breeds. </span><span>Blood group determining variants, which are relevant clinically and in cat breeding, were genotyped to assess the across breed distribution of blood types A, B and AB.</span> <span>Extensive panel testing identified 13 disease-associated variants in 48 breeds or breed types for which the variant had not previously been observed, strengthening the argument for panel testing across populations. The study also indicates that multiple breed clubs have effectively used DNA testing to reduce disease-associated genetic variants within certain pedigreed cat populations. Appearance-associated genetic variation in all cats is also discussed. Additionally, we combined genotypic data with phenotype information and clinical documentation</span><span>, actively conducted owner and veterinarian interviews, and recruited cats for clinical examination</span><span> to investigate the causality of a number of</span><span> tested variants across different breed backgrounds</span><span>. Lastly, genome-wide informative SNP heterozygosity levels were calculated to obtain a comparable measure of the genetic diversity in different cat breeds.</span></p> <p><span>This study represents the first comprehensive exploration of informative Mendelian variants in felines by screening over 10,000 domestic cats. The results qualitatively contribute to the understanding of feline variant heritage and genetic diversity and demonstrate the clinical utility and importance of such information in supporting breeding programs and the research community. The work also highlights the crucial commitment of pedigreed cat breeders and registries in supporting the establishment of large genomic databases that when combined with phenotype information can advance scientific understanding and provide insights that can be applied to improve the health and welfare of cats.</span></p>
Data from: Comparative genomics of Methicillin-resistant Staphylococcus aureus ST239: distinct geographical variants in Beijing and Hong Kong
Background: The ST239 lineage is a globally disseminated, multiply drug-resistant hospital-associated methicillin-resistant Staphylococcus aureus (HA-MRSA). We performed whole-genome sequencing of representative HA-MRSA isolates of the ST239 lineage from bacteremic patients in hospitals in Hong Kong (HK) and Beijing (BJ) and compared them with three published complete genomes of ST239, namely T0131, TW20 and JKD6008. Orthologous gene group (OGG) analyses of the Hong Kong and Beijing cluster strains were also undertaken. Results: Homology analysis, based on highest-percentage nucleotide identity, indicated that HK isolates were closely related to TW20, whereas BJ isolates were more closely related to T0131 from Tianjin. Phylogenetic analysis, incorporating a total of 30 isolates from different continents, revealed that strains from HK clustered with TW20 into the 'Asian clade', whereas BJ isolates and T0131 clustered closely with strains of the 'Turkish clade' from Eastern Europe. HK isolates contained the typical φSPβ-like prophage with the SasX gene similar to TW20. In contrast, BJ isolates contained a unique 15 kb PT1028-like prophage but lacked φSPβ-like and φSA1 prophages. Besides distinct mobile genetic elements (MGE) in the two clusters, OGG analyses and whole-genome alignment of these clusters highlighted differences in genes located in the core genome, including the identification of single nucleotide deletions in several genes, resulting in frameshift mutations and the subsequent predicted truncation of encoded proteins involved in metabolism and antimicrobial resistance. Conclusions: Comparative genomics, based on de novo assembly and deep sequencing of HK and BJ strains, revealed different origins of the ST239 lineage in northern and southern China and identified differences between the two clades at single nucleotide polymorphism (SNP), core gene and MGE levels. The results suggest that ST239 strains isolated in Hong Kong since the 1990s belong to the Asian clade, present mainly in southern Asia, whereas those that emerged in northern China were of a distinct origin, reflecting the complexity of dissemination and the dynamic evolution of this ST239 lineage.
VCF of structural variant calls of Nanopore data aligned to dm6 reference genome
<p>Heterozygous chromosome inversions suppress meiotic crossover (CO) formation within an inversion, potentially because they lead to gross chromosome rearrangements that produce inviable gametes. Interestingly, COs are also severely reduced in regions nearby but outside of inversion breakpoints even though COs in these regions do not result in rearrangements. Our mechanistic understanding of why COs are suppressed outside of inversion breakpoints is limited by a lack of data on the frequency of noncrossover gene conversions (NCOGCs) in these regions. To address this critical gap, we mapped the location and frequency of rare CO and NCOGC events that occurred outside of the <em>dl</em>-<em>49</em> <em>chrX</em> inversion in <em>D</em>. <em>melanogaster</em>. We created full-sibling wildtype and inversion stocks and recovered COs and NCOGCs in the syntenic regions of both stocks, allowing us to directly compare rates and distributions of recombination events. We show that COs are completely suppressed within 500 kb of inversion breakpoints, are severely reduced within 2 Mb of an inversion breakpoint, and increase above wildtype levels 2–4 Mb from the breakpoint. We find that NCOGCs occur evenly throughout the chromosome and, importantly, occur at wild-type levels near inversion breakpoints. We propose a model in which COs are suppressed by inversion breakpoints in a distance-dependent manner through mechanisms that influence DNA double-strand break repair outcome but not double-strand break location or frequency. We suggest that subtle changes in the synaptonemal complex and chromosome pairing might lead to unstable interhomolog interactions during recombination that permits NCOGC formation but not CO formation.</p>
Dataset for the manuscript "In silico evaluation of variant calling methods for bacterial whole genome sequencing"
<p>Input data and associated analysis code for reproducing results reported in the manuscript.</p>
Data from: Comparative genomics of Methicillin-resistant Staphylococcus aureus ST239: distinct geographical variants in Beijing and Hong Kong
Open the record for dataset details and reuse information.
Data from: Genome-wide association analysis in dogs implicates 99 loci as risk variants for anterior cruciate ligament rupture
Open the record for dataset details and reuse information.
Genetic epidemiology of blood type, disease and trait variants, and genome-wide genetic diversity in over 11,000 domestic cats
Open the record for dataset details and reuse information.
Data from: Candidate variants for additive and interactive effects on bioenergy traits in switchgrass (Panicum virgatum L.) identified by genome-wide association analyses
Open the record for dataset details and reuse information.
Association of common genetic variants with brain microbleeds: A genome-wide association study
Open the record for dataset details and reuse information.
VCF of structural variant calls of Nanopore data aligned to dm6 reference genome
Open the record for dataset details and reuse information.
Elimination of mitochondrial DNA variants by nuclear genome transfer in human oocytes
GEO Series GSE42077. Homo sapiens. 11 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.