Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

198

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

198 results for “variant analysis”

Learn how ShareScore rates datasets ↗
dryad36/100

X-linked multi-ancestry meta-analysis reveals tuberculosis susceptibility variants

<p>Globally, tuberculosis (TB) presents with a clear male bias that cannot be completely accounted for by environment, behaviour, socioeconomic factors, or the impact of sex hormones on the immune system. This suggests that genetic and biological differences, which may be mediated by the X chromosome, further influence the observed male sex bias. The X chromosome is heavily implicated in immune function and yet has largely been ignored in previous association studies. Here we report the first multi-ancestry X chromosome specific meta-analysis on TB susceptibility. We identified X-linked TB susceptibility variants using seven genotyping data sets and 20,255 individuals from diverse genetic ancestries. Sex-specific effects were also identified in polygenic heritability between males and females along with enhanced concordance in direction of genetic effects for males but not females. These sex-specific genetic effects were supported by a sex-stratified and combined meta-analysis conducted using the X chromosome specific XWAS software and a multi-ancestry analysis using the MR-MEGA software. Seven significant associations were identified. Two in the overall analysis (rs6610096, rs7888114) and a second for the female specific analysis (rs4465088) including all data sets. For the ancestry specific meta-analysis three significant associations were identified for males in the Asian cohorts (rs1726176, rs5939510, rs1726203) and one in females for the African cohort (rs2428212). Several genomic regions previously associated with TB susceptibility were reproduced in this study, along with strong ancestry-specific effects. These results support the hypothesis that the X chromosome and sex-specific effects could significantly impact the observed male bias in TB incidence rates globally.   </p>

opencc-zeroJun 2024View details →
zenodo36/100

Transcriptome analysis of T47D cells and H2A.J-KO derivatives for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial gland cells

<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater ampliconic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater ampliconic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater metatranscriptomic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater metatranscriptomic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Comparison between the results from JGA analysis somatic short variant discovery workflow and those from the compatible Terra workflow

<p>Files starting from <code>HCC1143.somatic</code> are the results from <a href="https://github.com/ddbj/jga-analysis/tree/main/somatic-short-variant">JGA analysis somatic short variant discovery workflow</a>. Files starting from <code>submissions_</code> are the results from the compatible Terra workflow.</p> <p>VCFs are identical between two workflows except for the header lines. MAFs are also identical except for the header lines.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Comparison between the results from JGA analysis mitochondrial short variant discovery workflow and those from the compatible Terra workflow

<p>Files starting from <code>NA12878.chrM</code> are the results from <a href="https://github.com/ddbj/jga-analysis/tree/mitocondrial-variant">JGA analysis mitochondrial short variant discovery workflow</a>. Files starting from <code>submissions_</code> are the results from the compatible Terra workflow.</p> <p>VCFs are identical between two workflows except for the header lines.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Enhancing stock price data analysis through variants of principal component analysis

<p>The dataset used in the research titled &quot;Enhancing stock price data analysis through variants of principal component analysis&quot;. It includes the daily closing prices of&nbsp;top 100 stocks in S&amp;P500 from 29th March 2020 to 28th March 2023.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Supplemental Material Outcomes of SDHB pathogenic variant carriers a systematic review and meta-analysis

<p>Supplemental Material Outcomes of SDHB pathogenic variant carriers a systematic review and meta-analysis</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

X-linked multi-ancestry meta-analysis reveals tuberculosis susceptibility variants

Open the record for dataset details and reuse information.

publicJun 2024View details →
dryad36/100

Raw data: Association and functional analysis of angiotensin-converting enzyme 2 gene genetic variants with the pathogenesis of pre-eclampsia

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad36/100

Data from: Association and function analysis of genetic variants and the risk of gestational diabetes mellitus in a southern Chinese population

Open the record for dataset details and reuse information.

publicDec 2024View details →
zenodo32/100

Variant analysis of SARS-CoV-2 genomes

<p>These are supplemental files accompanying a publication.</p>

opencc-by-4.0May 2020View details →
zenodo32/100

Transcriptome analysis of WT versus H2A.J-KO MEFs for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences

<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>We further tested a role for H2A.J in Interferon-Stimulated Gene&nbsp;expression by analyzing the transcriptome of WT and H2A.J MEFs induced into senescence by etoposide. TruSeq stranded DNA libraries were prepared from polyA-selected RNA and sequenced as 43 bp paired-end reads. The fastq sequences were mapped to Gencode.vM24.transcripts.fa.gz (GRCm38 transcriptome) with salmon.&nbsp;Read counts were then aggregated to the gene level with tximeta, and differential gene expression was analysed with DESeq2, edgeR, and limma-voom. Gene set enrichment analysis was performed with camera.</p> <p>The transciptomes of&nbsp; senescent WT and H2AFJ-KO showed strong separation from proliferating MEFs, and a weaker separation distinguished WT and H2A.J-KO MEFs. Strikingly, gene set enrichment analysis indicated highly significant defects in Interferon Response Gene Expression in the H2A.J-KO MEFs in senescence with significant down-regulation in senescent H2A.J-KO cells of a series of oligoadenylate synthase genes (Oas1g, Oas1a, Oasl1, Oas2, Oasl2) and several ISGs. Thus, H2A.J also contributes to ISG expression in the heterologous context of senescent MEFs.</p>

opencc-by-4.0Nov 2020View details →
dryad32/100

Data from: Multivariate analysis of dopaminergic gene variants as risk factors of heroin dependence

BACKGROUND: Heroin dependence is a debilitating psychiatric disorder with complex inheritance. Since the dopaminergic system has a key role in rewarding mechanism of the brain, which is directly or indirectly targeted by most drugs of abuse, we focus on the effects and interactions among dopaminergic gene variants. OBJECTIVE: To study the potential association between allelic variants of dopamine D2 receptor (DRD2), ANKK1 (ankyrin repeat and kinase domain containing 1), dopamine D4 receptor (DRD4), Catechol-O-methyl transferase (COMT) and dopamine transporter (SLC6A3) genes and heroin dependence in Hungarian patients. METHODS: 303 heroin dependent subjects and 555 healthy controls were genotyped for 7 single nucleotide polymorphisms (SNPs): rs4680 of the COMT gene; rs1079597 and rs1800498 of the DRD2 gene; rs1800497 of the ANKK1 gene; rs1800955, rs936462 and rs747302 of the DRD4 gene. Four variable number of tandem repeats (VNTRs) were also genotyped: 120 bp duplication and 48 bp VNTR in exon 3 of DRD4 and 40 bp VNTR and intron 8 VNTR of SLC6A3. We also provide a multivariate model for the associations among them implying Bayesian networks in Bayesian multilevel analysis. FINDINGS AND CONCLUSIONS: In single marker analysis the TaqIA (rs1800497) and TaqIB (rs1079597) variants were associated with heroin dependence. Moreover, -521 C/T SNP (rs1800955) of the DRD4 gene showed nominal association with a possible protective effect of the C allele. After applying the Bonferroni correction TaqIB was still significant suggesting that the minor (A) allele of the TaqIB SNP is a risk component in the genetic background of heroin dependence. The findings of the additional multiple marker analysis are consistent with the results of the single marker analysis, but this method was able to reveal an indirect effect of a promoter polymorphism (rs936462) of the DRD4 gene and this effect is mediated through the -521 C/T (rs1800955) polymorphism in the promoter.

opencc-zeroDec 2012View details →
zenodo32/100

Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".

<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

In silico analysis dataset for HPDL missense variants

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

SARS-CoV-2 exposure in Malawian blood donors: an analysis of seroprevalence and variant dynamics between January 2020 and July 2021

<p>Results of population-based age stratified seroepidemiological investigation in Malawi.</p>

opencc-by-4.0Dec 2020View details →
zenodo32/100

Epigenetic-gene-variant-dynamics-analysis

<p>The dataset and analysis pipeline provided here are part of the study titled "Mapping Epigenetic Gene Variant Dynamics: Comparative Analysis of Frequency, Functional Impact, and Trait Associations in African and European Populations." This study investigates the frequency and functional impact of genetic variants in epigenetic genes across different populations, focusing on African and European ancestries. The files included are essential for replicating the analysis described in the manuscript. Key files include data from the GWAS Catalog and phenotype from the Pan-UK Biobank study and the UK Biobank, facilitating the investigation of trait associations with epigenetic variants. The complete code and description of additional data files necessary for the analysis can be accessed on GitHub at<span>&nbsp;</span><a href="https://github.com/smsinks/epigenetic-gene-variant-dynamics-analysis" target="_new">https://github.com/smsinks/epigenetic-gene-variant-dynamics-analysis</a>.</p>

opencc-by-4.0Jul 2024View details →
dryad32/100

Data from: Whole genome sequencing and rare variant analysis in essential tremor families

Essential tremor (ET) is one of the most common movement disorders. The etiology of ET remains largely unexplained. Whole genome sequencing (WGS) is likely to be of value in understanding a large proportion of ET with Mendelian and complex disease inheritance patterns. In ET families with Mendelian inheritance patterns, WGS may lead to gene identification where WES analysis failed to identify the causative single nucleotide variant (SNV) or indel due to incomplete coverage of the entire coding region of the genome, in addition to accurate detection of larger structural variants (SVs) and copy number variants (CNVs). Alternatively, in ET families with complex disease inheritance patterns with gene x gene and gene x environment interactions enrichment of functional rare coding and non-coding variants may explain the heritability of ET. We performed WGS in eight ET families (n=40 individuals) enrolled in the Family Study of Essential Tremor. The analysis included filtering WGS data based on allele frequency in population databases, rare SNV and indel classification and association testing using the Mixed-Model Kernel Based Adaptive Cluster (MM-KBAC) test. A separate analysis of rare SV and CNVs segregating within ET families was also performed. Prioritization of candidate genes identified within families was performed using phenolyzer. WGS analysis identified candidate genes for ET in 5/8 (62.5%) of the families analyzed. WES analysis in a subset of these families in our previously published study failed to identify candidate genes. In one family, we identified a deleterious and damaging variant (c.1367G&gt;A, p.(Arg456Gln)) in the candidate gene, CACNA1G, which encodes the pore forming subunit of T-type Ca(2+) channels, CaV3.1, and is expressed in various motor pathways and has been previously implicated in neuronal autorhythmicity and ET. Other candidate genes identified include SLIT3 which encodes an axon guidance molecule and in three families, phenolyzer prioritized genes that are associated with hereditary neuropathies (family A, KARS, family B, KIF5A and family F, NTRK1). Functional studies of CACNA1G and SLIT3 suggest a role for these genes in ET disease pathogenesis.

opencc-zeroAug 2019View details →
zenodo32/100

FIGURE. Canonical variate analysis of Corybas hypogaeus, C. obscurus ("darkie"), C. wallii ("triwhite") and two variants of the C. trilobus aggregate ("Rimutaka", "Trotters") in Five new species of Corybas (Diurideae, Orchidaceae) endemic to New Zealand and phylogeny of the Nematoceras clade

FIGURE. Canonical variate analysis of Corybas hypogaeus, C. obscurus ("darkie"), C. wallii ("triwhite") and two variants of the C. trilobus aggregate ("Rimutaka", "Trotters")

opennotspecifiedAug 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record