Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

14 results for “TWAS”

Learn how ShareScore rates datasets ↗
zenodo44/100

COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS

<p>We provide&nbsp;significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS.&nbsp;We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within &plusmn;1Mb, which we assumed they act through cis mechanisms. We include&nbsp;the model&rsquo;s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value &lt; 0.05 and R<sup>2</sup> &gt; 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Common factor GWAS and TWAS output for nociplastic type pain

<p>file: GSEM_commonFactorGWAS_COPC_6trait_30MAY2023.csv.gz</p> <p>description: Common factor GWAS output for GenomicSEM analyses of 6 COPC traits (see&nbsp;doi:&nbsp;https://doi.org/10.1101/2023.06.27.23291959)</p> <p>columns:</p> <p>SNP = rsID SNP identifier&nbsp;</p> <p>CHR = chromosome</p> <p>BP = base pair position</p> <p>MAF = minor allele frequency</p> <p>A1 = effect allele</p> <p>A2 = other allele</p> <p>i = index (1 - n SNPs)</p> <p>lhs = left hand side of equation&nbsp;</p> <p>op = equation operator (lavaan syntax)</p> <p>rhs = right hand side of equation</p> <p>est = effect size (beta)</p> <p>se_c = standard error of effect estimate</p> <p>Z_Estimate = Z value</p> <p>Pval_Estimate = p value of effect</p> <p>Q = Q (heterogeneity) value</p> <p>Q_df = degrees of freedom for Q</p> <p>Q_pval = Q p value</p> <p>fail = GSEM fail message if applicable</p> <p>warning = GSEM warning message if applicable&nbsp;</p> <p>Z_smooth = smoothing parameter if applicable&nbsp;</p> <p>N_estimate = N estimate&nbsp;</p> <p>&nbsp;</p> <p>file: GSEM_commonFactorTWAS_COPC_6trait_30MAY2023.csv.gz</p> <p>description: Common factor TWAS output for GenomicSEM analyses of 6 COPC traits (see&nbsp;doi:&nbsp;https://doi.org/10.1101/2023.06.27.23291959)</p> <p>columns:</p> <p>Gene = ensembl gene ID&nbsp;</p> <p>Panel = which model (tissue+gene) i.e. reference weights file</p> <p>HSQ = gene heritability&nbsp;</p> <p>i = index (1 - n gene-tissue models)</p> <p>lhs = equation left hand side</p> <p>op = operator (lavaan syntax)</p> <p>rhs = equation right hand side</p> <p>est = association estimate (beta)</p> <p>se_c = standard error of beta</p> <p>Z_Estimate = Z value&nbsp;</p> <p>Pval_Estimate = p value of association test</p> <p>Q = Q (heterogeneity) value</p> <p>Q_df = degrees of freedom for Q</p> <p>Q_pval = p value for Q</p> <p>fail = GSEM fail message if applicable&nbsp;</p> <p>warning = GSEM warning message if applicable&nbsp;</p> <p>tissue = tissue</p> <p>p_bonf_tissue = adjusted p value - bonferroni adjustment within tissue&nbsp;</p> <p>p_fdr_tissue = adjusted p value - false discovery rate adjustment within tissue</p> <p>threshold_bonf_tissue = p value threshold for bonferroni adjustment within tissue&nbsp;</p> <p>p_bonf_experiment = adjusted p value - bonferroni adjustment experiment-wide</p> <p>p_threshold_bonf_experiment = p value threshold (bonferroni, experiment-wide)</p> <p>Q_bonf_tissue = adjusted p value for Q, bonferroni within-tissue&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Significant sQTLs and splicing TWAS reference panels (AMP-AD brain and EADB Belgian LCL cohorts)

<p>This dataset is part of the manuscript &quot;<em><strong>New insights into the genetic etiology of Alzheimer&rsquo;s disease and related dementias</strong></em>&quot; by Bellenguez, K&uuml;&ccedil;&uuml;kali, et al. Nature Genetics 2022.</p> <p>Publication link:&nbsp;<a href="https://www.nature.com/articles/s41588-022-01024-z">https://www.nature.com/articles/s41588-022-01024-z</a></p> <p>GitHub repository of all QTL/TWAS data shared&nbsp;for this study:&nbsp;<a href="https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS">https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS</a></p> <p>For details, please see the publication.&nbsp;For any questions, please contact Fahri K&uuml;&ccedil;&uuml;kali (<a href="mailto:fahri.kucukali@uantwerpen.vib.be">fahri.kucukali@uantwerpen.vib.be</a>) and Kristel Sleegers (<a href="mailto:Kristel.Sleegers@uantwerpen.vib.be">Kristel.Sleegers@uantwerpen.vib.be</a>).</p> <p>Significant sQTL catalogues are compressed with <em>gzip </em>and tar achieve of splicing TWAS reference panels are compressed with <em>bzip2</em>.</p> <p><strong>sQTL catalogues</strong></p> <p>The files show significant sQTL - splice junction pairs mapped in AMP-AD brain and EADB Belgian LCL cohorts. The catalogues are in hg38/GRCh38 human genome build.&nbsp;Most of the columns in the files&nbsp;are based on FastQTL output (<a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a>).</p> <p><em>sQTL file columns:</em></p> <ol> <li>variant_id - ID of the significant sQTL variant based on the dbSNPv151 rsID annotation or hg38/GRCh38 CHR_POS_REF_ALT ID if rsID not available.</li> <li>junc_id - sJunction ID assigned by Regtools/Leafcutter pipeline. For stranded datasets (ROSMAP and EADB Belgian) strand info is provided with &quot;+&quot; or &quot;-&quot; symbols, and if not stranded, &quot;?&quot; symbol is used.</li> <li>junc_distance - Genomic distance between sQTL variant and splice junction start</li> <li>ma_samples - Number of samples carrying the minor allele</li> <li>ma_count - Total count of minor alleles</li> <li>maf - Minor allele frequency</li> <li>pval_nominal - Nominal P-value&nbsp;of the association</li> <li>slope - Slope of the association&nbsp;with respect to alternative (ALT)&nbsp;allele indicated on column 14</li> <li>slope_se - Standard error of the slope</li> <li>pval_nominal_threshold - Nominal&nbsp;<em>P</em>-value significant threshold for thissJunction</li> <li>min_pval_nominal - Most significant&nbsp;<em>P</em>-value observed for this sJunction</li> <li>pval_beta -&nbsp;permutation&nbsp;<em>P</em>-value&nbsp;obtained via beta approximation and later used to calculate Storey q-values</li> <li>junction - sJunction in chr:start-end splice junction format</li> <li>cluster - The splice cluster of this sJunction</li> <li>genes - Genes overlapping with this sJunction (if any), based on GENCODEv24 (AMP-AD) and GENCODEv32 (EADB Belgian)</li> <li>GRCh38_chr_pos - Genomic position of the variant, separated by underscore</li> <li>ref_alt - Reference (REF) and alternative (ALT) allele of the variant, separated by &quot;&gt;&quot; sign. ALT is the tested (A1) allele</li> </ol> <p>Of note, we also mapped the significant eQTLs in the same datasets (please see the data availability section of the manuscript or the GitHub repository).</p> <p><strong>Splicing TWAS reference panels</strong></p> <p>Custom splicing TWAS reference panels prepared using FUSION pipeline (<a href="http://gusevlab.org/projects/fusion/">http://gusevlab.org/projects/fusion/</a>) in AMP-AD brain and EADB Belgian LCL cohorts. All data in hg38/GRCh38 genome build.&nbsp;In each directory, you will find &quot;.pos&quot;, .&quot;profile&quot;, and &quot;.profile.err&quot; files. These are explained in the FUSION website as well, but briefly these are:</p> <ol> <li><strong>.pos:</strong> This is a position file that describes the 1Mb extended splice junction start and end coordinates for each calculated weight file for splice junction phenotype. Used for scanning the variants in those coordinates for TWAS.</li> <li><strong>.profile: </strong>This informs about all prediction weights calculated, in terms of number of variants in the model, heritability information, and R2 info for each prediction model used (top1, blup, enet, bslmm, lasso; bslmm was not used therefore has NA values).</li> <li><strong>.profile.err: </strong>This summarizes the reference panel in terms of average hsq (with SD), and which model is the best performing.</li> </ol> <p>Each TWAS weight&nbsp;is provided in a .RDat file under <strong>All_Splicing_Weights</strong>, and in this data we included all calculated functional weights independent of the fact that they are heritable features or not. In our TWAS analyses, we included the heritable functional weights at a hsq&nbsp;<em>P</em>-value &le; 0.05 level.</p> <p>Please also see the expression TWAS reference panels we prepared in the same datasets (see the data availability section of the manuscript or the GitHub repository). If you need an LD reference data in hg38/GRCh38 genome build based on 1000 Genomes NFE samples (whose variant ID annotation are matching to these functional weights), suitable for running the TWAS/FUSION pipeline, please contact us.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

CONTENT -- Multi-context genetic modeling TWAS and eAssociation summary statistics

<p>We provide the summary statistics of running CONTENT, the context-by-context approach, and UTMOST on over 22 phenotypes. The phenotypes are listed in the manuscript, and their respective studies and sample size can be found in a table under the supplementary section of the manuscript. All 3 methods were trained on GTEx v7 as well as CLUES, a single-cell RNA sequencing dataset of PBMCs. The data include the gene name, model, cross-validated R^2, prediction pvalue, TWAS p value, TWAS Z score, and a column titled &quot;hFDR&quot; indicating whether the association was statistically significant while employing hierarchical FDR. The benefits of employing such an approach for all methods can be found in the manuscript.</p> <p>&nbsp;</p> <p>We also include the eAssociations that we obtain by training prediction models on GTEx and CLUES alone. For the CxC and UTMOST approaches, these files contain the gene, context, pvalue and adjusted R^2. For CONTENT, these include the gene, context and pvalue and adjusted R^2 for each CONTENT model--the column names are described like a regression of y~x, rsq_y_x, so rsq_observed_full is the adjusted R^2 from regressing the observed expression onto the cross-validated full model predictions. In cases where the R^2 is higher from the specific or shared models, it&#39;s best to use either of those rather than the full model for out of sample prediction.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

TWAS models for eQTLGen

<p>TWAS models based on eQTLGen summary statistics. Instructions on how these weights can be used for TWAS can be found here: TBC.</p> <p>Please cite:</p> <ul> <li>Study describing how the TWAS models were generated: TBC</li> <li>Original study for eQTLGen in which the eQTL summary statistics were generated:&nbsp;https://doi.org/10.1038/s41588-021-00913-z</li> </ul> <p>Contact Oliver Pain for further information (oliver.pain@kcl.ac.uk).</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

MOSTWAS models, TWAS summary statistics, and simulation results for Bhattacharya and Love, 2020

<p>This compressed folder contains three sub-folders that all pertain to models and results generated with Multi-Omic Strategies for Transcriptome-Wide Association Studies (MOSTWAS):</p> <ol> <li><em>MOSTWAS_Models</em> contains compressed folders for MOSTWAS models trains on TCGA breast cancer and ROS/MAP pre-frontal cortex multi-omic data.</li> <li><em>Simulations&nbsp;</em>contains Simulation_Results_MOSTWAS.xlsx that provides full simulation results outlined in Bhattacharya and Love, 2020 (paper accompanying MOSTWAS).</li> <li><em>TWAS_associations</em>&nbsp;contains four Excel files that provide MOSTWAS and local-only TWAS associations for breast cancer-specific survival (using iCOGs GWAS summary statistics), late-onset Alzheimer&#39;s disease risk (using IGAP GWAS summary statistics), and major depressive disorder risk (using PGC GWAS and UK Biobank GWAX summary statistics).</li> </ol>

opencc-by-4.0Apr 2020View details →
zenodo36/100

m6A-TWAS weights for "Genetic Analyses Support the Contribution of mRNA N6-methyladenosine (m6A) Modification to Human Diseases Heritability"

<p>We included the m6A-TWAS weights (trained using FUSION) for our manuscript &quot;Genetic Analyses Support the Contribution of mRNA&nbsp;<em>N</em>6-methyladenosine (m6A) Modification to Human Diseases Heritability&quot;.&nbsp;</p> <p>This dataset is a&nbsp;supplement to the m6A QTL&nbsp;summary statistics data and imputed genotype data at&nbsp;https:// doi.org/10.5281/zenodo.3870952.</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

MOSTWAS models, TWAS summary statistics, and simulation results for Bhattacharya and Love, 2020

<p>MOSTWAS models, TWAS results, simulation results, and comparison to BGW-TWAS results</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

TWAS models for MetaBrain

<p>TWAS models based on MetaBrain summary statistics. Instructions on how these weights can be used for TWAS can be found here: TBC.</p> <p>Please cite:</p> <ul> <li>Study describing how the TWAS models were generated: TBC</li> <li>Original study for MetaBrain in which the eQTL summary statistics were generated:&nbsp;https://doi.org/10.1101/2021.03.01.433439</li> </ul> <p>Contact Oliver Pain for further information (oliver.pain@kcl.ac.uk).</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

TWAS weights for expression and splicing in three human spinal cord segments

<p>TWAS weights files from three human spinal cord segments, created with FUSION.</p> <p>Each file consists of a tar.gz compressed directory containing RData files, readable with the R programming language.</p> <p>For expression weights, each RData file is a set of TWAS models for a single gene.</p> <p>For splicing weights, each RData file is a set of TWAS models for a single mRNA splicing junction read, created using Leafcutter.</p>

opencc-by-4.0Aug 2021View details →
zenodo28/100

Significant eQTLs and expression TWAS reference panels (AMP-AD brain and EADB Belgian LCL cohorts)

<p>This dataset is part of the manuscript &quot;<em><strong>New insights into the genetic etiology of Alzheimer&rsquo;s disease and related dementias</strong></em>&quot; by Bellenguez, K&uuml;&ccedil;&uuml;kali, et al. Nature Genetics 2022.</p> <p>Publication link:&nbsp;<a href="https://www.nature.com/articles/s41588-022-01024-z">https://www.nature.com/articles/s41588-022-01024-z</a></p> <p>GitHub repository of all QTL/TWAS data shared&nbsp;for this study:&nbsp;<a href="https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS">https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS</a></p> <p>For details, please see the publication.&nbsp;For any questions, please contact Fahri K&uuml;&ccedil;&uuml;kali (<a href="mailto:fahri.kucukali@uantwerpen.vib.be">fahri.kucukali@uantwerpen.vib.be</a>) and Kristel Sleegers (<a href="mailto:Kristel.Sleegers@uantwerpen.vib.be">Kristel.Sleegers@uantwerpen.vib.be</a>).</p> <p>Significant eQTL catalogues are compressed with <em>gzip </em>and tar achieve of expression TWAS reference panels are compressed with <em>bzip2</em>.</p> <p><strong>eQTL catalogues</strong></p> <p>The files show significant eQTL - gene pairs mapped in AMP-AD brain and EADB Belgian LCL cohorts. The catalogues are in hg38/GRCh38 human genome build.&nbsp;Most of the columns in the files&nbsp;are based on FastQTL output (<a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a>).</p> <p><em>eQTL file columns:</em></p> <ol> <li>variant_id - ID of the significant eQTL variant based on the dbSNPv151 rsID annotation or hg38/GRCh38 CHR_POS_REF_ALT ID if rsID not available.</li> <li>gene_id - eGene ENSG gene ID based on GENCODEv24 (AMP-AD) or GENCODEv32 (EADB Belgian)</li> <li>tss_distance - Genomic distance between eQTL variant and eGene</li> <li>ma_samples - Number of samples carrying the minor allele</li> <li>ma_count - Total count of minor alleles</li> <li>maf - Minor allele frequency</li> <li>pval_nominal - Nominal <em>P</em>-value&nbsp;of the association</li> <li>slope - Slope of the association&nbsp;with respect to alternative (ALT)&nbsp;allele indicated on column 14</li> <li>slope_se - Standard error of the slope</li> <li>pval_nominal_threshold - Nominal&nbsp;<em>P</em>-value significant threshold for this eGene</li> <li>min_pval_nominal - Most significant&nbsp;<em>P</em>-value observed for this eGene</li> <li>pval_beta -&nbsp;permutation&nbsp;<em>P</em>-value&nbsp;obtained via beta approximation and later used to calculate Storey q-values</li> <li>gene_name - eGene name based on GENCODEv24 (AMP-AD) and GENCODEv32 (EADB Belgian)</li> <li>GRCh38_chr_pos - Genomic position of the variant, separated by underscore</li> <li>ref_alt - Reference (REF) and alternative (ALT) allele of the variant, separated by &quot;&gt;&quot; sign. ALT is the tested (A1) allele</li> </ol> <p>Of note, we also mapped the significant sQTLs in the same datasets (please see the data availability section of the manuscript or the GitHub repository).</p> <p><strong>Expression TWAS reference panels</strong></p> <p>Custom expression TWAS reference panels prepared using FUSION pipeline (<a href="http://gusevlab.org/projects/fusion/">http://gusevlab.org/projects/fusion/</a>) in AMP-AD brain and EADB Belgian LCL cohorts. All data in hg38/GRCh38 genome build.&nbsp;In each directory, you will find &quot;.pos&quot;, .&quot;profile&quot;, and &quot;.profile.err&quot; files. These are explained in the FUSION website as well, but briefly these are:</p> <ol> <li><strong>.pos:</strong> This is a position file that describes the 1Mb extended gene start and end coordinates for each calculated weight file for each gene expression&nbsp;phenotype. Used for scanning the variants in those coordinates for TWAS.</li> <li><strong>.profile: </strong>This informs about all prediction weights calculated, in terms of number of variants in the model, heritability information, and R2 info for each prediction model used (top1, blup, enet, bslmm, lasso; bslmm was not used therefore has NA values).</li> <li><strong>.profile.err: </strong>This summarizes the reference panel in terms of average hsq (with SD), and which model is the best performing.</li> </ol> <p>Each TWAS weight&nbsp;is provided in a .RDat file under <strong>All_Expression_Weights</strong>, and in this data we included all calculated functional weights independent of the fact that they are heritable features&nbsp;or not. In our TWAS analyses, we included the heritable functional weights at a hsq&nbsp;<em>P</em>-value &le; 0.05 level.</p> <p>Please also see the splicing TWAS reference panels we&nbsp;prepared in the same datasets (see the data availability section of the manuscript or the GitHub repository). If you need an LD reference data in hg38/GRCh38 genome build based on 1000 Genomes NFE samples (whose variant ID annotation are matching to these functional weights), suitable for running the TWAS/FUSION pipeline, please contact us.</p>

opencc-by-4.0Dec 2021View details →
zenodo24/100

Chemotherapy Toxicity TWAS Summary Statistics

<p>TWAS Summary Statistics</p>

opencc-by-4.0Jul 2020View details →
zenodo24/100

Population Matched Transcriptome Prediction Increases TWAS Discovery and Replication Rate

<p>Files from the analyses performed in &quot;Population Matched Transcriptome Prediction Increases TWAS Discovery and Replication Rate.&quot; Please note that the original PAGE Wojcik et al 2019 GWAS Summary Statistics are not included in this Zenodo repository.&nbsp;</p>

opencc-by-4.0Sep 2020View details →
geo24/100

TWAS-based translational genomics approach identifies IL10RB as the top candidate gene for COVID-19 host susceptibility and severity

GEO Series GSE180622. Homo sapiens. 48 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record