Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “GWAS GTEx”

Learn how ShareScore rates datasets ↗
zenodo40/100

GWAS and GTEx QTL integration

<p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication &quot;Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits&quot; (https://doi.org/10.1101/814350).</p> <p>&nbsp;</p> <pre><em># GTEx GWAS integration </em> This package contains the application of several GWAS-QTL integration methods. The results were analyzed in [this preprint](<em>https://www.biorxiv.org/content/10.1101/814350v1</em>) about GTEx v8 application to several GWAS traits. <em>``` </em><em>. </em><em>|-- colocalization </em><em>| |-- coloc </em><em>| | `-- coloc_enloc_priors_eqtl.tar.gz </em><em>| |-- enloc </em><em>| | |-- enloc_eqtl_eur.tar.gz </em><em>| | `-- enloc_sqtl_eur.tar.gz </em><em>| `-- eur_ld.bed.gz </em><em>|-- prediction_models </em><em>| |-- gtex_v8_expression_mashr_snp_smultixcan_covariance.txt.gz </em><em>| |-- gtex_v8_splicing_mashr_snp_smultixcan_covariance.txt.gz </em><em>| |-- mashr_eqtl.tar </em><em>| `-- mashr_sqtl.tar </em><em>|-- smr </em><em>| |-- SMR_gtex_v8_README.txt </em><em>| `-- SMRresults_GTEx_v8_peQTL5e-08.tar.gz </em><em>|-- smultixcan </em><em>| |-- smultixcan_eqtl.tar.gz </em><em>| `-- smultixcan_sqtl.tar.gz </em><em>`-- spredixcan </em><em> |-- spredixcan_eqtl.tar.gz </em><em> `-- spredixcan_sqtl.tar.gz </em> <em> ``` </em><em> </em>You can uncompress gzipped tarball packages <em>`*.tar.gz` </em>in a UNIX command line with an instruction such as: <em>```bash </em><em>tar -xzvpf smultixcan_eqtl.tar.gz </em><em>``` </em>, and the tar packages (<em>`*.tar`</em>) with an analogous instruction: <em>```bash </em><em>tar -xvpf mashr_eqtl.tar </em><em>``` </em> <em>## Preliminaries </em> <strong>**</strong>Finemapping<strong>** </strong>results are contained in a separate release due to size constraints. GWAS summary statistics for 114 traits were harmonized and imputed to GTEx v8 variants with MAF&gt;0.01 using only european samples. (summary imputation software [here](<em>https://github.com/hakyimlab/summary-gwas-imputation</em>)). Some of the following analyses used the full set of 114 traits, while some focused only on 87 traits whose imputed associations showed no deflation (the imputation algorithm is conservative, and studies with too few available variants have a depleted distribution of association p-values after imputation). The harmonized and imputed GWAS summary statistics are contained in a separate release due to size constraints. For completeness&#39; sake, the imputed summary statistics look like: <em>``` </em><em>variant_id panel_variant_id chromosome position effect_allele non_effect_allele current_build frequency sample_size zscore pvalue effect_size standard_error imputation_status n_cases </em><em>rs554008981 chr1_13550_G_A_b38 chr1 13550 A G hg38 0.017316017316017316 336474 -2.2919929353647097 0.021906050841240293 NA NA imputed NA </em><em>rs201055865 chr1_14671_G_C_b38 chr1 14671 C G hg38 0.012987012987012988 336474 -0.9559192804440632 0.33911301727494103 NA NA imputed NA </em><em>... </em><em>``` </em> The GWAS were split in approximately independent LD regions (Berisa-Pickrell)/ GWAS regions are defined in <em>`eur_ld.bed.gz` </em>(note that a few of them are ill-defined in hg38 and where ignored; only completely defined regions were used). <em>## Colocalization </em> <em>### Enloc </em> ENLOC ([see fotware here](<em>https://github.com/xqwen/integrative</em>)) was run for sQTLs and eQTLs using individuals of european ancestry and DAP-G QTL enrichment results on 87 traits. Result files are included in <em>`enloc_eqtl_eur.tar.gz` </em>and <em>`enloc_sqtl_eur.tar.gz` </em>Each file contains a particular tissue-trait combination. Each row details colocalization between a GWAS region (Berisa-Pickrell) and gene&#39;s or intron&#39;s cis-window. A region might overlap multiple genes/introns or viceversa. Each ENLOC file contains the following columns: <strong>* </strong>gwas_locus: GWAS LD region <strong>* </strong>molecular_qtl_trait: gene or intron <strong>* </strong>locus_gwas_pip: posterior inclusion probability of variants in the GWAS LD region <strong>* </strong>locus_rcp: regional colocalization probability (main colocalization measure) <strong>* </strong>lead_coloc_SNP: snp with highest RCP <strong>* </strong>lead_snp_rcp: rcp of the lead coloc snp <em>### Coloc </em> Coloc ([see software here](<em>https://cran.r-project.org/web/packages/coloc/index.html</em>)) was run using prior probabilities estimated from QTL enrichment of GWAS variants (computed via ENLOC). Results for eQTL are available in <em>`coloc_enloc_priors_eqtl.tar.gz`</em>. Each file contains results for a trait-tissue combination. Columns are: <strong>* </strong>gene_id: gene or intron id <strong>* </strong>p0: probability that neither QTL nor GWAS contain a causal variant <strong>* </strong>p1: probability that only GWAS contains a causal variant <strong>* </strong>p2: probability that only QTL has a causal variant <strong>* </strong>p3: probability that GWAS and QTL have a causal variant and it&#39;s distinct <strong>* </strong>p4: probability that GWAS and QTL have a causal variant and it&#39;s the same (main colocalization measure) <em>## PrediXcan </em> <em>`mashr_eqtl.tar` </em>and <em>`mashr_sqtl.tar` </em>contain prediction models (trained on expression or splicing data respectively, for 49 GTEx tissues) and LD compilations to be used with PrediXcan, S-PrediXcan, MultiXcan and S-MultiXcan. For every tissue, the <em>`mashr_{tissue}.db` </em>file is a SQLite file with the prediction model definitions. <em>`mashr_{tissue}.txt.gz` </em>is a gzipped-text file with the upper triangular matrices of covariance between snps within a gene/intron prediction model. Many variants in these models don&#39;t have an rsid. To fully leverage the information in these models, it is advised to at least harmonize to GTEx variants, and if possible impute as we did [here](<em>https://github.com/hakyimlab/summary-gwas-imputation</em>). <em>### S-PrediXcan </em> S-PrediXcan was run for the 114 harmonized and imputed traits, on eQTL and sQTL mashr prediction models. All of the GWAS traits had the same format, so that the following format parameters were used with S-PrediXcan: <em>``` </em><em>--snp_column panel_variant_id --effect_allele_column effect_allele --non_effect_allele_column non_effect_allele --zscore_column zscore \ </em><em>--keep_non_rsid --additional_output --model_db_snp_key varID \ </em><em>``` </em> Each file is a CSV, with each row containing a gene/intron association at a given trait-tissue combination: <strong>* </strong>gene: ENSEMBLE ID or intron id <strong>* </strong>gene_name: HUGO name or intron id <strong>* </strong>zscore: predicted association z-score <strong>* </strong>effect_size: estimated effect size <strong>* </strong>pvalue: association p-value <strong>* </strong>var_g: estimated variance of predicted expression or splicing <strong>* </strong>pred_perf_r2: prediction model cross-validated performance <strong>* </strong>pred_perf_pval: prediction model cross-validated performance <strong>* </strong>pred_perf_qval: deprecated, empty field left for compatibility <strong>* </strong>n_snps_used: number of snps in the intersection of GWAS and model <strong>* </strong>n_snps_in_cov: number of snps in the LD compilation <strong>* </strong>n_snps_in_model: number of snps in the model <strong>* </strong>best_gwas_p: smallest p-value acros GWAS snps used in this model <strong>* </strong>largest_weight: largest prediction model weight <em>### S-Multixcan </em> S-MultiXcan results were generated from the above S-PrediXcan results. Each fiel contains multi-tissue associations for a given trait: <strong>* </strong>gene: ENSEMBLE ID or intron id <strong>* </strong>gene_name: HUGO name or intron id <strong>* </strong>pvalue: multi-tissue association p-value <strong>* </strong>n: number of models avialble for this gene/intron <strong>* </strong>n_indep: number of independent components of variation in predicted expression/splicing (surviving principal components) <strong>* </strong>p_i_best: highest single-tissue p-value (S-PrediXcan) <strong>* </strong>t_i_best: tissue of highest p-value <strong>* </strong>p_i_worst: lowest single-tissue p-value (S-PrediXcan) <strong>* </strong>t_i_worst: tissue of lowest p-value <strong>* </strong>eigen_max: maximum eigenvalue of SVD <strong>* </strong>eigen_min: minimum eigenvalue of SVD <strong>* </strong>eigen_min_kept: smallest eigenvalue retained after discarding smallest variations <strong>* </strong>z_min: minimum single-tissue z-score <strong>* </strong>z_max: maximum single-tissue z-score <strong>* </strong>z_mean: mean single-tissue zscre <strong>* </strong>z_sd: standard deviation of the single-tissue z-scores <strong>* </strong>tmi: trace of M * M_i where M is predicted expression/splicing covariance across tissues for a gene, and M_i is its SVD pseudo-inverse <strong>* </strong>status: computation status, 0 if no errors <em>## SMR </em> See <em>`SMR_gtex_v8_README.txt` </em>for details.</pre> <p>&nbsp;</p> <p>&nbsp;</p> <p># Disclaimer</p> <p>The data is provided &quot;as is&quot;, and the authors assume no responsibility for errors or omissions. &nbsp;<br> The User assumes the entire risk associated with its use of these data. &nbsp;<br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. &nbsp;<br> The User bears all responsibility in determining whether these data are fit for the User&#39;s intended use. &nbsp;</p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. &nbsp;<br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. &nbsp;<br> &nbsp;<br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia&#39;s requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>

opencc-by-4.0Oct 2019View details →
zenodo32/100

Extended data: Impact of admixture and ancestry on eQTL analysis and GWAS colocalization in GTEx

<p>eQTL summary statistics and GWAS colocalization posterior probabilities&nbsp;from eQTL calling in an admixed subcohort of GTEx v8 with local&nbsp;and global&nbsp;ancestry adjustments. For the original, non-peer-reviewed preprint, see&nbsp;<a href="https://www.biorxiv.org/content/10.1101/836825v1">https://www.biorxiv.org/content/10.1101/836825v1</a>.&nbsp;</p> <p>For the related source code, see&nbsp;<a href="https://doi.org/10.5281/zenodo.3924788">https://doi.org/10.5281/zenodo.3924788</a>&nbsp;or&nbsp;<a href="https://github.com/nicolerg/gtex-admixture-la">https://github.com/nicolerg/gtex-admixture-la</a>.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Publicly available GWAS summary statistics, harmonized and imputed to GTEx v8' variant reference

<p># harmonized and imputed GWAS summary statistics</p> <p>&nbsp;</p> <p>* `harmonized_imputed_gwas.tar` contains 114 publicly available GWAS traits, harmonized and imputed to GTEx v8 reference</p> <p>&nbsp;</p> <p>* `gwas_metadata.txt` is a table with useful information about each trait, such as:</p> <p>- Tag: trait name (also in the file name)</p> <p>-&nbsp; PUBMED_Paper_Link: PUBMED or publication URL (if available)</p> <p>- Portal: URL to web portal from which data was downloaded</p> <p>- Consortium: GWAS Consortium authoring the data</p> <p>- Sample_Size: number of individuals covered in the study</p> <p>- Population: individuals&#39;ancestry (EUR, EAS, etc)</p> <p>-&nbsp; abbreviation: short name used for figures</p> <p>-&nbsp; new_abbreviation: alternative name for additional figures</p> <p>-&nbsp; Deflation: whether imputed summary statistics exhibited deflation (i.e. association p-values are lower than expected by chance. The summary statistics imputation method is conservative, and in public GWAS with few observed variants (&lt;2M), the distribution of p-values lags towards lower significance spectrums.</p> <p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication &quot;Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits&quot; (https://doi.org/10.1101/814350).</p> <p># Disclaimer</p> <p>The data is provided &quot;as is&quot;, and the authors assume no responsibility for errors or omissions. &nbsp;<br> The User assumes the entire risk associated with its use of these data. &nbsp;<br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. &nbsp;<br> The User bears all responsibility in determining whether these data are fit for the User&#39;s intended use. &nbsp;</p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. &nbsp;<br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. &nbsp;<br> &nbsp;<br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia&#39;s requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>

opencc-by-4.0Jan 2020View details →
zenodo24/100

SMR added - GWAS and GTEx QTL integration

<p>The eQTL SMR results had been left out of the following repository due to an error</p> <p>https://zenodo.org/record/3518299</p> <p>sQTL SMR results have also been uploaded to https://zenodo.org/record/3525070</p> <p>&nbsp;</p> <p>----------</p> <p>sqlite version of these results are also available here https://github.com/abhiramrao/gtex_v8_GPMs</p> <p>----------</p> <p>The GTEx Consortium</p> <p>SMR results using eQTL and sQTLs from GTEx (release 8)</p> <p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication &quot;Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits&quot; (https://doi.org/10.1101/814350).</p> <p># Disclaimer</p> <p>The data is provided &quot;as is&quot;, and the authors assume no responsibility for errors or omissions. &nbsp;<br> The User assumes the entire risk associated with its use of these data. &nbsp;<br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. &nbsp;<br> The User bears all responsibility in determining whether these data are fit for the User&#39;s intended use. &nbsp;</p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. &nbsp;<br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. &nbsp;<br> &nbsp;<br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia&#39;s requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>

opencc-by-4.0Oct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record