Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “eQTL”
Plasma circulating microRNA-expression quantitative trait loci (eQTLs) data in the Rotterdam Study
<p>The dataset contains GWAS summary statistics for 2,083 plasma circulating microRNAs, obtained from nearly 2,178 participants of the Rotterdam Study. The dataset includes three files, as outlined below:</p> <p><strong>File1: SNP_reference_file_maf0.01_Rsq0.7.txt</strong></p> <p>A reference file for SNPs with good imputation quality (Rsq > 0.7) and minor allele frequency > 0.01 among participants included in our GWAS in the Rotterdam Study (N=2,178). The headers are:</p> <p>SNP: rsID</p> <p>chr: chromosome number according to GRCh37</p> <p>bp: basepair position according to GRCh37</p> <p>effect_allele: effect allele</p> <p>other_allele: other allele</p> <p>eaf: effect allele frequency</p> <p><strong>File2: miReQTLs_1e-5_maf0.01_Rsq0.7.txt</strong></p> <p>Summary statistics for all SNPs significantly associated with 2083 miRNAs (p-value < 1e-5), filtered by minor allele frequency > 0.01 and Rsq > 0.7. The headers are:</p> <p>SNP: rsID</p> <p>beta: effect estimate</p> <p>se: standard error</p> <p>pval: p-value</p> <p>miRNA: miRNA ID</p> <p><strong>File3: miReQTLs_nominal_sig.csv.gz</strong></p> <p>Summary statistics for all SNPs nominally associated with 2083 miRNAs (p-value < 0.05). The headers are:</p> <p>RSID: SNP ID</p> <p>p-value: p-value</p> <p>phenotype: miRNA</p> <p>SE: standard error</p> <p>BETA: effect estimate</p> <p> </p> <p>The SNP allelic information and frequency can be found in the reference file (<strong>File1</strong>). </p> <p><br>For more information, please contact: m.ghanbari@erasmusmc.nl</p>
HGSVC2 full eQTL results
<p>Full summary statistics of the QTL mappings performed on the GEUVADIS and deep 1000GP RNA-sequencing samples presented in the: "Structural variation characterization from the de novo assembly of 64 haplotype-resolved human genomes of diverse ancestry" paper.<br> <br> **Updated release now including sQTL and updated eQTL results</p>
Heart Failure eQTLs companion to "Pathologic gene network rewiring implicates PPP1R3A as a central cardioprotective factor in pressure overload heart failure"
<p>These are the results of a QTL analysis companion to "Pathologic gene network rewiring implicates PPP1R3A as a central cardioprotective factor in pressure overload heart failure". We performed RNA expression measurements and obtained genotype information in genome-wide markers for 313 patients (177 failing hearts , 136 donor, non-failing [control] hearts) using Affymetrix expression and Affymetrix Human 6.0 respectively.<strong> </strong>Prior to eQTL discovery, we used PEER to find hidden covariates that could confound signals in our data as well as filtering any genotypes with major allele frequencies less than 5%. To test associations between gene expression in each cohort separately, we used QTLTools with an additive model accounting for gender, age, sample site, and the PEER factors as covariates. We corrected for eQTL multiple association testing using a 10000 permutations per locus in a 2 megabase window and a false discovery rate cutoff of 5%. To select the number of PEER factors, we performed the full analysis multiple times from 1 to 15 PEER factors and observed a saturation of new QTLs being discovered when using 10 factors.</p> <p>Four files are provided, two for each cohort (cases and controls):</p> <p>- peer_[cases|controls]_nominal.txt: Nominal associations with a p-value threshold of 0.001</p> <p>- peer_[cases|controls]_permutations_all.significant.txt: All significant associations detected after the QTLtools permutation test.</p> <p>The column names are those from QTLtools, in order:</p> <p><br> 1. The phenotype ID<br> 2. The chromosome ID of the phenotype<br> 3. The start position of the phenotype<br> 4. The end position of the phenotype<br> 5. The strand orientation of the phenotype<br> 6. The total number of variants tested in cis<br> 7. The distance between the phenotype and the tested variant (accounting for strand orientation)<br> 8. The ID of the tested variant ( in Affy 6.0 SNP ids)<br> 9. The chromosome ID of the variant<br> 10. The start position of the variant<br> 11. The end position of the variant<br> 12. The nominal P-value of association between the variant and the phenotype<br> 13. The corresponding regression slope<br> 14. A binary flag equal to 1 is the variant is the top variant in cis</p>
Summary statistics of stimulated fibroblast eQTLs
<p>This dataset comprises summary statistics from an eQTL mapping study of stimulated fibroblast eQTLs demonstrated in the manuscript entitled "Mapping interindividual dynamics of innate immune response at single-cell resolution" by Kumasaka et al.</p>
Results for eQTL and sQTL meta-analysis and colocalization
<p>This dataset is part of the manuscript: "<em>Atlas of genetic effects in human microglia transcriptome across brain regions, aging and disease pathologies</em>", by Lopes KP, Snijders GJL, Humphrey J, et al.</p> <p> </p> <p>Description of files:</p> <p><em>COLOC_supp_table_all_results.tsv.gz - </em>Table with results from <strong>COLOC</strong><em> </em>(gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>disease - disease name (Alzheimer’s disease - AD, Bipolar Disorder - BPD, Multiple sclerosis - MS, Parkinson’s disease - PD, Schizohphrenia - SCZ)</li> <li>GWAS - GWAS study (IMSGC_2019, Jansen_2018, Kunkle_2019, Lambert_2013, Marioni_2018, Nalls23andMe_2019, Ripke_2014, Stahl_2019)</li> <li>locus - locus id according to each GWAS study</li> <li>GWAS_SNP - SNP reported in the GWAS study</li> <li>GWAS_P - <em>P</em>-value of the GWAS_SNP reported in the GWAS study</li> <li>GWAS_chr - chromosome of the GWAS_SNP (hg38)</li> <li>GWAS_pos - genomic position in the chromosome of the GWAS_SNP (hg38)</li> <li>QTL - id for the QTL study</li> <li>type - the type of QTL (eQTL or sQTL)</li> <li>QTL_SNP - SNP id from the QTL association</li> <li>QTL_P - <em>P</em>-value for the QTL association </li> <li>QTL_Beta - Slope (beta) for the QTL association</li> <li>QTL_MAF - minor allele frequency for the QTL_SNP in each QTL study. If not available, values were obtained from the European superpopulation of 1000 Genomes phase 3</li> <li>QTL_chr - chromosome for the QTL_SNP (hg38)</li> <li>QTL_pos - genomic position in the chromosome of the QTL_SNP (hg38)</li> <li>QTL_junction - splicing junction tested in the association (for sQTLs only)</li> <li>QTL_Gene - gene name for the QTL association</li> <li>QTL_Ensembl - Ensembl gene id for the QTL_gene (GENCODE v30)</li> <li>nsnps - number of SNPs tested </li> <li>PP.H0.abf - posterior probability for H0 (no causal variant)</li> <li>PP.H1.abf - posterior probability for H1 (causal variant for trait 1 only)</li> <li>PP.H2.abf - posterior probability for H2 (causal variant for trait 2 only)</li> <li>PP.H3.abf - posterior probability for H3 (two distinct causal variants)</li> <li>PP.H4.abf - posterior probability for H4 (one common causal variant)</li> <li>cell_type - cell type of the QTL study</li> <li>SNP_distance - the absolute distance between GWAS_SNP and QTL_SNP</li> <li>LD - linkage disequilibrium between the GWAS_SNP and the QTL_SNP according to 1000 genomes phase 3 European reference panel 3 (only for PP4>0.5, -Inf otherwise)</li> </ol> <p><em>mashR_lfsr_eQTL.txt.gz - </em><strong>mashR </strong>results for <strong>eQTL</strong><em> </em>(gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>ensembl_snp - Ensembl ID and the SNP prioritized by mashR (best SNP per gene)</li> <li>MFG_eur_expression_peer10.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the MFG region</li> <li>STG_eur_expression_peer10.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the STG region</li> <li>SVZ_eur_expression_peer5.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the SVZ region</li> <li>THA_eur_expression_peer10.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the THA region</li> </ol> <p><em>mashR_lfsr_eQTL.txt.gz - </em><strong>mashR </strong>results for <strong>sQTL</strong><em> </em>(gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>pos_ensembl_rsnp - splicing junction coordinates, Ensembl ID, and SNP ID prioritized by mashR (best SNP per junction)</li> <li>MFG_eur_rsplicing_peer5_gene.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the MFG region</li> <li>STG_eur_rsplicing_peer5_gene.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the STG region</li> <li>SVZ_eur_rsplicing_peer0_gene.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the SVZ region</li> <li>THA_eur_rsplicing_peer5_gene.cis_qtl_nominal - local false sign rate (lfsr) of the gene-SNP pair for the THA region</li> </ol> <p><em>out_mfg_stg_svz_tha.metasoft.gz - </em><strong>METASOFT</strong> results<strong> </strong>for <strong>eQTLs</strong> meta-analysis from MiGA four brain regions<em> </em>(gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>RSID - Id composed by gene Ensembl and SNP ID separated by an underscore for each gene-SNP pair tested in the eQTL study</li> <li>#STUDY - number of studies included in the meta-analysis</li> <li>PVALUE_FE - <em>P</em>-value of the fixed-effects model (FE) according to METASOFT</li> <li>BETA_FE - Estimated Beta under the fixed-effects model according to METASOFT</li> <li>STD_FE - Standard error of BETA_FE</li> <li>PVALUE_RE - <em>P</em>-value of the random effects model (RE) according to METASOFT</li> <li>BETA_RE - Estimated Beta under the random-effects model (RE) according to METASOFT</li> <li>STD_RE - Standard error of BETA_RE</li> <li>PVALUE_RE2 - <em>P</em>-value of the Han and Eskin's Random Effects model (RE2) according to METASOFT</li> <li>STAT1_RE2 - RE2 statistic mean effect part</li> <li>STAT2_RE2 - RE2 statistic heterogeneity part</li> <li>PVALUE_BE - BE P-value (“NA” in all row, -binary_effects option is not used)</li> <li>I_SQUARE - I-square heterogeneity statistic</li> <li>Q - Cochran's Q statistic</li> <li>PVALUE_Q - Cochran's Q statistic's <em>P</em>-value</li> <li>TAU_SQUARE - Tau-square heterogeneity estimator of DerSimonian-Laird</li> <li>PVALUES_OF_STUDIES(Tab_delimitered) - <em>P</em>-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA</li> <li>MVALUES_OF_STUDIES(Tab_delimitered) - M-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA</li> </ol> <p><em>out_miga_young_mynd_fairfax.metasoft.gz - </em><strong>METASOFT</strong> results<strong> </strong>for <strong>eQTL</strong> meta-analysis from MiGA four brain regions plus microglia eQTL from Young et al. (2019), and monocytes eQTL from Navarro et al. (2020) and Fairfax et al. (2014)<em> </em>(gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>RSID - Id composed by gene Ensembl and SNP ID separated by an underscore for each gene-SNP pair tested in the eQTL study</li> <li>#STUDY - number of studies included in the meta-analysis</li> <li>PVALUE_FE - <em>P</em>-value of the fixed-effects model (FE) according to METASOFT</li> <li>BETA_FE - Estimated Beta under the fixed-effects model according to METASOFT</li> <li>STD_FE - Standard error of BETA_FE</li> <li>PVALUE_RE - <em>P</em>-value of the random effects model (RE) according to METASOFT</li> <li>BETA_RE - Estimated Beta under the random-effects model (RE) according to METASOFT</li> <li>STD_RE - Standard error of BETA_RE</li> <li>PVALUE_RE2 - <em>P</em>-value of the Han and Eskin's Random Effects model (RE2) according to METASOFT</li> <li>STAT1_RE2 - RE2 statistic mean effect part</li> <li>STAT2_RE2 - RE2 statistic heterogeneity part</li> <li>PVALUE_BE - BE P-value (“NA” in all row, -binary_effects option is not used)</li> <li>I_SQUARE - I-square heterogeneity statistic</li> <li>Q - Cochran's Q statistic</li> <li>PVALUE_Q - Cochran's Q statistic's <em>P</em>-value</li> <li>TAU_SQUARE - Tau-square heterogeneity estimator of DerSimonian-Laird</li> <li>PVALUES_OF_STUDIES(Tab_delimitered) - <em>P</em>-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA, 5-Young et al., 6-Navarro et al., 7-Fairfax et al.</li> <li>MVALUES_OF_STUDIES(Tab_delimitered) - M-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA, 5-Young et al., 6-Navarro et al., 7-Fairfax et al.</li> </ol> <p><em>out_mfg_stg_svz_tha_sClusters.metasoft.gz - </em><strong>METASOFT</strong> results<strong> </strong>for <strong>sQTLs</strong> meta-analysis from MiGA four brain regions (gzip-compressed). Table columns are formatted as follows:</p> <ol> <li>RSID - Id composed by splicing junction coordinates, gene Ensembl ID, and SNP ID separated by underscores for each junction-SNP pair tested in the sQTL study (e.g. chr1_962047_962355_ENSG00000187961.14_1:11008:C:G)</li> <li>#STUDY - number of studies included in the meta-analysis</li> <li>PVALUE_FE - <em>P</em>-value of the fixed-effects model (FE) according to METASOFT</li> <li>BETA_FE - Estimated Beta under the fixed-effects model according to METASOFT</li> <li>STD_FE - Standard error of BETA_FE</li> <li>PVALUE_RE - <em>P</em>-value of the random effects model (RE) according to METASOFT</li> <li>BETA_RE - Estimated Beta under the random-effects model (RE) according to METASOFT</li> <li>STD_RE - Standard error of BETA_RE</li> <li>PVALUE_RE2 - <em>P</em>-value of the Han and Eskin's Random Effects model (RE2) according to METASOFT</li> <li>STAT1_RE2 - RE2 statistic mean effect part</li> <li>STAT2_RE2 - RE2 statistic heterogeneity part</li> <li>PVALUE_BE - BE P-value (“NA” in all row, -binary_effects option is not used)</li> <li>I_SQUARE - I-square heterogeneity statistic</li> <li>Q - Cochran's Q statistic</li> <li>PVALUE_Q - Cochran's Q statistic's <em>P</em>-value</li> <li>TAU_SQUARE - Tau-square heterogeneity estimator of DerSimonian-Laird</li> <li>PVALUES_OF_STUDIES(Tab_delimitered) - <em>P</em>-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA</li> <li>MVALUES_OF_STUDIES(Tab_delimitered) - M-values of each study in the respective order 1-MFG, 2-STG, 3-SVZ, 4-THA</li> </ol> <p><strong>NOTE:</strong> The effect sizes of eQTLs and sQTL are defined as the effect of the alternative allele (ALT) relative to the reference (REF) allele in the human genome reference (GRCh38). A file containing that information for all alleles tested is available at 10.5281/zenodo.4301005</p>
Full eQTL summary statistics for the PAUSE trial
<p>This dataset contains full eQTL summary statistics for the study 'Immunosuppression causes dynamic changes in expression QTLs in psoriatic skin' (<em>Nat.Comm., </em>2023). The study analyzes 375 skin samples from patients with psoriasis. For each SNP-gene pair, the dataset provides information listed below.</p> <ol> <li><strong>ID</strong>: the variant identifier, with chromosome number, position, ref and alt alleles.</li> <li><strong>Estimate</strong>: estimate</li> <li><strong>Std.Error</strong>: standard error</li> <li><strong>df</strong>: degrees of freedom</li> <li><strong>t value</strong>: t-value of the association</li> <li><strong>Pr(>|t|): </strong>p-value of the association</li> <li><strong>Gene: </strong>Ensembl gene ID</li> <li><strong>CHROM: </strong>the variant chromosome</li> </ol>
Natural Killer cells demonstrate distinct eQTL and transcriptome-wide disease associations, highlighting their role in autoimmunity.
<p><strong>Abstract</strong> </p> <p>Natural Killer (NK) cells are innate lymphocytes with central roles in immunosurveillance and are implicated in autoimmune pathogenesis. The degree to which regulatory variants affect NK gene expression is poorly understood. We performed expression quantitative trait locus (eQTL) mapping of negatively selected NK cells from a population of healthy Europeans (n=245). We find a significant subset of genes demonstrate eQTL specific to NK cells and these are highly informative of human disease, in particular autoimmunity. An NK cell transcriptome-wide association study (TWAS) across five common autoimmune diseases identified further novel associations at 27 genes. In addition to these <em>cis</em> observations, we find novel master-regulatory regions impacting expression of <em>trans</em> gene networks at regions including 19q13.4, the Killer cell Immunoglobulin-like Receptor (KIR) Region, <em>GNLY</em>, <em>MC1R</em> and<em> UVSSA</em>. Our findings provide new insights into the unique biology of NK cells, demonstrating markedly different eQTL from other immune cells, with implications for disease mechanisms.</p> <p><strong>Preprint</strong></p> <p>https://www.biorxiv.org/content/10.1101/2021.05.10.443088v1</p> <p><strong>Dataset</strong></p> <p>nk_raw_for_zenodo.txt: Matrix of raw gene expression at 47,209 probes in primary human NK cells from 245 healthy individuals of European ancestry. Gene expression is quantified using the Illumina HumanHT-12 v4 BeadChip gene expression array platform. Column names represent Array Address ID for each probe, and row names represent pseudonymised sample identifiers, which can be matched to sample genotypes. Sample genotypes are available at the European Genome-Phenome Archive with accession ID EGAS00000000109).</p> <p>probes_passing_QC.txt: List of probes passing quality control; probe sequences mapping to a unique genomic locus, and probe sequences not containing common genomic variation (minor allele frequency >1%), n=29,002. Column names are Array Address ID (probeID), Ensembl ID (ensembl), Gene ID (gene), and Illumina probe ID (ilmn). </p>
Distinct signals of clinal and seasonal allele frequency change at eQTLs in Drosophila melanogaster
<p>Populations of short-lived organisms can respond to spatial and temporal environmental heterogeneity through local adaptation. However, the comparative signals of local adaptation across space and time remains poorly understood. Here, we examined patterns of allele frequency change across a latitudinal cline and between seasons at previously reported expression quantitative trait loci (eQTLs). We divided eQTLs into groups by utilizing differential expression profiles of fly populations collected across latitudinal clines or exposed to different environmental conditions. In general, we find that eQTLs are enriched for clinally varying polymorphisms, and that these eQTLs change in frequency in concordant ways across the cline and in response to starvation and chill-coma. The enrichment of eQTLs among seasonally varying polymorphisms is more subtle, and the direction of allele frequency change at eQTLs appears to be somewhat idiosyncratic. Taken together, we suggest that clinal adaptation at eQTLs is at least partially distinct from seasonal adaptation.</p>
GTEx v8 fine mapping on eQTL and sQTL
<p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication "Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits" (https://doi.org/10.1101/814350).</p> <p># GTEx-GWAS integration: Finemapping</p> <p>This package contains DAP-G results on GTEx v8 eQTL and sQTL data.<br> See ([DAP-G software](https://github.com/xqwen/dap)) for details.<br> We used only European individuals and variants with MAF>0.01, on genes that are annotated as `protein_coding` or `lncRNA`. <br> DAP-G `ld_control` parameter was 0.75.</p> <p>The results were analyzed in [this preprint](https://www.biorxiv.org/content/10.1101/814350v1)</p> <p>## Contents</p> <p>```<br> finemapping/<br> |-- README_finemapping.md<br> |-- dapg_eqtl.tar<br> `-- dapg_sqtl.tar<br> ```<br> Unpack each tarball with a command like `tar -xvpf dapg_sqtl.tar`</p> <p>For every tissue:</p> <p>* `{tissue}.variants_pip.txt.gz` contains the variants' posterior inclusion probabilities at being causal for every gene.<br> * gene: gene id (or intron id)<br> * rank: ranking of the variant according to its PIP (see below)<br> * variant_id: gtex variant id<br> * pip: posterior inclusion probability of the variant in the causal models<br> * log10_abf: approximate Bayes factor (-log10)<br> * cluster_id: id of cluster to which the variant belongs <br> * `{tissue}.models_variants.txt.gz` contains, for every model contemplated by DAPG, the list of variants involved. Most of them have single variant.<br> * `{tissue}.model_summary.txt.gz` contains, for every analized gene, a summary of the modes such as expected number of causal variants<br> * gene: gene id (or intron id)<br> * pes: posterior expected model size (i.e. number of causal variants)<br> * pse_se: standard error of the above<br> * log_nc: dapg undocumented statistic<br> * log10_nc: dapg undocumented statistic<br> * `{tissue}.models.txt.gz` for every analyzed gene:<br> * gene: gene id (or intron id)<br> * model: number (serving as a model name)<br> * n: number of variants (0 for null model)<br> * pp: posterior inclusion probability of the model<br> * ps: posterior score<br> * `{tissue}.clusters.txt.gz` for every analyzed gene:<br> * gene: gene id (or intron id)<br> * cluster: number (serving as cluster name)<br> * n_snps: number of variants in the cluster<br> * pip: posterior inclusion probability<br> * average_r2: average correlation within the cluster<br> * `{tissue}.cluster_correlations.txt.gz`: upper triangular matrix of correlations among clusters </p> <p> </p> <p> </p> <p># Disclaimer</p> <p>The data is provided "as is", and the authors assume no responsibility for errors or omissions. <br> The User assumes the entire risk associated with its use of these data. <br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. <br> The User bears all responsibility in determining whether these data are fit for the User's intended use. </p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. <br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. <br> <br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia's requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>
MTClass: Identification and annotation of multi-phenotype cis-eQTLs using machine learning
<p>This dataset contains the aggregated results from three iterations of MTClass. Brief descriptions of the file names are below:</p> <ul> <li><strong>multi-tissue.zip</strong>: Multi-tissue study containing the 9-tissue, 13 brain tissue, and 48-tissue results from MTClass, MultiPhen, and MANOVA</li> <li><strong>multi-exon.zip</strong>: Multi-exon study containing the multi-exon results from the 13 individual brain tissues (MTClass, MultiPhen, and MANOVA for each tissue)</li> <li><strong>2D_exon_tissue.zip</strong>: Multi-tissue/multi-exon combined eQTL study, done in 9 tissues using multi-layer perceptron. MTClass was only run once on this dataset due to the relatively higher computational burden.</li> <li><strong>PsychENCODE_isoQTL.zip</strong>: Multi-isoform eQTL study, done in human prefrontal cortex using PsychENCODE data.</li> <li><strong>OneK1K_scRNAseq.zip</strong>: Multi-cell-type eQTL study, done using scRNA-seq data from the OneK1K cohort.</li> </ul>
Summary statistics of cell-type specific cis-eQTLs in eight brain cell-types
<p>This dataset contains eQTL summary statistics for all SNPs-gene pairs in 8 major brain cell types (within 1MB window surrounding the TSS of each expressed gene). For each cell type, there is one file per chromosome.</p> <p>Column description:</p> <p>1. Gene_id</p> <p>2. SNP_id</p> <p>3. Distance to TSS</p> <p>4. Nominal p-value</p> <p>5. Beta</p> <p>In addition, a file contains the SNP positions (snp_pos.txt) and tested allele.</p> <p>Update February 2023: The summary statistics from our 'tissue-like' analysis were added (pb[1-22].gz). These file contain eQTL summary statistics after aggregating all reads from all nuclei for each individual (instead of per cell type).</p>
Metadata for various molecular traits included in the eQTL Catalogue
<p>Metadata for various molecular traits included in the eQTL Catalogue. </p> <p>Metadata files for the Leafcutter datasets can be found <a href="https://doi.org/10.5281/zenodo.7850746">here</a>.</p> <p>The tab-separated files contain the following columns:</p> <ul> <li><strong>phenotype_id</strong> - ID of the molecular trait that has been quantified. This can be either the gene ID (RNA-eq eQTLs), probe ID (microarray eQTLs), transcript ID (full-length transcript usage QTLs), splice junction ID (Leafcutter), exon id (exon-level QTLs) or any other molecular trait that has been quantified.</li> <li><strong>quant_id</strong> - Used to quantify relative transcript usage or relative transcriptional event usage (in txrevise). </li> <li><strong>group_id</strong> - Used for transcript usage and splicing phenotypes. Overlapping phenotypes whose relative expression is quantified belong to the same group (e.g. alternative spliced exons form clusters in Leafcutter). QTLTools permutation p-values are calculated accross all phenotypes within a group and only the phenotype with the smallest permutation p-value is reported.</li> <li><strong>gene_id</strong> - Ensembl gene id</li> <li><strong>chromosome</strong> - Chromosome of the gene</li> <li><strong>gene_start</strong> - End coordinate of the gene (GRCh38)</li> <li><strong>gene_end</strong> - Start coordinate of the gene (GRCh38)</li> <li><strong>strand</strong> - Strand of the gene</li> <li><strong>gene_name</strong> - Gene name extracted from Ensembl biomart.</li> <li><strong>gene_type</strong> - Gene type extracted from Ensembl biomart.</li> <li><strong>gene_gc_content</strong> - Percentage GC content of the gene. Extracted from Ensembl biomart and used as a covariate in cqn normalisation. Calculated with bedtools nuc for exons.</li> <li><strong>gene_version</strong> - Ensembl gene version</li> <li><strong>phenotype_pos</strong> - Genomic position used to determine the centre point of the <em>cis-</em>window for QTL mapping. By default this is the beginning of the gene (either gene start or gene end, depending on the strand of the gene).</li> <li><strong>phenotype_length</strong> - (optional) - Length of the gene or exon in basepairs. Required to properly normalise featureCounts quantification results with cqn.</li> </ul>
Genotype of expression quantitative loci (eQTL) analyses for bovine blood and liver
<p>To identify expression quantitative loci (eQTL) operating in bovine blood and liver, 238 animals were genotyped uisng Illumina BovineHD genotyping arrary. For this file minor allele is used as ref as the default in plink.</p>
Distinct signals of clinal and seasonal allele frequency change at eQTLs in Drosophila melanogaster
Open the record for dataset details and reuse information.
Raw microarray gene expression datasets included in the eQTL Catalogue
<p>Raw microarray intensity values for five datasets:</p> <ul> <li>CEDAR</li> <li>Fairfax_2012</li> <li>Fairfax_2014</li> <li>Naranbhai_2015</li> <li>Kasela_2017</li> </ul>
aFCn for independent cis-eQTLs across 49 tissues in the GTEx v8 release
<p>The aFC-n model is the multi-variant generalization of the aFC approach (<a href="https://github.com/secastel/aFC">https://github.com/secastel/aFC</a>), to estimate the <i>cis</i>-regulatory effect size in genes associated with multiple conditionally independent eQTLs. Notably, aFC-n is the first method that allows for predicting genetically regulated gene expression in a haplotype-specific fashion, paving the way for a future class of genome association studies that can systematically incorporate allele dosage effects.</p><p>This dataset is generated using aFC-n software package (<a href="https://doi.org/10.5281/zenodo.8412460">10.5281/zenodo.8412460</a>) to calculate effect sizes for ~half a million independent <i>cis</i>-eQTLs across 49 tissues in the GTEx v8 release as described in the <a href="https://www.biorxiv.org/content/10.1101/2022.01.28.478116v1">manuscript</a>.</p><p>GTEx_v8_aFCn contains:</p><ul><li>gene_id : Ensembl gene ID</li><li>variant_id : eQTL ID associated to the gene</li><li>log2_aFC : The log2 of eQTL effect sizes measured as allelic Fold Change (aFC).</li><li>log2_aFC_min_95_interv : Lower bound 95% confidence interval</li><li>log2_aFC_plus_95_interv: Upper bound 95% confidence interval</li></ul><p>GTEx_v8<i>_</i>aFCn<i>_</i>combined contains:</p><ul><li>gene_id : Ensembl gene ID</li><li>variant_id : eQTL ID associated to the gene</li><li>rest of the columns : The log2 of eQTL effect sizes measured as allelic Fold Change (aFC) in the tissue specified by the column label capped at ±log2(100) <ul><li>nan = The variant is not an eQTL for the tissue / the effect size is not calculated for variants on chrX</li></ul></li></ul><p> </p>
Full eQTL and caQTL summary statistics from RASQUAL
<p>The summary statistics files contain the following columns:</p> <ol> <li>Feature ID</li> <li>Variant ID</li> <li>Chromosome</li> <li>Variant position</li> <li>Allele frequency (not MAF!)</li> <li>HWE Chi-square statistic</li> <li>Imputation quality score (IA)</li> <li>Chi square statistic (2 x log Likelihood ratio)</li> <li>Effect size (Pi)</li> <li>Sequencing/mapping error rate (Delta)</li> <li>Reference allele mapping bias (Phi)</li> <li>Overdispersion</li> <li>No. of feature SNPs</li> <li>No. of tested SNPs</li> <li>Convergence status (0=success)</li> <li>Squared correlation between prior and posterior genotypes (fSNPs)</li> <li>Squared correlation between prior and posterior genotypes (rSNP)</li> </ol>
Full summary statistics from eQTL mapping in heterogeneous differentiating cultures (HDCs)
<p>Full summary statistics, and significant gene-variant pairs, for the cis-eQTL mapping results presented in the manuscript: <em>Cell-type and dynamic state govern genetic regulation of gene expression in heterogeneous differentiating cultures</em>. These include cell-type specific eQTL mapping in 29 cell types, dynamic eQTL mapping in 3 trajectories, and topic eQTL mapping with 10 topics. We additionally include a mash object, which includes posterior effect size estimates and local false sign rate for the analysis of eQTL effects across cell types. See README and manuscript for details.</p>
Trans-eQTL effects on risk of type 1 diabetes: a test of the sparse effector (omnigenic) hypothesis of complex trait genetics (supplementary data)
<p>This repository contains summary-level data generated by performing <a href="https://github.com/molepi-precmed/trans-qtls">Genomewide aggregated trans- effects (GATE) analysis</a> in case-control study of Type 1 Diabetes (T1D).</p>
Table S1. List of relevant studies of gene expression with available eQTL data.
<p><strong>Supplementary Table S1. List of relevant studies of gene expression with available eQTL data.</strong> The table contains information of gene expression eQTL data used in the current study: source name (CEDAR, GTEx, blood eQTL by Westra et al.), tissue name and corresponding article DOI.</p> <p>Part of the article: Williams FMK et al. "Sequence variation at 8q24.21 and risk of back pain"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.