Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

255

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

255 results for “gwas”

Learn how ShareScore rates datasets ↗
zenodo40/100

Effectively controlling for sample relatedness in large-scale GWAS: application to 79 EHR-derived longitudinal traits

<p><span>Sample relatedness is a major confounder in genome-wide association studies (GWAS), potentially leading to inflated type I error rates if not appropriately controlled. A common strategy is to incorporate a random effect related to&nbsp;genetic relatedness matrix (GRM) into regression models. However, this approach is challenging for large-scale GWAS of complex traits, such as longitudinal traits. Here we propose a scalable and accurate analysis framework, SPA<sub>GRM</sub>, which controls for sample relatedness via a precise approximation of the joint distribution of genotypes. SPA<sub>GRM</sub> can utilize GRM-free models and thus is applicable to various trait types and statistical methods, including linear mixed models and generalized estimation equations for longitudinal traits. A hybrid strategy incorporating saddlepoint approximation greatly increases the accuracy to analyze low-frequency and rare genetic variants, especially in unbalanced phenotypic distributions. We also introduce SPA<sub>GRM(CCT)</sub> to aggregate the results following different models via Cauchy combination test. Extensive simulations and real data analyses demonstrated that SPA<sub>GRM</sub> maintains well-controlled type I error rates and SPA<sub>GRM(CCT)</sub> can serve as a broadly effective method. Applying SPA<sub>GRM</sub> to 79 longitudinal traits extracted from</span><span> </span><span>UK Biobank primary care data, we identified 7,463 genetic loci, making a pioneering attempt to conduct GWAS for these traits as longitudinal traits.</span></p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Meta-GWAS for age-related hearing impairement

<p>The dataset comprises summary statistics from the meta-GWAS of 17 studies on age-related hearing impairement. The dataset accompanies the following paper:</p> <p><strong>Genome-wide association meta-analysis identifies 48 risk variants and highlights the role of the stria vascularis in age-related hearing impairment</strong></p> <p>Please cite the paper if using this dataset.</p> <p>Phenotype of ARHI was established using ICD diagnoses and self-reported hearing loss. The study comprised 148,152 cases and 575,472 controls or European ancestry.&nbsp;Adult male and female participants were included from the following 17 population-based cohort studies: Age, Genes/Environment Susceptibility - Reykjavik (AGES), the Danish Twin Registry (DTR), the Estonian Genome Center at the University of Tartu (EGCUT), FinnGen, Framingham Heart Study (FHS), Health Aging and Body Composition (HABC), Italian Network of Genetic Isolates - Friuli Venezia Giulia (INGI-FVG), the Rotterdam Study (RS, cohorts 1 - 3), the Salus in Apulia study (SA; formerly known as Great Age study), Screening Across the Lifespan Twin (SALT&nbsp;and SALTY - young), Screening Twin Adults: Genes and Environment (STAGE), TwinsUK, UK Biobank (UKBB), and the Women&rsquo;s Genome Health Study (WGHS). &nbsp;</p> <p>UK Biobank data have been used under project #11516.</p> <p>Individual GWASs have been QC&#39;d and harmonyzed using EasyQC followed by fixed-effects IVW meta-analysis using METAL. The dataset includes the results of meta-analysis for n = 8,244,938 SNV with MAF &gt;0.001 and present in at least 9 cohorts.<br> <br> <strong>Dataset columns:</strong></p> <p>SNP, rsID</p> <p>CHR, chromosome</p> <p>BP, genomic position (hg19)</p> <p>Allele1, effect allele</p> <p>Allele2, other allele</p> <p>Freq1, mean frequency of Allele1</p> <p>FreqSE, standard error of Freq1</p> <p>MinFreq, minimal frequency of Allele1 in the study cohorts</p> <p>MaxFreq, maximal frequency of Allele1 in the study cohorts</p> <p>Effect, effect size from the meta-analysis for Allele1</p> <p>StdErr, standard error of Effect</p> <p>P.value, corresponding p-value for meta-analysis</p> <p>Direction, direction of effects in individual studies</p> <p>HetISq, I<sup>2</sup>&nbsp;statistic for heterogeneity between studies</p> <p>HetChiSq, chi<sup>2</sup>&nbsp;statistic for heterogeneity between studies</p> <p>HetDf, degrees of freedom for the&nbsp;chi<sup>2</sup>&nbsp;statistic</p> <p>HetPval, p-value for heterogeneity between studies</p> <p>N, summary sample size</p> <p><strong>Dataset columns description:</strong></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Seventeen studies included:</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Manhattan and QQ plots of GWAS on salt stress responses in root system architecture parameters of wild tomato (S. pimpinellifolium)

<p>The population of +/2 200 accessions of wild tomato was screened with the protocol described&nbsp;<a href="https://www.protocols.io/view/studying-root-system-architecture-changes-in-tomat-2mqgc5w">here</a>&nbsp;with the only exception that the plants were transferred 4 days after germination (rather than 3 - described in the protocol). The images were analyzed using the&nbsp;<a href="https://smartroot.github.io/">SmartRoot</a>&nbsp;for days 0, 1, 2, 3, and 4 after transfer to treatment plates (0 or 100 mM NaCl, 1/4 MS, 0.5% sucrose, 0.1% MES, 1% Dashin agar). The data analysis was performed as described&nbsp;<a href="https://rpubs.com/mjulkowska/BIGpimp_RSA_salt">here</a>, while the pareto front calculations were done according to Chandrasekhar &amp; Julkowska paper (<a href="https://www.biorxiv.org/content/10.1101/2021.08.12.456185v1">preprint here</a>). The GWAS was performed using the ASReml script similar to&nbsp;<a href="https://onlinelibrary.wiley.com/doi/10.1111/tpj.15310">Awlia et al. (2021)</a>. The raw GWAS outputs can be found <a href="https://zenodo.org/badge/DOI/10.5281/zenodo.5856310.svg">here</a>. This dataset represents Manhattan plots and QQ plots made out of the data.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

FUNCTIONAL ANALYSIS OF LITTING SIZE AND NUMBER OF TEATS IN PIGS: FROM GWAS TO POST-GWAS

<p>Reproductive traits, such as number of teats and litter size, are essential for animal breeding programs due to the importance for the production chain, since they influence the maternal ability of sow and can affect the number of weaned piglets. Our objective was to identify candidate genes associated with reproductive traits in pigs, using GWAS data from a systematic review combined with sequencing data, to build networks of biological processes and TFs (transcription factors) from the identified genes, in order to highlight the most candidate genes for litter size and number of teats. In the systematic review only peer-reviewed articles were used, with descriptors related to the evaluated traits, and selected based on eligibility criteria. Fourteen papers were selected and classified into groups for functional analysis of gene networks with 2,077 candidate genes identified. After combining with the list of genes presenting known structural variants in the 5&#39;UTR and/or coding region, 306 genes remained to be used to build the networks of biological processes and TFs genes, highlighting processes associated with litter size (e.g., ionotropic glutamate receptor signaling pathway and blastocyte growth) and number of teats (e.g., growth hormone receptor, regulation of the BMP - Bone Morphogenetic Proteins signaling pathway and blood vessel proliferation). Two most candidate genes for litter size trait (<em>GRID2 </em>and <em>PALB2</em>) and six most candidate genes for number of teats (<em>GHR, IFT80,</em> <em>FSTL3, SKOR1, SMURF1</em> and <em>AKT3</em>) were prioritized. TFs associated with candidate genes were also identified for litter size (<em>PALB2</em> and <em>GRID2</em>) and number of teats (<em>RIN, LTBP2</em> and <em>COL6A6</em>). Thus, it is suggested that the most candidate genes and TFs presented in this study may play an important role in the traits studied, being important for genetic studies and animal breeding.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

GWAS summary statistics and code for "Sequence variants affecting voice pitch in humans"

<p>Contents:&nbsp;GWAS summary statistics for voice pitch (median F0 in reading) and code for acoustic analysis</p> <p>Please refer to the corresponding publication:</p> <p>Gisladottir et al. Sequence variants affecting voice pitch in humans.&nbsp;<em>Science Advances</em></p> <p>The GWAS summary statistics is also available at:&nbsp;https://www.decode.com/summarydata/</p> <p>The code for acoustic analysis is also available at:&nbsp;https://github.com/cadia-lvl/deCODE</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

QR GWAS summary statistics for 39 quantitative traits in the UK Biobank

<p>Quantile regression (QR) GWAS summary statistics from the study "Genome-wide discovery for biomarkers using quantile regression at biobank scale". The preprint is available at <a href="https://doi.org/10.1101/2023.06.05.543699" target="_blank" rel="noopener">https://doi.org/10.1101/2023.06.05.543699</a>.&nbsp;</p> <p><strong>List of traits</strong></p> <p>A comma-delimited text file, QRGWAS.Traits_n39.csv, includes the list of 39 quantitative traits from the UK Biobank reported in the QR GWAS analyses above.</p> <p><strong>Summary statistics</strong></p> <p>The tab-delimited text files are QR GWAS summary statistics, which are bgzip compressed (.tsv.gz files) and tabix indexed (.tbi files).</p> <ul> <li>Column "CHR": chromosome</li> <li>Column "POS": based pair position</li> <li>Column "ID": variant ID</li> <li>Column "REF": non-effect allele</li> <li>Column "ALT": effect allele tested in GWAS</li> <li>Column "EAF": frequency of the effect allele</li> <li>Column "N": sample size</li> <li>Column "P_QR": integrated p-value of the quantile regression (QR) model across multiple quantile levels.</li> <li>Column "P_LR": p-value of the linear regression (LR) association statistic</li> <li>Columns from "P_Q10" to "P_Q90": quantile-specific QR p-value for the quantile levels 0.1, 0.2, ..., 0.9 (10th, 20th, ..., 90th quantiles).</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo40/100

GWAS and GTEx QTL integration

<p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication &quot;Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits&quot; (https://doi.org/10.1101/814350).</p> <p>&nbsp;</p> <pre><em># GTEx GWAS integration </em> This package contains the application of several GWAS-QTL integration methods. The results were analyzed in [this preprint](<em>https://www.biorxiv.org/content/10.1101/814350v1</em>) about GTEx v8 application to several GWAS traits. <em>``` </em><em>. </em><em>|-- colocalization </em><em>| |-- coloc </em><em>| | `-- coloc_enloc_priors_eqtl.tar.gz </em><em>| |-- enloc </em><em>| | |-- enloc_eqtl_eur.tar.gz </em><em>| | `-- enloc_sqtl_eur.tar.gz </em><em>| `-- eur_ld.bed.gz </em><em>|-- prediction_models </em><em>| |-- gtex_v8_expression_mashr_snp_smultixcan_covariance.txt.gz </em><em>| |-- gtex_v8_splicing_mashr_snp_smultixcan_covariance.txt.gz </em><em>| |-- mashr_eqtl.tar </em><em>| `-- mashr_sqtl.tar </em><em>|-- smr </em><em>| |-- SMR_gtex_v8_README.txt </em><em>| `-- SMRresults_GTEx_v8_peQTL5e-08.tar.gz </em><em>|-- smultixcan </em><em>| |-- smultixcan_eqtl.tar.gz </em><em>| `-- smultixcan_sqtl.tar.gz </em><em>`-- spredixcan </em><em> |-- spredixcan_eqtl.tar.gz </em><em> `-- spredixcan_sqtl.tar.gz </em> <em> ``` </em><em> </em>You can uncompress gzipped tarball packages <em>`*.tar.gz` </em>in a UNIX command line with an instruction such as: <em>```bash </em><em>tar -xzvpf smultixcan_eqtl.tar.gz </em><em>``` </em>, and the tar packages (<em>`*.tar`</em>) with an analogous instruction: <em>```bash </em><em>tar -xvpf mashr_eqtl.tar </em><em>``` </em> <em>## Preliminaries </em> <strong>**</strong>Finemapping<strong>** </strong>results are contained in a separate release due to size constraints. GWAS summary statistics for 114 traits were harmonized and imputed to GTEx v8 variants with MAF&gt;0.01 using only european samples. (summary imputation software [here](<em>https://github.com/hakyimlab/summary-gwas-imputation</em>)). Some of the following analyses used the full set of 114 traits, while some focused only on 87 traits whose imputed associations showed no deflation (the imputation algorithm is conservative, and studies with too few available variants have a depleted distribution of association p-values after imputation). The harmonized and imputed GWAS summary statistics are contained in a separate release due to size constraints. For completeness&#39; sake, the imputed summary statistics look like: <em>``` </em><em>variant_id panel_variant_id chromosome position effect_allele non_effect_allele current_build frequency sample_size zscore pvalue effect_size standard_error imputation_status n_cases </em><em>rs554008981 chr1_13550_G_A_b38 chr1 13550 A G hg38 0.017316017316017316 336474 -2.2919929353647097 0.021906050841240293 NA NA imputed NA </em><em>rs201055865 chr1_14671_G_C_b38 chr1 14671 C G hg38 0.012987012987012988 336474 -0.9559192804440632 0.33911301727494103 NA NA imputed NA </em><em>... </em><em>``` </em> The GWAS were split in approximately independent LD regions (Berisa-Pickrell)/ GWAS regions are defined in <em>`eur_ld.bed.gz` </em>(note that a few of them are ill-defined in hg38 and where ignored; only completely defined regions were used). <em>## Colocalization </em> <em>### Enloc </em> ENLOC ([see fotware here](<em>https://github.com/xqwen/integrative</em>)) was run for sQTLs and eQTLs using individuals of european ancestry and DAP-G QTL enrichment results on 87 traits. Result files are included in <em>`enloc_eqtl_eur.tar.gz` </em>and <em>`enloc_sqtl_eur.tar.gz` </em>Each file contains a particular tissue-trait combination. Each row details colocalization between a GWAS region (Berisa-Pickrell) and gene&#39;s or intron&#39;s cis-window. A region might overlap multiple genes/introns or viceversa. Each ENLOC file contains the following columns: <strong>* </strong>gwas_locus: GWAS LD region <strong>* </strong>molecular_qtl_trait: gene or intron <strong>* </strong>locus_gwas_pip: posterior inclusion probability of variants in the GWAS LD region <strong>* </strong>locus_rcp: regional colocalization probability (main colocalization measure) <strong>* </strong>lead_coloc_SNP: snp with highest RCP <strong>* </strong>lead_snp_rcp: rcp of the lead coloc snp <em>### Coloc </em> Coloc ([see software here](<em>https://cran.r-project.org/web/packages/coloc/index.html</em>)) was run using prior probabilities estimated from QTL enrichment of GWAS variants (computed via ENLOC). Results for eQTL are available in <em>`coloc_enloc_priors_eqtl.tar.gz`</em>. Each file contains results for a trait-tissue combination. Columns are: <strong>* </strong>gene_id: gene or intron id <strong>* </strong>p0: probability that neither QTL nor GWAS contain a causal variant <strong>* </strong>p1: probability that only GWAS contains a causal variant <strong>* </strong>p2: probability that only QTL has a causal variant <strong>* </strong>p3: probability that GWAS and QTL have a causal variant and it&#39;s distinct <strong>* </strong>p4: probability that GWAS and QTL have a causal variant and it&#39;s the same (main colocalization measure) <em>## PrediXcan </em> <em>`mashr_eqtl.tar` </em>and <em>`mashr_sqtl.tar` </em>contain prediction models (trained on expression or splicing data respectively, for 49 GTEx tissues) and LD compilations to be used with PrediXcan, S-PrediXcan, MultiXcan and S-MultiXcan. For every tissue, the <em>`mashr_{tissue}.db` </em>file is a SQLite file with the prediction model definitions. <em>`mashr_{tissue}.txt.gz` </em>is a gzipped-text file with the upper triangular matrices of covariance between snps within a gene/intron prediction model. Many variants in these models don&#39;t have an rsid. To fully leverage the information in these models, it is advised to at least harmonize to GTEx variants, and if possible impute as we did [here](<em>https://github.com/hakyimlab/summary-gwas-imputation</em>). <em>### S-PrediXcan </em> S-PrediXcan was run for the 114 harmonized and imputed traits, on eQTL and sQTL mashr prediction models. All of the GWAS traits had the same format, so that the following format parameters were used with S-PrediXcan: <em>``` </em><em>--snp_column panel_variant_id --effect_allele_column effect_allele --non_effect_allele_column non_effect_allele --zscore_column zscore \ </em><em>--keep_non_rsid --additional_output --model_db_snp_key varID \ </em><em>``` </em> Each file is a CSV, with each row containing a gene/intron association at a given trait-tissue combination: <strong>* </strong>gene: ENSEMBLE ID or intron id <strong>* </strong>gene_name: HUGO name or intron id <strong>* </strong>zscore: predicted association z-score <strong>* </strong>effect_size: estimated effect size <strong>* </strong>pvalue: association p-value <strong>* </strong>var_g: estimated variance of predicted expression or splicing <strong>* </strong>pred_perf_r2: prediction model cross-validated performance <strong>* </strong>pred_perf_pval: prediction model cross-validated performance <strong>* </strong>pred_perf_qval: deprecated, empty field left for compatibility <strong>* </strong>n_snps_used: number of snps in the intersection of GWAS and model <strong>* </strong>n_snps_in_cov: number of snps in the LD compilation <strong>* </strong>n_snps_in_model: number of snps in the model <strong>* </strong>best_gwas_p: smallest p-value acros GWAS snps used in this model <strong>* </strong>largest_weight: largest prediction model weight <em>### S-Multixcan </em> S-MultiXcan results were generated from the above S-PrediXcan results. Each fiel contains multi-tissue associations for a given trait: <strong>* </strong>gene: ENSEMBLE ID or intron id <strong>* </strong>gene_name: HUGO name or intron id <strong>* </strong>pvalue: multi-tissue association p-value <strong>* </strong>n: number of models avialble for this gene/intron <strong>* </strong>n_indep: number of independent components of variation in predicted expression/splicing (surviving principal components) <strong>* </strong>p_i_best: highest single-tissue p-value (S-PrediXcan) <strong>* </strong>t_i_best: tissue of highest p-value <strong>* </strong>p_i_worst: lowest single-tissue p-value (S-PrediXcan) <strong>* </strong>t_i_worst: tissue of lowest p-value <strong>* </strong>eigen_max: maximum eigenvalue of SVD <strong>* </strong>eigen_min: minimum eigenvalue of SVD <strong>* </strong>eigen_min_kept: smallest eigenvalue retained after discarding smallest variations <strong>* </strong>z_min: minimum single-tissue z-score <strong>* </strong>z_max: maximum single-tissue z-score <strong>* </strong>z_mean: mean single-tissue zscre <strong>* </strong>z_sd: standard deviation of the single-tissue z-scores <strong>* </strong>tmi: trace of M * M_i where M is predicted expression/splicing covariance across tissues for a gene, and M_i is its SVD pseudo-inverse <strong>* </strong>status: computation status, 0 if no errors <em>## SMR </em> See <em>`SMR_gtex_v8_README.txt` </em>for details.</pre> <p>&nbsp;</p> <p>&nbsp;</p> <p># Disclaimer</p> <p>The data is provided &quot;as is&quot;, and the authors assume no responsibility for errors or omissions. &nbsp;<br> The User assumes the entire risk associated with its use of these data. &nbsp;<br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. &nbsp;<br> The User bears all responsibility in determining whether these data are fit for the User&#39;s intended use. &nbsp;</p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. &nbsp;<br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. &nbsp;<br> &nbsp;<br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia&#39;s requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Haplotype analysis of GWAS candidates identified for root:shoot ratio changes under salt stress in Arabidopsis

<p>The haplotype analysis was performed on 7 loci identified through GWAS by Magdalena Julkowska, while she was a PostDoc at KAUST, Saudi Arabia, workin in the lab of Dr. Mark Tester.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Screen of Arabidopsis mutants homologous to Cowpea GWAS peaks

<p>This dataset was collected for T-DNA insertion mutants of Arabidopsis that were selected based on the sequence homology with the candidate genes in cowpea (<em>Vigna unguiculata</em>) identified through GWAS for drought induced changes in growth, evapotranspiration and photosynthetic efficiency. The T-DNA insertion lines were germinated on agar plates and transferred to soil - where the seedlings were exposed to drought stress (10% soil water-holding capacity).&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Cowpea GWAS drought stress in early vegetative stage

<p>The GWAS outputs for Cowpea (<em>Vigna unguiculata</em>) responses to drought stress at early vegetative stress. The cowpea seedlings (miniCore population) were exposed to drought stress at 17 days after germination using the weight of the pot and AAWEsmo device, developed in Julkowska Lab, Boyce Thompson Institute. The seedlings were kept at 60 and 10% of soil water holding capacity for 2 weeks and the data on cowpea shoot size, evapotranspiration and photosystem II efficiency was collected. The data was assembled and curated (https://rpubs.com/mjulkowska/Cowpea2022alltraits), and subsequently used for GWAS. The GWAS data was analyzed, and the most interesting associations were selected (https://rpubs.com/mjulkowska/Cowpea2022Gwas).&nbsp;</p> <p>The raw data was collected by Hayley Sussman, with help of Olga Khmelnitsky, while GWAS was performed by Magdalena Julkowska, using ASReml script developed by Arthur Korte (https://github.com/arthurkorte/GWAS), adapted for cowpea.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Solanum pimpinellifolium input data collected from the TPA - used for GWAS

<p>The GWAS input data used for mapping the candidate genes in S. pimpinellifolium collection, exposed to salt stress in The Plant Accelerator experiment (TPA). The phenotypic data was collected in an experiment was performed by Mitchell Morton while being a PhD student in the group of Prof. Mark Tester at KAUST. The genotypic data was collected by Magdalena Julkowska, and used for sequencing. The SNPs were called by Elodie Ray, and subsequently curated by Magdalena Julkowska for GWAS analysis. The GWAS was conducted by Magdalena Julkowska.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

GWAS analysis for tolerance to heat stress and milk production, in subtropical Egyptian goats

<p>The conservation of local Egyptian goat genotypes has been a national program since 2009. Our study investigated different subtropical Egyptian goat populations (367 does) from different harsh ecological zones(hot Upper Egypt, Coastal Zone of Western Desert, and Wahati Desert Oasis) to identify genes associated with heat stress. We examined the physiological response of animals that were exposed to simulated summer grazing conditions, and physiological parameters including respiration rate, gas volume, rectal temperature, and skin temperature. Temperature Humidity Index ranged from 98.6 to 109.3, indicating that the animals were under severe heat stress. Results showed significant differences between the populations in their tolerance to heat stress, with Saidi goats being the most adapted to hot-dry conditions. Respiration rate was found to be the most reliable physiological trait for differentiating between animals in their tolerance to heat stress. The GWAS analysis involved 157 genotypes and 54,032 marker-SNPs, revealing 90 SNPs associated with heat stress and 70 SNPs associated with milk production. Chromosome 1 had the highest number of SNPs associated with heat-stress traits, while chromosome 4 exhibited the largest number of significant SNPs associated with milk production. Several genes were associated with heat stress response and tolerance in goats. These genes were categorized into three main groups: heat stress response, stress response, and reproduction, and feed intake. The first group includes USP54, KDM6A, and ETNPPL, which have multiple functions, such as steroid biosynthesis, metabolism, and stress response, and are also involved in key biological processes such as reproduction, immune response, and metabolism. The second group, which is mostly associated with immune response, includes GLTSCR2 and NAALADL2. Finally, a group of genes that may control animal feed intakes, including TRPM3 and ZBTB8A. Additionally, several genes, such as FHIT, GALNT18, RAPGEF5, and RBFOX1, linked to milk production, are known to be linked fertility, and fatty acid composition. These findings provide insights into the potential roles of these genes in heat stress response and tolerance in goats, and the genetic basis for improved milk production and animal welfare and can be used as selection markers in the ongoing breeding programs.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

GWAS results of selected binarized neurocognitive task performances from Philadelphia Neurodevelopmental Cohort

<p>This archive contains analysis results associated with the publication:</p> <p>Shraddha Pai, Shirley Hui, Philipp Weber, Soumil Narayan, Owen Whitley, Peipei Li, Viviane Labrie, Jan Baumbach, Anne L Wheeler, Gary D Bader. (2023). Multi-scale systems genomics analysis predicts pathways, cell types and drug targets involved in normative variation in peri-adolescent human cognition. Cerebral Cortex.</p> <p>-----------------------------------------------------------<br> *loco.mlma files: &nbsp;Univariate GWAS summary statistics for binarized performance of computerized neurocognitive tests in the Philadelphia Neurodevelopmental Cohorts. GWAS was performed on individuals ascertained to be of European genetic ancestry. For methods, please refer to the publication above.</p> <p>Phenotype codes:<br> lnb_tp2 &nbsp; &nbsp;Working memory &nbsp; &nbsp;LNB: Number of Correct Responses to 2-Back Trials (TP)<br> pcpt_t_tp &nbsp; &nbsp;Attention &nbsp; &nbsp;PCPT: Total of Correct Responses to Number Trials (TP) and Letter Trials (TP)<br> peit_cr &nbsp; &nbsp;Emotion identification &nbsp; &nbsp;PEIT: Total Correct Responses for All Test Trials, by genus<br> pfmt_ifac_tot &nbsp; &nbsp;Face memory &nbsp; &nbsp;PFMT: Total Correct Responses for All Test Trials<br> plot_tc &nbsp; &nbsp;Spatial reasoning &nbsp; &nbsp;PLOT: Total Correct Responses for All Test Trials, by genus<br> pmat_pc &nbsp; &nbsp;Nonverbal reasoning &nbsp; &nbsp;PMAT: Percent of Correct Responses for All Test Trials, by genus<br> pvrt_cr &nbsp; &nbsp;Verbal reasoning &nbsp; &nbsp;PVRT: Total Correct Responses for All Test Trials, by genus<br> pwmt_kiwrd_tot &nbsp; &nbsp;Word memory &nbsp; &nbsp;PWMT: Total Correct Responses for All Test Trials<br> volt_svt &nbsp; &nbsp;Object memory &nbsp; &nbsp;VOLT: Total Correct Responses for All Test Trial</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

GWAS Summary Statistics of 7 chronic pain types

<p>Summary Statistics obtained from .fastGWA files after running GWAS analyses using GCTA package&nbsp;for 7 chronic pain types: back, neck, hip, knee, facial and abdominal pain as well as&nbsp;chronic headaches.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Data from the manuscript 'Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context'

<p>Data sets from the manuscript &#39;Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context&#39;. These include the test statistics from analyses of real and simulated data, and the data used to generate the figures relating to the goodness-of-fit of the generalised extreme value distribution to the GPS test statistics under the null. Please see the enclosed README for more details.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Input data for running a GWAS on penicillin resistance in Streptococcus pneumoniae

<p>Results from running the pyseer tutorial at&nbsp;https://pyseer.readthedocs.io/en/master/tutorial.html</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Meta-GWAS of Age at Type 1 Diabetes Diagnosis

<p>7,923 subjects with type 1 diabetes from five studies (SDRNT1BIO, DCCT, CACTI, WESDR and EDC) were included in this analysis. This dataset includes summary stats for 8,154,711 autosomal SNPs.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Meta-GWAS of C-peptide in Type 1 Diabetes

<p>7,252 subjects with type 1 diabetes from four studies (SDRNT1BIO, DCCT, CACTI and WESDR) were included in this analysis. This dataset includes summary stats for 8,150,646 autosomal SNPs.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

GWAS summary statistics imputation support data and integration with PrediXcan MASHR

<p># GWAS summary statistics imputation, integration with PrediXcan MASHR-M</p> <p>&nbsp;</p> <p>The file `sample_data.tar` contains all necessary files to perform imputation of GWAS summary statistics to the GTEx v8 QTL data set.</p> <p>It includes 1000 Genomes individuals&#39; genotypes as reference panel.</p> <p>The `.tar` archive, upon uncompression, contains the following folder structure:</p> <p>```</p> <p>data<br> |-- coordinate_map<br> |-- gwas<br> |-- liftover<br> |-- models<br> |&nbsp;&nbsp; |-- eqtl<br> |&nbsp;&nbsp; |&nbsp;&nbsp; `-- mashr<br> |&nbsp;&nbsp; `-- sqtl<br> |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; `-- mashr<br> |-- reference_panel_1000G<br> `-- ucsc</p> <p>```</p> <p>&nbsp;</p> <p>`data/eur_ld.bed.gz` contains definitions of approximately independent LD-regions in hg38 (Berisa-Pickrell regions, lifted over)</p> <p>`data/gtex_v8_eur_filtered_maf0.01_monoallelic_variants.txt.gz` is a snp annotation file, listing all GTEx v8 variants with MAF&gt;0.01 in europeans.</p> <p>`data/coordinate_map` contains precomputed mapping tables that MetaXcan tools can use to convert GWAS&#39; genomic coordinates in GWAS between genome assemblies.</p> <p>`data/gwas` contains a sample GWAS file for the purposes of a tutorial (data obtained from Nikpay et al (Nat Gen 2016) https://www.ncbi.nlm.nih.gov/pubmed/26343387</p> <p>`data/liftover` contains Liftover chains to map coordinates between human genome assemblies (used by full harmonization tools)</p> <p>`data/models` contains PrediXcan MASHR-M models, and cross-tissue S-MultiXcan LD compilation, from eQTL and sQTL.</p> <p>`data/reference_panel_1000G` contains 1000G hg38 genotypes, in parquet format, to be used by imputation tools.</p> <p>`data/ucsc` contains genomic coordinates of rsids in hg17, hg18 and hg19. You can use these to add chromosome and start position information to a GWAS based on its rsids. (column `end` is not used)</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Annotated genes harboring major effect markers (R2 ≥ 15%). Highlighted in green are genes annotated from Rhodes et al. 2014,2017, in orange genes annotated as similar to Peroxidase, in yellow new annotations from sorghum genome in Atlas. In the first three columns start and stop position on the sorghum genome and transcript name, followed by the nearest marker name and the distance of the gene from the nearest marker, then a column where are shown the GWAS methods and target traits for which the linked SNP was significant, the last column shows the category of the genes.

<p><strong>We conducted a comprehensive genomics study to map genomic loci determining the production of antioxidants in sorghum grains. Encouraging results were obtained and published in peer-reviewed article with impact factor (https://doi.org/10.1371/journal.pone.0225979). Annotated genes harboring major effect markers (R<sup>2</sup> &ge; 15%) were identified and will be of worldwide interest. </strong></p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record