Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
215
datasets available to search
ShareScore release 0.9.0
Dataset results
215 results for “Summary Statistics”
Genome-wide association summary statistics for human blood plasma glycome
<p>The dataset contains results of genome-wide association study of human blood plasma glycome. The 113 files contain association summary statistics for 113 glycome traits, of which 36 were directly measured by UPLC technology and 77 were derived glycome traits. Description of each glycome trait can be found in the <strong>Additional notes</strong> section. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Sharapov, S. Z., Tsepilov, Y. A., Klaric, L., Mangino, M., Thareja, G., Shadrina, A. S., … Aulchenko, Y. (2019). Defining the genetic control of human blood plasma N-glycome using genome-wide association study. <em>Human Molecular Genetics</em>. http://doi.org/10.1093/hmg/ddz054</li> <li>Sodbo Sharapov, Yakov Tsepilov, Lucija Klaric, Massimo Mangino, Gaurav Thareja, Mirna Simurina, Concetta Dagostino, Julia Dmitrieva, Marija Vilaj, FranoVuckovic, Tamara Pavic, Jerko Stambuk, Irena Trbojevic-Akmacic, Jasminka Kristic, Jelena Simunovic, Ana Momcilovic, Harry Campbell, Malcolm Dunlop, Susan Farrington, Maria Pucic-Bakovic, Christian Gieger, Massimo Allegri, Edouard Louis, Michel Georges, Karsten Suhre, Tim Spector, Frances MK Williams, Gordan Lauc, Yurii Aulchenko. (2018). Genome-wide association summary statistics for human blood plasma glycome (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1298406</li> </ol> <p><strong>Funding</strong></p> <p>This work was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736) and by the European Structural and Investments funding for the "Croatian National Centre of Research Excellence in Personalized Healthcare" (contract #KK.01.1.1.01.0010).</p> <p>The work of SSh was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.</p> <p>The work of YT was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).</p> <p>Karsten Suhre and Gaurav Thareja are supported by ‘Biomedical Research Program’ funds at Weill Cornell Medicine - Qatar, a program funded by the Qatar Foundation. We thank all staff at Weill Cornell Medicine - Qatar and Hamad Medical Corporation, and especially all study participants who made the QMDiab study possible.</p> <p>The SOCCS study was supported by grants from Cancer Research UK (C348/A3758, C348/A8896, C348/ A18927); Scottish Government Chief Scientist Office (K/OPR/2/2/D333, CZB/4/94); Medical Research Council (G0000657-53203, MR/K018647/1); Centre Grant from CORE as part of the Digestive Cancer Campaign (<a href="http://www.corecharity.org.uk">http://www.corecharity.org.uk</a>).</p> <p>TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility and Biomedical Research Centre based at Guy’s and St Thomas’ NHS Foundation Trust in partnership with King’s College London.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNP: SNP rsID</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>OTHER_ALLELE: reference allele (coded as "0")</li> <li>EFFECT_ALLELE: effective allele (coded as "1")</li> <li>EAF: effective allele frequency </li> <li>N: sample size</li> <li>BETA: effect size of effective allele</li> <li>SE: standard error of effect size</li> <li>PVAL: P-value of association (without GC correction)</li> <li>IMPUTATION: imputation quality</li> </ol>
Genome-wide association summary statistics for human healthspan
<p>The dataset contains genome-wide association summary statistics computed for heathspan. The UKB sub-population of 300,447 genetically Caucasian, British individuals were analyzed. For more details see [1].</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Zenin, A., Tsepilov, Y., Sharapov, S., Getmantsev, E., Menshikov, L. I., Fedichev, P. O., & Aulchenko, Y. (2019). Identification of 12 genetic loci associated with human healthspan. <em>Communications Biology</em>, <em>2</em>(1), 41. http://doi.org/10.1038/s42003-019-0290-0</li> <li>Aleksandr Zenin, Yakov Tsepilov, Sodbo Sharapov, Evgeny Getmantsev, Leonid Menshikov, Peter Fedichev, & Yurii Aulchenko. (2018). Genome-wide association summary statistics for human healthspan (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1302861</li> </ol> <p><strong>Funding</strong></p> <p>The work was supported by Russian Ministry of Science and Education under 5-100 Excellence Programme. <br> The work was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017). <br> This research has been conducted using the UK Biobank Resource. <br> The study has been funded by Gero LLC.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNPID - SNP rsID</li> <li>chr - chromosome</li> <li>pos - position (GRCh37 build / hg19)</li> <li>EA - effective allele (coded as "1")</li> <li>RA - reference allele (coded as "0")</li> <li>EAF - effective allele frequency</li> <li>beta - effect size of effective allele</li> <li>se - standard error of effect size</li> <li>Z - Z-value of association</li> <li>-log10(p-value) - minus log10(P-value) of association</li> </ol>
Genome-wide association summary statistics for back pain
<p>The dataset contains results of a genome-wide association study of back pain. Two files contain association summary statistics for discovery GWAS based on the analysis of 350,000 white British individuals from the UK Biobank and meta-analysis GWAS based on the meta-analysis of the same 350,000 individuals and additional 103,862 individuals of European Ancestry from the UK biobank (total N = 453,862). The phenotype of back pain was defined by the answer provided by the UK biobank participants to the following question: "Pain type(s) experienced in last month". Those who reported “Back pain”, were considered as cases, all the rest were considered as controls. Individuals who did not reply or replied: "Prefer not to answer" or "Pain all over the body" were excluded. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org/">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Insight into the genetic architecture of back pain and its risk factors from a study of 509,000 individuals. Freidin, Maxim; Tsepilov, Yakov; Palmer, Melody; Karssen, Lennart; Suri, Pradeep; Aulchenko, Yurii; Williams, Frances MK,# CHARGE Musculoskeletal Working Group. PAIN: February 06, 2019 - Volume Articles in Press - Issue - p<br> doi: 10.1097/j.pain.0000000000001514</li> <li>Maxim B Freidin, Yakov A Tsepilov, Melody Palmer, Lennart Karssen, CHARGE Musculoskeletal Working Group, Pradeep Suri, … Frances MK Williams. (2018). Genome-wide association summary statistics for back pain (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1319332</li> </ol> <p><strong>Funding:</strong></p> <p>This study was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736). <br> The research has been conducted using the UK Biobank Resource (project # 18219).</p> <p>The development of software implementing SMR/HEIDI test and database for GWAS results was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Program”.</p> <p>Dr. Suri’s time for this work was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development Service. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p>Dr. Tsepilov’s time for this work was supported in part by the Russian Ministry of Science and Education under the 5-100 Excellence Program.</p> <p><strong>Column headers - discovery (350K)</strong></p> <ol> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>ID: SNP rsID</li> <li>REF: reference allele (coded as "0")</li> <li>ALT: effect allele (coded as "1")</li> <li>CASE_ALLELE_CT: allele observation count in cases</li> <li>CTRL_ALLELE_CT: allele observation count in controls</li> <li>ALT_FREQ: effect allele frequency </li> <li>MACH_R2: imputation quality</li> <li>TEST: model of association test (additive)</li> <li>OBS_CT: sample size</li> <li>BETA: effect size of effect allele</li> <li>SE: standard error of effect size</li> <li>T_STAT: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> <li>MAF: minor allele frequency</li> </ol> <p><strong>Column headers - meta-analysis (450K)</strong></p> <ol> <li>MarkerName: SNP rsID</li> <li>Allele1: effect allele (coded as "1")</li> <li>Allele2: reference allele (coded as "0")</li> <li>Freq1: effect allele frequency</li> <li>FreqSE: standard error of effect allele frequency</li> <li>Effect: effect size of effect allele</li> <li>StdErr: standard error of effect size</li> <li>P-value: P-value of association (without GC correction)</li> <li>Direction: sign of effect in discovery and replication samples</li> <li>n_total: Total sample size</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>MACH_R2_discovery: imputation quality in discovery sample</li> </ol>
Seasonal and annual summary statistics of urbanization, vegetation, land surface temperature, and bioclimatic variables derived from remotely-sensed imagery in areas surrounding long-term bird monitoring locations in the greater Phoenix, Arizona, USA metropolitan area (1997-2023)
This data package consists of 26 years (1998-2023) of environmental data and 22 years (2000-2022) years of bioclimatic data associated with CAP-LTER long-term point-count bird censusing sites (https://doi.org/10.6073/pasta/4777d7f0a899f506d6d4f9b5d535ba09), temporally aggregated by year and by four meteorological seasons (Winter, Spring, Summer, Fall). The environmental variables include land surface temperature (LST), three spectral indices of vegetation and water – the normalized difference vegetation index (NDVI), the soil adjusted vegetation index (SAVI), and modified normalized difference water index (MNDWI) – and four spectral indices of impervious surface/urbanization. Impervious surface indices include the normalized difference built-up index (NDBI), the normalized difference impervious surface index (NDISI), the enhanced normalized differences impervious surface index (ENDISI), and the normalized impervious surface index (NISI). LST and all spectral indices were derived from annual and seasonal composites of 30-m resolution Landsat 5-9 Level-2 Surface Reflectance imagery. The seven bioclimatic variables (e.g., air temperature, precipitation) were sourced from 1-km resolution gridded estimates of daily climatic data from NASA Daymet V4. We created temporally-aggregated Daymet raster images by calculating mean pixel-values for each season and year, as well as seasonally and annually summed precipitation. We summarized the values of each environmental variable by generating variously-sized (100-m, 500-m, 1000-m) buffers around each bird point count location and extracting weighted mean values of each environmental variable, with each pixel's values weighted by the proportion of its area falling within the buffer. All imagery retrieval and data processing were completed with Google Earth Engine (Gorelick et al. 2017) and program R. A complete description of data processing methods, including the aggregation of imagery by year and season and the calculation of s
Summary statistics and annual trends in chloride concentration in urban Minnesota lakes and streams
These data tables describe statistical summaries of chloride concentration, temporal trends in annual chloride, and projected risk of future chloride pollution in lakes and streams of the 17 most urban counties in Minnesota. Data were summarized separately for lakes and for streams, and include statistics (mean, median, standard deviation, upper and lower confidence intervals, and maximum) over the entire data record, over the warm season (May - October), and over the most recent 5 years (i.e., since 2018). Trends were computed on annual means and medians. For lakes, data were aggregated by lake basin (MN DNR Lake ID, or DOW) as well as by depth of sample (surface and deep). For streams, data were aggregated by individual site level as well as by stream reach (per MN Pollution Control Agency assessment units). Risk of chloride pollution was also determined for sites with longer records (10+ years) based on current concentration, number of exceedances of chronic standards, and projected chloride concentration based on current trends. Raw data were extracted from two sources: (1) the National Water Quality Portal (USGS & EPA) and (2) the Metropolitan Council Environmental Information Management System. The data were originally collected by a large number of entities, including watershed management authorities in the state of Minnesota, tribal groups, the Minnesota Pollution Control Agency, municipalities, university researchers, private consultants, and the Metropolitan Council. Some data records begin as early as the 1950's or 1960's, with many sites still including active data collection. A total of approximately 45,000 observations of chloride were included for lakes and wetlands, and approximately 70,000 observations for streams. Nearly 1600 stream sites and 700 lake/wetland sites were represented in the raw data, with 356 stream sites and 600 lakes represented in the summaries (after filtering out sites with less than 1 year of data collection). Primary data retr
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
Genome- and transcriptome-wide association summary statistics for outcome from traumatic brain injury
<p>The dataset contains summary statistics for the genome- and transcriptome-wide association studies (GWAS, TWAS) of genetic effects on outcome in traumatic brain injury (TBI). The study participants attended hospital within 24 hours of TBI, and underwent head computed tomography imaging.</p> <p><strong>Study participants</strong></p> <p>European ancestry data set contains 4710 individuals; multi-ethnic cohort 5268 individuals, including Europeans (n = 4710), Africans (n = 245) and Admixed Americans (n = 313).</p> <p>The largest European population contribution was from CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research, https://www.center-tbi.eu), where each participating center (60 centers from 20 countries in Europe) recruited patients between December 2013 and December 2017. The patients recruited in CENTER-TBI were supplemented by subjects from cohorts recruited at two European centres (Cambridge, UK, and Turku, Finland).</p> <p>The majority of patients in the US cohort were recruited between 2014 and 2018 to TRACK-TBI (Transforming Research and Clinical Knowledge in TBI, https://tracktbi.ucsf.edu) by the 18 US participant sites. The subjects recruited to the US cohort from TRACK-TBI were supplemented by patients recruited to an institutional research initiative at Mass General Brigham (MGB).</p> <p><strong>Outcome definition</strong></p> <p>Outcomes were measured using the extended Glasgow Outcome Scale (GOSE), ranging from 1 (dead) to 8 (upper good recovery), measured 6 months post-TBI. TBI severity was specified using the Glasgow Coma Score (GCS), with TBI classified as mild (GCS 13-15), moderate (GCS 9-12), or severe (GCS 3-8).</p> <p>To account for the effect of injury severity on outcome, sliding dichotomization was used to categorize outcome as favourable or unfavourable. A GOSE ≤ 4 was used to define an unfavourable outcome for patients with either moderate (GCS 9-12) or severe (GCS 3-8) TBI, while the unfavourable group was extended to patients with GOSE ≤ 7 if they had mild (GCS 13-15) TBI.</p> <p><strong>Genotype data and imputation</strong></p> <p>Genotyping was completed at FIMM Technology Center for CENTER-TBI, Cambridge, Turku patients and the Broad Institute for TRACK-TBI, using the Illumina Global Screening Array (GSA-24v2-0 + Multi-Disease). The MGB cohort were genotyped using Illumina’s Multi-Ethnic Global array (MEGA) and the pre-releases forms, including MEGA and MEGA-Ex arrays at Illumina at the MGB Translational Genomics Core.</p> <p>A unified quality control procedure was applied for each study cohort and the array-based genotypes were imputed using the Haplotype Reference Consortium panel. Autosomal chromosomes were considered, post-imputation data was filtered by imputation quality (INFO > 0.4 for CENTER-TBI, Cambridge and Turku; R2 > 0.4 for TRACK-TBI and MGB) and MAF > 1%.</p> <p><strong>Genome-wide association analysis and meta-analysis</strong></p> <p>Genome-wide single-marker scans were performed using a penalized likelihood-based Firth logistic regression, and implemented in PLINK v2.0. Using favourable outcome as reference, models were fitted on the basis of imputed allelic dosages. Age, sex, major extracranial injury, pupillary reactivity, and the first 10 principal components were included as covariates. Study cohort (CENTER-TBI, Cambridge, Turku) was an additional covariate in the CENTER-TBI GWAS.</p> <p>Fixed-effects meta-analysis of the three European ancestry GWAS was performed using METAL. For trans-ethnic meta-analysis, summary statistics of five GWASs in patients of European, African and Admixed Americans were aggregated via MR-MEGA.</p> <p><strong>Transcriptome-wide association study</strong></p> <p>Genetically regulated gene expression (GREx) was imputed using a regression model fitted on a separate gene expression database. Elastic net models provided by PrediXcan for all available GTEx brain tissues and whole blood were used. For TWAS, the same sliding dichotomy model for outcome with the same set of covariates as in the GWAS, but PCA components were replaced with the top five principal components of the respective gene expression data. </p> <p><strong>Column headers - GWAS</strong></p> <p>rsID: variant rsID<br> Chrom: chromosome<br> Pos: position (build GRCh38)<br> A1: effect allele<br> A2: reference allele<br> EAF: allele frequency of effect allele<br> Effect: effect size of effect allele<br> StdErr: standard error of effect size<br> P: p value of association (with genomic correction)<br> N: sample size</p> <p>Note. 'Effect' and 'StdErr' are only available for the European ancestry meta-analysis.</p> <p><br> <strong>Column headers - TWAS</strong></p> <p>tissue: GTEx tissue type<br> id: ensembl gene id<br> coef: model coefficient<br> se: model standard error for coefficient<br> p: model-based p value<br> symbol: gene symbol<br> name: gene name written out<br> chr: chromosome<br> start: gene start position (build GRCh38)</p>
Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)
<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., & Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>. </p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p> </p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>
Meningitis hGWAS results (summary statistics)
<p>Summary statistics for association between genetic variation and meningitis phenotypes. Contains human genome association and interaction effects (pGWAS.tar.bz2).</p> <p>Unpack with `tar xf `. Contents are described in the README.</p>
Voxel-level summary statistics of hippocampus shape, white matter microstructure, and cortical surface curvature in UK Biobank (n=33,324)
<p>This deposit hosts GWAS summary statistics of hippocampus shape (n=33,324), white matter microstructure (n=33,324), and cortical surface curvature (n=15,752) using UKB unrelated white subjects. The data was generated by using the highly efficient imaging genetics (<a href="https://github.com/Zhiwen-Owen-Jiang/heig">HEIG v1.1.0</a>) framework where only the triplets - summary statistics of low-dimensional representations (LDRs), the functional bases, and the variance-covariance matrix LDRs - are shared, which is sufficient to recover all voxel-variant pairs as well as to conduct voxel-level heritability and (cross-trait) genetic correlation analysis. Check the <a href="https://github.com/Zhiwen-Owen-Jiang/heig/wiki">tutorial</a> and the <a href="../records/13770930">example data</a> used in the tutorial. </p> <p>The shared data includes:</p> <p>1. Triplets for hippocampus shape measured by the radial distance from the medial model for each vertex. The original images contain 30,000 vertices while the shared data contains 49 LDRs. Left and right hemispheres were analyzed separately, each with 15,000 vertices.</p> <p>2. Triplets for 21 white matter tracts measured by fractional anisotropy. The original images contain 32,217 voxels and each tract contains 88 ~ 3503 voxels while the shared data contains 1,034 LDRs. Tracts were analyzed separately.</p> <p>3. Triplets for cortical surface curvature. The original images contain 59,412 vertices while the shared data contains 1,750 LDRs. The entire brain was analyzed as a whole.</p> <p>4. LD matrix and its inverse for 22 chromosomes including 460k genotyped SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 8.4k white unrelated subjects in UKB. Two regularization levels are provided: {85%, 80%} for heritability and genetic correlations within images and {75%, 70%} for cross-trait genetic correlations.</p> <p>5. LD matrix and its inverse for 22 chromosomes including 1.2 million imputed HapMap3 SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 42k white unrelated subjects in UKB. Two regularization levels are provided: {98%, 95%} for heritability and genetic correlations within images and {90%, 85%} for cross-trait genetic correlations.</p>
GWAS summary statistics for waist-to-hip ratio and body principal components
<p>This dataset contains genome-wide association summary statistics for waist-to-hip ratio (WHR), as well as those for body principal components (PCs). A subset of 387,139 unrelated, white British individuals were analyzed for WHR. PCs were combined from the summary statistics for WHR and 13 other anthropometric traits (body mass index, standing height, weight, hip circumference, waist circumference, arm lean mass (left), arm fat mass (left), leg lean mass (left), leg fat mass (left), trunk lean mass, trunk fat mass, body fat percentage, basal metabolic rate) provided by the Neale lab (http://www.nealelab.is/uk-biobank). All traits were inverse-rank normal transformed (by the Neale lab or ourselves for WHR).</p> <p>All effect sizes, including those for PCs, are standardized, i.e. they represent the effects on a trait with variance 1.</p> <p>The zip files contain the data to run the sample pipeline and the shiny app, both available from <a href="https://github.com/JonSulc/PCA_Cross-sex_MR">https://github.com/JonSulc/PCA_Cross-sex_MR</a>.</p>
QTL summary statistics from the DIRECT consortium
<p>These are the complete summary statistics for DIRECT genotype-phenotypes associations (QTLs). The project performed genotype-phenotype associations for gene expression (RNAseq), targeted proteins (Olink), targeted metabolites (Biocrates) and untargeted metabolites (Metabolon) derived from 3,029 blood and plasma samples from the DIRECT cohort. This submission includes supplementary files and nominal pvalues (as uncorrected pvalues) for all associations included in the manuscript. Trans associations included are typically limited to pvalues <1e-04. Network tables are also included, with information to load and use Cytoscape to visualize them. This is version 2, some files were missing on version 1.</p>
Genome-wide association summary statistics of chronic musculoskeletal pain at four anatomic sites and their genetically independent components
<p>The dataset contains results of a genome-wide association study of distinct chronic musculoskeletal pain conditions: back pain, knee pain, neck pain, and hip pain. Additionally, there are genome-wide association summary statistics for four genetically independent components of pain conditions, listed above. For more details, please, read the paper XXX.</p> <p>All files contain association summary statistics for genome-wide association meta-analysis of the 265,000 white British individuals from the UK Biobank and additional 191,580 individuals of European Ancestry from the UK biobank (total N = 456,580). Cases and controls were defined based on questionnaire responses. First, participants responded to “Pain type(s) experienced in the last months” followed by questions inquiring if the specific pain had been present for more than 3 months. Those who reported back, neck or shoulder, hip, or knee pain lasting more than 3 months were considered chronic back, neck/shoulder, hip, and knee pain cases, respectively. Participants reporting no such pain lasting longer than 3 months were considered controls (regardless of whether they had another regional chronic pain, such as abdominal pain, or not). Individuals who preferred not to answer were excluded from the study. Besides this, we excluded individuals who reported more than 3 months of pain all over the body.</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilization of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite the corresponding paper and this repository:</strong></p> <ol> <li>Tsepilov et al 2020</li> </ol> <p><strong>Funding:</strong></p> <p>The work of YSA and SZS was supported by the Russian Ministry of Education and Science under the 5-100 Excellence Programme and by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project 0324-2019-0040). The work of YAT, ASSh, and EEE was supported by the Russian Foundation for Basic Research (project 19-015-00151). The contribution of LСK was funded by PolyOmica. Dr. Suri was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development (RR&D) Service. Dr. Suri is a Staff Physician at the VA Puget Sound Health Care System. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p><strong>List of files:</strong></p> <ol> <li>Back_output_done.csv: GWAS summary statistics for the chronic back pain</li> <li>gpc1_output_done.csv: GWAS summary statistics for the GIP1</li> <li>gpc2_output_done.csv: GWAS summary statistics for the GIP2</li> <li>gpc3_output_done.csv: GWAS summary statistics for the GIP3</li> <li>gpc4_output_done.csv: GWAS summary statistics for the GIP4</li> <li>Hip_output_done.csv: GWAS summary statistics for the chronic hip pain</li> <li>Knee_output_done.csv: GWAS summary statistics for the chronic knee pain</li> <li>Neck_output_done.csv: GWAS summary statistics for the chronic neck pain</li> </ol> <p><strong>Column headers:</strong></p> <ol> <li>gwas_id: uninformative field</li> <li>rs_id: dbSNP rsID (GRCh37 build) </li> <li>snp_num: uninformative field</li> <li>chr: chromosome (GRCh37 build) </li> <li>bp: position (GRCh37 build) </li> <li>ea: effect allele (coded as "1")</li> <li>ra: reference allele (coded as "0")</li> <li>eaf: effect allele frequency</li> <li>af_ref: uninformative field</li> <li>beta: effect size of effect allele</li> <li>se: standard error of effect size</li> <li>p: P-value of association (without GC correction)</li> <li>n:Total sample size</li> <li>z: Z-statistic of association</li> <li>info: uninformative field</li> <li>af_outlier: uninformative field</li> <li>pz_outlier: uninformative field</li> </ol>
Data Set of Extracted Summary Statistics from Equipment Sensor Data
<p>This data set was generated in accordance with the semiconductor industry and contains values of summary statistics from sensor recordings of the high-precision and high-tech production equipment. Basically, the semiconductor production consists of hundreds of process steps performing physical and chemical operations on so-called wafers, i.e. slices based on semiconductor material. In the production chain, each process equipment is equipped with several sensors recording physical parameters like gas flow, temperature, voltage, etc., resulting in so-called sensor data. Out of the sensor data, values of summary statistics are extracted. These are values like mean, standard deviation and gradients. To keep the entire production as stable as possible, these values are used to monitor the whole production in order to intervene in case of deviations.</p> <p>After the production, each device on the wafer is tested in the most careful way resulting in so-called wafer test data. In some cases, suspicious patterns occur in the wafer test data potentially leading to failure. In this case the root cause must be found in the production chain. For this purpose, the given data is provided. The aim is to find correlations between the wafer test data and the values of summary statistics in order to identify the root cause.</p> <p>The given data is divided into four data sets: "XTrain.csv", "YTrain.csv", "XTest.csv" and "YTest.csv". "XTrain.csv" and "XTest.csv" represent the values of summary statistics originating in the production chain separated for the purpose of training and validating a statistical model. Included are 114 observations of 77 parameters (values of summary statistics). The "YTrain.csv" and "YTest.csv" contain the corresponding wafer test data (144 observations of one parameter).</p>
Summary statistic of a Trans ancestry multi-trait GWAS
<p>Trans ancestry multi-trait GWAS by adapting the omnibus test to the trans ancestry setting. </p><p>Genome wide summary statistics for 19 blood count traits were retrieved from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel)</a></p><p>They curated using the JASS (Joint Analysis of Summary Statistics) pipeline https://gitlab.pasteur.fr/statistical-genetics/jass_suite_pipeline</p><p>See Troubat et al preprint for all details on the obtention of this dataset https://doi.org/10.1101/2023.06.23.546248</p>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Regional summary statistics for 1107 protein targets based on the Olink technology
<p>This data set contains regional summary statistics (±500kb around the protein coding gene) for a total of 1107 protein - gene combinations as measured by the Olink Proximity Extension Assay in the Fenland study (https://www.mrc-epid.cam.ac.uk/research/studies/fenland/) among 485 individuals. A detailed description of the genetic analysis can be found here https://www.nature.com/articles/s41467-021-27164-0. </p>
Summary Statistics from "Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2"
<p>GWAMA summary statistics of PCSK9 levels using fixed-effect model. Genome-wide data is given for Europeans with statin adjustment and Europeans without statin treatment only (subset of the population). In addition, locus-wide data of the PCSK9 gene locus for African-Americans without statin treatment is listed.</p> <p>When using this data, please cite: Pott J, Gadin J, Theusch E, et al.. Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2. Hum Mol Genet. 2021 Sep 30:ddab279. doi: 10.1093/hmg/ddab279. PMID: 34590679</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"
<p>These are the burden test results (summary statistics) for the publication:</p> <p>"Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer’s Disease",</p> <p>Nature Genetics, 2022.</p> <p> </p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on 6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL>=75, LOF+REVEL>=50, LOF+REVEL>=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p> </p>
nextGEMS cycle3 datasets: statistical summaries for streamed data from climate simulations
<p>This Zenodo holds the datasets used in the paper "Statistical summaries for streamed data from climate simulations" by Katherine Grayson. All the data comes from the nextGEMS cycle 3 and has been regridded for plotting purposes with resolution given in the title of each data set. The wind speed data set has been made by taking the square root of the squared and summed v10 and u10 components respectively. All data has been retrieved and regridded through the AQUA reader on the Levante supercomputer, developed as part of the Destination Earth initative. The source code to create all the figures using this data can be found in https://github.com/kat-grayson/one_pass_algorithms_paper/tree/main </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.