Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Genomics statistics”
Genome-wide association summary statistics for human blood plasma glycome
<p>The dataset contains results of genome-wide association study of human blood plasma glycome. The 113 files contain association summary statistics for 113 glycome traits, of which 36 were directly measured by UPLC technology and 77 were derived glycome traits. Description of each glycome trait can be found in the <strong>Additional notes</strong> section. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Sharapov, S. Z., Tsepilov, Y. A., Klaric, L., Mangino, M., Thareja, G., Shadrina, A. S., … Aulchenko, Y. (2019). Defining the genetic control of human blood plasma N-glycome using genome-wide association study. <em>Human Molecular Genetics</em>. http://doi.org/10.1093/hmg/ddz054</li> <li>Sodbo Sharapov, Yakov Tsepilov, Lucija Klaric, Massimo Mangino, Gaurav Thareja, Mirna Simurina, Concetta Dagostino, Julia Dmitrieva, Marija Vilaj, FranoVuckovic, Tamara Pavic, Jerko Stambuk, Irena Trbojevic-Akmacic, Jasminka Kristic, Jelena Simunovic, Ana Momcilovic, Harry Campbell, Malcolm Dunlop, Susan Farrington, Maria Pucic-Bakovic, Christian Gieger, Massimo Allegri, Edouard Louis, Michel Georges, Karsten Suhre, Tim Spector, Frances MK Williams, Gordan Lauc, Yurii Aulchenko. (2018). Genome-wide association summary statistics for human blood plasma glycome (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1298406</li> </ol> <p><strong>Funding</strong></p> <p>This work was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736) and by the European Structural and Investments funding for the "Croatian National Centre of Research Excellence in Personalized Healthcare" (contract #KK.01.1.1.01.0010).</p> <p>The work of SSh was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.</p> <p>The work of YT was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).</p> <p>Karsten Suhre and Gaurav Thareja are supported by ‘Biomedical Research Program’ funds at Weill Cornell Medicine - Qatar, a program funded by the Qatar Foundation. We thank all staff at Weill Cornell Medicine - Qatar and Hamad Medical Corporation, and especially all study participants who made the QMDiab study possible.</p> <p>The SOCCS study was supported by grants from Cancer Research UK (C348/A3758, C348/A8896, C348/ A18927); Scottish Government Chief Scientist Office (K/OPR/2/2/D333, CZB/4/94); Medical Research Council (G0000657-53203, MR/K018647/1); Centre Grant from CORE as part of the Digestive Cancer Campaign (<a href="http://www.corecharity.org.uk">http://www.corecharity.org.uk</a>).</p> <p>TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility and Biomedical Research Centre based at Guy’s and St Thomas’ NHS Foundation Trust in partnership with King’s College London.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNP: SNP rsID</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>OTHER_ALLELE: reference allele (coded as "0")</li> <li>EFFECT_ALLELE: effective allele (coded as "1")</li> <li>EAF: effective allele frequency </li> <li>N: sample size</li> <li>BETA: effect size of effective allele</li> <li>SE: standard error of effect size</li> <li>PVAL: P-value of association (without GC correction)</li> <li>IMPUTATION: imputation quality</li> </ol>
Genome-wide association summary statistics for human healthspan
<p>The dataset contains genome-wide association summary statistics computed for heathspan. The UKB sub-population of 300,447 genetically Caucasian, British individuals were analyzed. For more details see [1].</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Zenin, A., Tsepilov, Y., Sharapov, S., Getmantsev, E., Menshikov, L. I., Fedichev, P. O., & Aulchenko, Y. (2019). Identification of 12 genetic loci associated with human healthspan. <em>Communications Biology</em>, <em>2</em>(1), 41. http://doi.org/10.1038/s42003-019-0290-0</li> <li>Aleksandr Zenin, Yakov Tsepilov, Sodbo Sharapov, Evgeny Getmantsev, Leonid Menshikov, Peter Fedichev, & Yurii Aulchenko. (2018). Genome-wide association summary statistics for human healthspan (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1302861</li> </ol> <p><strong>Funding</strong></p> <p>The work was supported by Russian Ministry of Science and Education under 5-100 Excellence Programme. <br> The work was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017). <br> This research has been conducted using the UK Biobank Resource. <br> The study has been funded by Gero LLC.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNPID - SNP rsID</li> <li>chr - chromosome</li> <li>pos - position (GRCh37 build / hg19)</li> <li>EA - effective allele (coded as "1")</li> <li>RA - reference allele (coded as "0")</li> <li>EAF - effective allele frequency</li> <li>beta - effect size of effective allele</li> <li>se - standard error of effect size</li> <li>Z - Z-value of association</li> <li>-log10(p-value) - minus log10(P-value) of association</li> </ol>
Genome-wide association summary statistics for back pain
<p>The dataset contains results of a genome-wide association study of back pain. Two files contain association summary statistics for discovery GWAS based on the analysis of 350,000 white British individuals from the UK Biobank and meta-analysis GWAS based on the meta-analysis of the same 350,000 individuals and additional 103,862 individuals of European Ancestry from the UK biobank (total N = 453,862). The phenotype of back pain was defined by the answer provided by the UK biobank participants to the following question: "Pain type(s) experienced in last month". Those who reported “Back pain”, were considered as cases, all the rest were considered as controls. Individuals who did not reply or replied: "Prefer not to answer" or "Pain all over the body" were excluded. This dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org/">http://gwasarchive.org</a>. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Insight into the genetic architecture of back pain and its risk factors from a study of 509,000 individuals. Freidin, Maxim; Tsepilov, Yakov; Palmer, Melody; Karssen, Lennart; Suri, Pradeep; Aulchenko, Yurii; Williams, Frances MK,# CHARGE Musculoskeletal Working Group. PAIN: February 06, 2019 - Volume Articles in Press - Issue - p<br> doi: 10.1097/j.pain.0000000000001514</li> <li>Maxim B Freidin, Yakov A Tsepilov, Melody Palmer, Lennart Karssen, CHARGE Musculoskeletal Working Group, Pradeep Suri, … Frances MK Williams. (2018). Genome-wide association summary statistics for back pain (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1319332</li> </ol> <p><strong>Funding:</strong></p> <p>This study was supported by the European Community’s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736). <br> The research has been conducted using the UK Biobank Resource (project # 18219).</p> <p>The development of software implementing SMR/HEIDI test and database for GWAS results was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Program”.</p> <p>Dr. Suri’s time for this work was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development Service. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p>Dr. Tsepilov’s time for this work was supported in part by the Russian Ministry of Science and Education under the 5-100 Excellence Program.</p> <p><strong>Column headers - discovery (350K)</strong></p> <ol> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>ID: SNP rsID</li> <li>REF: reference allele (coded as "0")</li> <li>ALT: effect allele (coded as "1")</li> <li>CASE_ALLELE_CT: allele observation count in cases</li> <li>CTRL_ALLELE_CT: allele observation count in controls</li> <li>ALT_FREQ: effect allele frequency </li> <li>MACH_R2: imputation quality</li> <li>TEST: model of association test (additive)</li> <li>OBS_CT: sample size</li> <li>BETA: effect size of effect allele</li> <li>SE: standard error of effect size</li> <li>T_STAT: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> <li>MAF: minor allele frequency</li> </ol> <p><strong>Column headers - meta-analysis (450K)</strong></p> <ol> <li>MarkerName: SNP rsID</li> <li>Allele1: effect allele (coded as "1")</li> <li>Allele2: reference allele (coded as "0")</li> <li>Freq1: effect allele frequency</li> <li>FreqSE: standard error of effect allele frequency</li> <li>Effect: effect size of effect allele</li> <li>StdErr: standard error of effect size</li> <li>P-value: P-value of association (without GC correction)</li> <li>Direction: sign of effect in discovery and replication samples</li> <li>n_total: Total sample size</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build) </li> <li>MACH_R2_discovery: imputation quality in discovery sample</li> </ol>
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
Genome- and transcriptome-wide association summary statistics for outcome from traumatic brain injury
<p>The dataset contains summary statistics for the genome- and transcriptome-wide association studies (GWAS, TWAS) of genetic effects on outcome in traumatic brain injury (TBI). The study participants attended hospital within 24 hours of TBI, and underwent head computed tomography imaging.</p> <p><strong>Study participants</strong></p> <p>European ancestry data set contains 4710 individuals; multi-ethnic cohort 5268 individuals, including Europeans (n = 4710), Africans (n = 245) and Admixed Americans (n = 313).</p> <p>The largest European population contribution was from CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research, https://www.center-tbi.eu), where each participating center (60 centers from 20 countries in Europe) recruited patients between December 2013 and December 2017. The patients recruited in CENTER-TBI were supplemented by subjects from cohorts recruited at two European centres (Cambridge, UK, and Turku, Finland).</p> <p>The majority of patients in the US cohort were recruited between 2014 and 2018 to TRACK-TBI (Transforming Research and Clinical Knowledge in TBI, https://tracktbi.ucsf.edu) by the 18 US participant sites. The subjects recruited to the US cohort from TRACK-TBI were supplemented by patients recruited to an institutional research initiative at Mass General Brigham (MGB).</p> <p><strong>Outcome definition</strong></p> <p>Outcomes were measured using the extended Glasgow Outcome Scale (GOSE), ranging from 1 (dead) to 8 (upper good recovery), measured 6 months post-TBI. TBI severity was specified using the Glasgow Coma Score (GCS), with TBI classified as mild (GCS 13-15), moderate (GCS 9-12), or severe (GCS 3-8).</p> <p>To account for the effect of injury severity on outcome, sliding dichotomization was used to categorize outcome as favourable or unfavourable. A GOSE ≤ 4 was used to define an unfavourable outcome for patients with either moderate (GCS 9-12) or severe (GCS 3-8) TBI, while the unfavourable group was extended to patients with GOSE ≤ 7 if they had mild (GCS 13-15) TBI.</p> <p><strong>Genotype data and imputation</strong></p> <p>Genotyping was completed at FIMM Technology Center for CENTER-TBI, Cambridge, Turku patients and the Broad Institute for TRACK-TBI, using the Illumina Global Screening Array (GSA-24v2-0 + Multi-Disease). The MGB cohort were genotyped using Illumina’s Multi-Ethnic Global array (MEGA) and the pre-releases forms, including MEGA and MEGA-Ex arrays at Illumina at the MGB Translational Genomics Core.</p> <p>A unified quality control procedure was applied for each study cohort and the array-based genotypes were imputed using the Haplotype Reference Consortium panel. Autosomal chromosomes were considered, post-imputation data was filtered by imputation quality (INFO > 0.4 for CENTER-TBI, Cambridge and Turku; R2 > 0.4 for TRACK-TBI and MGB) and MAF > 1%.</p> <p><strong>Genome-wide association analysis and meta-analysis</strong></p> <p>Genome-wide single-marker scans were performed using a penalized likelihood-based Firth logistic regression, and implemented in PLINK v2.0. Using favourable outcome as reference, models were fitted on the basis of imputed allelic dosages. Age, sex, major extracranial injury, pupillary reactivity, and the first 10 principal components were included as covariates. Study cohort (CENTER-TBI, Cambridge, Turku) was an additional covariate in the CENTER-TBI GWAS.</p> <p>Fixed-effects meta-analysis of the three European ancestry GWAS was performed using METAL. For trans-ethnic meta-analysis, summary statistics of five GWASs in patients of European, African and Admixed Americans were aggregated via MR-MEGA.</p> <p><strong>Transcriptome-wide association study</strong></p> <p>Genetically regulated gene expression (GREx) was imputed using a regression model fitted on a separate gene expression database. Elastic net models provided by PrediXcan for all available GTEx brain tissues and whole blood were used. For TWAS, the same sliding dichotomy model for outcome with the same set of covariates as in the GWAS, but PCA components were replaced with the top five principal components of the respective gene expression data. </p> <p><strong>Column headers - GWAS</strong></p> <p>rsID: variant rsID<br> Chrom: chromosome<br> Pos: position (build GRCh38)<br> A1: effect allele<br> A2: reference allele<br> EAF: allele frequency of effect allele<br> Effect: effect size of effect allele<br> StdErr: standard error of effect size<br> P: p value of association (with genomic correction)<br> N: sample size</p> <p>Note. 'Effect' and 'StdErr' are only available for the European ancestry meta-analysis.</p> <p><br> <strong>Column headers - TWAS</strong></p> <p>tissue: GTEx tissue type<br> id: ensembl gene id<br> coef: model coefficient<br> se: model standard error for coefficient<br> p: model-based p value<br> symbol: gene symbol<br> name: gene name written out<br> chr: chromosome<br> start: gene start position (build GRCh38)</p>
Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)
<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., & Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>. </p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p> </p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>
Genome-wide association summary statistics of chronic musculoskeletal pain at four anatomic sites and their genetically independent components
<p>The dataset contains results of a genome-wide association study of distinct chronic musculoskeletal pain conditions: back pain, knee pain, neck pain, and hip pain. Additionally, there are genome-wide association summary statistics for four genetically independent components of pain conditions, listed above. For more details, please, read the paper XXX.</p> <p>All files contain association summary statistics for genome-wide association meta-analysis of the 265,000 white British individuals from the UK Biobank and additional 191,580 individuals of European Ancestry from the UK biobank (total N = 456,580). Cases and controls were defined based on questionnaire responses. First, participants responded to “Pain type(s) experienced in the last months” followed by questions inquiring if the specific pain had been present for more than 3 months. Those who reported back, neck or shoulder, hip, or knee pain lasting more than 3 months were considered chronic back, neck/shoulder, hip, and knee pain cases, respectively. Participants reporting no such pain lasting longer than 3 months were considered controls (regardless of whether they had another regional chronic pain, such as abdominal pain, or not). Individuals who preferred not to answer were excluded from the study. Besides this, we excluded individuals who reported more than 3 months of pain all over the body.</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilization of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite the corresponding paper and this repository:</strong></p> <ol> <li>Tsepilov et al 2020</li> </ol> <p><strong>Funding:</strong></p> <p>The work of YSA and SZS was supported by the Russian Ministry of Education and Science under the 5-100 Excellence Programme and by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project 0324-2019-0040). The work of YAT, ASSh, and EEE was supported by the Russian Foundation for Basic Research (project 19-015-00151). The contribution of LСK was funded by PolyOmica. Dr. Suri was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development (RR&D) Service. Dr. Suri is a Staff Physician at the VA Puget Sound Health Care System. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p><strong>List of files:</strong></p> <ol> <li>Back_output_done.csv: GWAS summary statistics for the chronic back pain</li> <li>gpc1_output_done.csv: GWAS summary statistics for the GIP1</li> <li>gpc2_output_done.csv: GWAS summary statistics for the GIP2</li> <li>gpc3_output_done.csv: GWAS summary statistics for the GIP3</li> <li>gpc4_output_done.csv: GWAS summary statistics for the GIP4</li> <li>Hip_output_done.csv: GWAS summary statistics for the chronic hip pain</li> <li>Knee_output_done.csv: GWAS summary statistics for the chronic knee pain</li> <li>Neck_output_done.csv: GWAS summary statistics for the chronic neck pain</li> </ol> <p><strong>Column headers:</strong></p> <ol> <li>gwas_id: uninformative field</li> <li>rs_id: dbSNP rsID (GRCh37 build) </li> <li>snp_num: uninformative field</li> <li>chr: chromosome (GRCh37 build) </li> <li>bp: position (GRCh37 build) </li> <li>ea: effect allele (coded as "1")</li> <li>ra: reference allele (coded as "0")</li> <li>eaf: effect allele frequency</li> <li>af_ref: uninformative field</li> <li>beta: effect size of effect allele</li> <li>se: standard error of effect size</li> <li>p: P-value of association (without GC correction)</li> <li>n:Total sample size</li> <li>z: Z-statistic of association</li> <li>info: uninformative field</li> <li>af_outlier: uninformative field</li> <li>pz_outlier: uninformative field</li> </ol>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
KMA Mapping and alignment statistics : livestock fecal metagenomes against ResFinder and genomes
<p>Three zip archives are included used in the analysis of the European livestock resistome.</p> <p>Two of them contain 'mapstat' files produced by the KMA software using the 'extended features' flag.<br> Each mapstat file thus summarize the mapping and alignment statistics when using KMA on a metagenome against a database.</p> <p>The last archive contains the 'refdata' file used to annotate the genomic mapstat hits. It encodes the taxonomic affilication of sequences hit by one or more samples.<br> </p>
Genome-wide association summary statistics for varicose veins of lower extremities
<p>The dataset contains summary statistics for the discovery and the replication stages of the large-scale genome-wide associations study for varicose veins of lower extremities. The discovery stage was based on genetic association data provided by the Neale Lab (<a href="https://vk.com/away.php?to=http%3A%2F%2Fwww.nealelab.is%2F&cc_key=">http://www.nealelab.is/</a>) for 337,199 UK biobank individuals. Phenotype “varicose veins of lower extremities” was defined based on International Classification of Disease (ICD-10) billing code “I83” present in the electronic patient record. Data were adjusted for two potential confounders – body mass index and deep venous thrombosis. A replication cohort (N=71,256) was generated by means of reverse meta-analysis of two overlapping datasets: genetic association data for 408,455 UK Biobank participants provided by the Gene ATLAS database (<a href="https://vk.com/away.php?to=http%3A%2F%2Fgeneatlas.roslin.ed.ac.uk%2F&cc_key=">http://geneatlas.roslin.ed.ac.uk/</a>), and the above mentioned data provided by the Neale Lab.</p> <p>Please, note, that in Shadrina et al (PLOS Genetics 2019) we only used "discovery" dataset, while in biorxiv preprint (https://doi.org/10.1101/368365) both discovery and replication datasets were used. </p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. </p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li> <p>Shadrina, A. S., Sharapov, S. Z., Shashkova, T. I. & Tsepilov, Y. A. Varicose veins of lower extremities: Insights from the first large-scale genetic study. <em>PLOS Genet.</em> <strong>15,</strong> e1008110 (2019).</p> </li> <li>Alexandra S. Shadrina, Sodbo Zh. Sharapov, Tatiana I. Shashkova, & Yakov A. Tsepilov. (2018). Genome-wide association summary statistics for varicose veins of lower extremities (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1323484</li> </ol> <p><strong>Funding:</strong></p> <p>The work of ASS was supported by the Russian Science Foundation [Project No 17-75-20223]. <br> The work of YAT was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme. <br> The work of SZS was supported by the Institute of Cytology and Genetics [Project No 0324-2018-0017].</p> <p><strong>Column headers - discovery</strong></p> <ol> <li>SNP: SNP rsID</li> <li>b: effect size of effect allele</li> <li>se: standard error of effect size</li> <li>chi2: T^2 value of effect allele</li> <li>Pval: P-value of association (without GC correction)</li> <li>N: sample size</li> <li>Chr: chromosome</li> <li>Pos: position (GRCh37 build)</li> <li>A1: effect allele (coded as "1")</li> <li>A2: reference allele (coded as "0")</li> </ol> <p><strong>Column headers - replication</strong></p> <ol> <li>SNP: SNP rsID</li> <li>A1: effect allele (coded as "1")</li> <li>A2: reference allele (coded as "0")</li> <li>N: Total sample size</li> <li>Z: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> </ol>
Genome-wide association statistics of Hearing Problems
<p>Genome-wide Association Statistics of Hearing Problems</p> <p>Citation: De Angelis F, Zeleznik OA, Wendt FR, Pathak GA, Tylee DS, De Lillo A, Koller D, Cabrera-Mendoza B, Clifford RE, Maihofer AX, Nievergelt CM, Curhan GC, Curhan SG, Polimanti R. Sex differences in the polygenic architecture of hearing problems in adults. Genome Med. https://doi.org/10.1186/s13073-023-01186-3</p> <p>COLUMN HEADERS<br> chromosome: chromosome<br> base_pair_location: position<br> effect_allele: effect allele (corresponds to the effect size’s sign; may not be the alternate allele)<br> other_allele: non-effect allele<br> beta: effect measured as beta, sign corresponds to the effect of the effect allele<br> standard_error: standard error of the effect<br> effect_allele_frequency: effect allele frequency in UK Biobank participants of European descent<br> p_value: p value of the association statistic<br> variant_id: variant identifier<br> rs_id: rsID of the variant<br> n: sample size per variant</p> <p> </p>
Genome-wide association summary statistics for sex- and age-specific analysis of chronic back pain
<p>The dataset comprises summary-level statistics for age- and sex-specific genome-wide association study of chronic back pain (cBP) in individuals of European descent from UK Biobank (<a href="https://www.ukbiobank.ac.uk/">https://www.ukbiobank.ac.uk/</a>). The study was carried out under UK Biobank approved project #18219. </p> <p><strong>The dataset accompanies the paper (please cite if using the dataset):</strong></p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/33021770/">Freidin, Maxim B.; Tsepilov, Yakov A.; Stanaway, Ian B.; Meng, Weihua; Hayward, Caroline; Smith, Blair H.; Khoury, Samar; Parisien, Marc; Bortsov, Andrey; Diatchenko, Luda; Børte, Sigrid; Winsvold, Bendik S.; Brumpton, Ben M.; Zwart, John-Anker; HUNT All-In Pain; Aulchenko, Yurii S.; Suri, Pradeep; Williams, Frances M.K. Sex- and age-specific genetic analysis of chronic back pain. Pain. 2020. doi:10.1097/j.pain.0000000000002100.</a></p> <p>The phenotype of cBP was defined as back pain for 3+ months. Linear mixed-effects additive model was fitted adjusting for age, genotyping array type, and 10 genetic PCs provided by UK Biobank. The following filters were applied: minor allele frequency >0.001, genotyping and individual call rates >0.98%, imputation quality score (INFO) >0.7. GWAS were carried out in males and females separately in the whole sample (<strong>allages</strong>) as well as in groups of younger than 65 years (<strong>under65</strong>) and 65+ years old (<strong>65plus</strong>) as detailed in the paper. Accordingly, 6 files are deposited here, corresponding to each group. </p> <p><strong>Column headers:</strong></p> <p>SNP, SNP rsID </p> <p>CHR, chromosome</p> <p>BP, genomic position (GRCh37 build)</p> <p>EA, effect allele (coded as "1")</p> <p>OTHER, other allele (coded as "0")</p> <p>A1FREQ, frequency of effect allele</p> <p>INFO, imputation quality</p> <p>BETA, effect size (for effect allele)</p> <p>SE, standard error of effect size</p> <p>PVAL, p-value for association</p>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Full Summary Statistics - TBDAR Genome-to-genome Study
<p><strong>TBDAR_G2G_Full_Summary_Stats.tar.gz: </strong>Full summary statistics (See README for details)</p> <p><strong>Mtb_Human_IDs.txt: </strong>Mapping between M.tb and human sequencing IDs, to faciliate joint analyses. </p> <p><strong>Supple_Data1.csv</strong>: <span lang="EN">G2G associations that meet the significance threshold of 5 × 10⁻⁸ </span></p>
Genome-wide determinants of mortality and clinical progression in Parkinson's disease - Summary statistics
<p>Summary statistics from "Genome-wide determinants of mortality and clinical progression in Parkinson’s disease".</p>
Multivariate Genome-wide association summary statistics for shared aging factor
<p>This dataset contains genome-wide summary statistics (autosomal variants) computed from a multivariate genome-wide association study of five aging-related phenotypes using Genomic Structural Equation Modeling (https://github.com/GenomicSEM/GenomicSEM). The effective sample size is calculated to be 1,958,774. Column descriptions are included in the accompanying README file. </p> <p>The summary statistics are provided on an "AS-IS" basis, without any type of warranty, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose.</p> <p>If investigators use these data, any and all consequences are entirely their responsibility. The user agrees that to cite the appropriate publication in any communications or publications arising directly or indirectly from these data by downloading and and using these data,. The user also agrees to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles and also agree that they will never attempt to identify any participant.</p> <p> </p>
Summary Statistics from "Genome-wide meta-analysis of phytosterols reveals five novel loci and a detrimental effect on coronary atherosclerosis"
<p>Summary statistics of meta GWAS of phytosterols.</p>
Genome-wide association summary statistics for G4 and G6
<p>This repository contains GWAS summary statistics for G4 (g4_GWAS_Sumstats_Cleaned.txt) and G6 (g6_GWAS_Sumstats_Cleaned.txt), which are part of the paper titled "Dynamics of cognitive variability with age and its genetic underpinning in NIHR BioResource Genes and Cognition cohort participants".</p> <p>For methodological details check the paper at: <span><a href="https://gbr01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41591-024-02960-5&data=05%7C02%7Cshafiqur.rahman%40mrc-bsu.cam.ac.uk%7Ceba3de1db9174976852b08dc6f7001a0%7C513def5bdf174107b5523dba009e5990%7C0%7C0%7C638507774058203135%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C&sdata=c8Nc%2B%2BqUFnfIoFkLqM7S6FS6PbXk8kKvIE8QvvOb2Zw%3D&reserved=0">https://www.nature.com/articles/s41591-024-02960-5</a></span></p> <p>The columns are as follows:</p> <ul> <li>SNP: rs identifier for the SNP</li> <li>CHR: chromosome (GRCh37 build) </li> <li>BP: base pair (GRCh37 build) </li> <li>A1: effect allele </li> <li>A2: reference allele </li> <li>FREQ: effect allele frequency </li> <li>INFO: imputation information</li> <li>P: p-value </li> <li>BETA: effect size of effect allele</li> <li>SE: standard error </li> <li>N: sample size</li> </ul>
Genome-wide association study Summary statistics of Invasive melanoma vs controls, In situ Melanoma vs controls and In situ vs invasive melanoma (case-case)
<p>Genome-wide association study Summary statistics of Invasive melanoma vs controls, In situ Melanoma vs controls and In situ vs invasive melanoma (case-case). The first GWAS meta-analysis combines GWAS summary statistics of invasive melanoma from UK Biobank (as of August 2022), FinnGen release 9, QSkin Sun and Health Study and The Queensland Study of Melanoma: environmental and genetic associations (Q-MEGA) study.</p> <p>The second GWAS meta-analysis combines GWAS summary statistics of in situ melanoma from UK Biobank (as of August 2022), FinnGen release 9, QSkin Sun and Health Study and The Queensland Study of Melanoma: environmental and genetic associations (Q-MEGA) study.</p> <p>The third GWAS meta-analysis combines GWAS summary statistics of in situ vs invasive (case-case; in situ code 0, invasive code 1) melanoma from UK Biobank (as of August 2022), QSkin Sun and Health Study and The Queensland Study of Melanoma: environmental and genetic associations (Q-MEGA) study.</p> <p>Columns</p> <p>CHR Chromosome</p> <p>SNP rsid</p> <p>POS Base position HG Build 37</p> <p>A1 effect allele</p> <p>A2 Non-effect allele</p> <p>A1FREQ Allele frequency of effect allele</p> <p>BETA effect estimate of effect allele</p> <p>SE standard error of effect estimate</p> <p>PVAL two-tailed p value</p> <p>DIRECTION the Direction of effect of the SNP in each cohort ( in the order UKBB, FINNGEN, QSKIN, QMEGA 610k, QMEGA OMNI)</p> <p>N sample size</p> <p>See </p>
Multi-ancestry genome-wide association statistics of anxiety
<p><strong>Multi-Ancestry Genome-Wide Association Statistics of Anxiety</strong></p> <p><strong>Citation</strong>: Friligkou E, Lokhammer S, Cabrera Mendoza B, Shen J, He J, Deiana G, Zanoaga MD, Asgel Z, Pilcher A, Di Lascio L, Makharashvili A, Koller D, Tylee D, Pathak GA, Polimanti R. Gene Discovery and Biological Insights into Anxiety Disorders from a Large-Scale Multi-Ancestry Genome-wide Association Study. Nat Genet. doi: 10.1038/s41588-024-01908-2.</p> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.