Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
zenodo52/100

Genome-wide association summary statistics for human blood plasma glycome

<p>The dataset&nbsp;contains results of genome-wide association study of human blood plasma&nbsp;glycome. The 113 files contain association summary statistics for 113 glycome traits, of which 36 were directly measured by UPLC technology and 77 were derived glycome traits. Description of each glycome trait can be found in the <strong>Additional notes</strong> section. This&nbsp;dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org">http://gwasarchive.org</a>.&nbsp;</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Sharapov, S. Z., Tsepilov, Y. A., Klaric, L., Mangino, M., Thareja, G., Shadrina, A. S., &hellip; Aulchenko, Y. (2019). Defining the genetic control of human blood plasma N-glycome using genome-wide association study. <em>Human Molecular Genetics</em>. http://doi.org/10.1093/hmg/ddz054</li> <li>Sodbo Sharapov, Yakov Tsepilov, Lucija Klaric, Massimo Mangino, Gaurav Thareja, Mirna Simurina, Concetta Dagostino, Julia Dmitrieva, Marija Vilaj, FranoVuckovic, Tamara Pavic, Jerko Stambuk, Irena Trbojevic-Akmacic, Jasminka Kristic, Jelena Simunovic, Ana Momcilovic, Harry Campbell, Malcolm Dunlop, Susan Farrington, Maria Pucic-Bakovic, Christian Gieger, Massimo Allegri, Edouard Louis, Michel Georges, Karsten Suhre, Tim Spector, Frances MK Williams, Gordan Lauc, Yurii Aulchenko. (2018). Genome-wide association summary statistics for human blood plasma glycome (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1298406</li> </ol> <p><strong>Funding</strong></p> <p>This work was supported by the European Community&rsquo;s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736) and by the European Structural and Investments funding for the &quot;Croatian National Centre of Research Excellence in Personalized Healthcare&quot; (contract #KK.01.1.1.01.0010).</p> <p>The work of SSh was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.</p> <p>The work of YT was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).</p> <p>Karsten Suhre and Gaurav Thareja are supported by &lsquo;Biomedical Research Program&rsquo; funds at Weill Cornell Medicine - Qatar, a program funded by the Qatar Foundation. We thank all staff at Weill Cornell Medicine - Qatar and Hamad Medical Corporation, and especially all study participants who made the QMDiab study possible.</p> <p>The SOCCS study was supported by grants from Cancer Research UK (C348/A3758, C348/A8896, C348/ A18927); Scottish Government Chief Scientist Office (K/OPR/2/2/D333, CZB/4/94); Medical Research Council (G0000657-53203, MR/K018647/1); Centre Grant from CORE as part of the Digestive Cancer Campaign (<a href="http://www.corecharity.org.uk">http://www.corecharity.org.uk</a>).</p> <p>TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility and Biomedical Research Centre based at Guy&rsquo;s and St Thomas&rsquo; NHS Foundation Trust in partnership with King&rsquo;s College London.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNP: SNP rsID</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build)&nbsp;</li> <li>OTHER_ALLELE: reference allele (coded as &quot;0&quot;)</li> <li>EFFECT_ALLELE: effective allele (coded as &quot;1&quot;)</li> <li>EAF: effective allele frequency&nbsp;</li> <li>N: sample size</li> <li>BETA: effect size of effective allele</li> <li>SE: standard error of effect size</li> <li>PVAL: P-value of association (without GC correction)</li> <li>IMPUTATION: imputation quality</li> </ol>

opencc-by-4.0Jun 2018View details →
zenodo52/100

Genome-wide association summary statistics for human healthspan

<p>The dataset contains genome-wide association summary statistics computed for heathspan. The UKB sub-population of 300,447 genetically Caucasian, British individuals were analyzed. For more details see [1].</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Zenin, A., Tsepilov, Y., Sharapov, S., Getmantsev, E., Menshikov, L. I., Fedichev, P. O., &amp; Aulchenko, Y. (2019). Identification of 12 genetic loci associated with human healthspan. <em>Communications Biology</em>, <em>2</em>(1), 41. http://doi.org/10.1038/s42003-019-0290-0</li> <li>Aleksandr Zenin, Yakov Tsepilov, Sodbo Sharapov, Evgeny Getmantsev, Leonid Menshikov, Peter Fedichev, &amp; Yurii Aulchenko. (2018). Genome-wide association summary statistics for human healthspan (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1302861</li> </ol> <p><strong>Funding</strong></p> <p>The work was supported by Russian Ministry of Science and Education under 5-100 Excellence Programme.&nbsp;<br> The work was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).&nbsp;<br> This research has been conducted using the UK Biobank Resource.&nbsp;<br> The study has been funded by Gero LLC.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNPID - SNP rsID</li> <li>chr - chromosome</li> <li>pos - position (GRCh37 build / hg19)</li> <li>EA - effective allele (coded as &quot;1&quot;)</li> <li>RA - reference allele (coded as &quot;0&quot;)</li> <li>EAF - effective allele frequency</li> <li>beta - effect size of effective allele</li> <li>se - standard error of effect size</li> <li>Z - Z-value of association</li> <li>-log10(p-value) - minus log10(P-value) of association</li> </ol>

opencc-by-4.0Jul 2018View details →
zenodo48/100

Human intestinal Bacteria Collection (HiBC): Isolates and genomes metadata

<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the taxonomy of the isolates, as well as metadata regarding their cultivation and isolation. We also provide metadata regarding the sequencing, genome assembly process and the biological sequences.</p> <p><strong>UPDATE v7</strong>: INSDC accession for <em>Segatella sinensis</em> CLA-AA-H117 was a missing value and is now the correct value of GCA_040324585.2.</p> <p><strong>UPDATE v6:&nbsp;</strong>The growth atmosphere is now indicated by anaerobic or aerobic instead of "Anaerobe/Aerobe" that was a misleading term. The risk group of these two isolates went from 1 to 2:</p> <ul> <li>CLA-AA-H205: <em>Anaerostipes caccae&nbsp;</em></li> <li>CLA-AA-H83: <em>Bacteroides fragilis</em></li> </ul> <p>The risk group of the following isolates has been updated (usually from unknown to 1, or from 2 to 1):</p> <ul> <li>CLA-SR-H026: <em>Aedoeadaptatus acetigenes</em></li> <li>CLA-KB-H139:<em> Bacteroides xylanisolvens</em></li> <li>CLA-SR-H015: <em>Bacteroides xylanisolvens</em></li> <li>CLA-AA-H187: <em>Blautia fusiformis</em></li> <li>CLA-AA-H274: <em>Brotaphodocola catenula</em></li> <li>CLA-AA-H286: <em>Butyricimonas faecihominis</em></li> <li>CLA-AA-H278:<em> Clostridium fessum</em></li> <li>CLA-AA-H147: <em>Dorea ammoniilytica</em></li> <li>CLA-SR-H027: D<em>orea formicigenerans</em></li> <li>CLA-KB-H89: <em>Dorea longicatena</em></li> <li>CLA-KB-H94: <em>Dorea longicatena</em></li> <li>CLA-SR-H022: <em>Enterococcus lactis</em></li> <li>CLA-AA-H250: <em>Hominenteromicrobium mulieris</em></li> <li>CLA-AA-H232: H<em>ominilimicola fabiformis</em></li> <li>CLA-AA-H246: <em>Hominisplanchenecus faecis</em></li> <li>CLA-AA-H276:<em> Hominiventricola filiformis</em></li> <li>CLA-AA-H213:<em> Oliverpabstia intestinalis</em></li> <li>CLA-AA-H241: <em>Oliverpabstia intestinalis</em></li> <li>CLA-AA-H58: <em>Pilosibacter fragilis</em></li> <li>CLA-KB-H110: <em>Ruthenibacterium lactatiformans</em></li> <li>CLA-AA-H174: <em>Segatella sinensis</em></li> <li>CLA-AA-H2: <em>Veillonella parvula</em></li> <li>CLA-AA-H273: <em>Waltera acetigignens</em></li> </ul> <p>Typos in media list have been fixed.&nbsp;</p> <p><strong>UPDATE v5</strong>: The accessions number for the genomes on INSDC databases are added under the column Accession. Plus two typos in the risk group column have been corrected as follow:</p> <ul> <li>CLA-AA-H173: from Risk Group 4 (!) to 2 like the other strain of <em>Sutterella wadsworthensis</em></li> <li>CLA-AA-H198: from Risk Group 4 (!) to 1 like the other <em>Bifidobacterium&nbsp;</em>species.</li> </ul> <p><strong>UPDATE v4</strong>: Only the taxonomy of a couple of isolates has been changed, as follow:</p> <ul> <li>CLA-ER-H4: <em>Collinsella sp900547855</em> instead of <em>Collinsella sp900544645</em></li> <li>CLA-AA-H142: <em>Pilosibacter fragilis</em> (<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov.</em> (<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H58: <em>Pilosibacter fragilis&nbsp;</em>(<em>f__Clostridiaceae</em>)&nbsp;instead of <em>Sakamotonia hominis gen. nov.&nbsp;</em>(<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H89B: <em>Lachnospira intestinalis sp. nov.</em> instead of <em>Lachnospira hominis sp. nov.</em></li> <li>CLA-JM-H10: <em>Lachnospira hominis sp. nov.</em> instead of <em>Lachnospira intestinalis sp. nov.</em></li> <li>CLA-JM-H7B: <em>Faecalibacterium taiwanense</em> instead of <em>Faecalibacterium faecis sp. nov.</em></li> <li>CLA-JM-H45: <em>Merdimmobilis hominis</em> instead of <em>Hominicola intestinalis gen. nov.</em></li> </ul> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

A cross-disorder dosage sensitivity map of the human genome

<p>This repository contains data from Collins et al., <em>A cross-disorder dosage sensitivity map of the human&nbsp;genome</em> (2022), including:</p> <p>1. <strong>Collins_rCNV_2022.dosage_sensitivity_scores.tsv.gz</strong>: This file contains predicted probabilities of haploinsufficiency (pHaplo) and triplosensitivity (pTriplo) for 18,641 autosomal protein-coding genes as defined in Gencode v19.</p> <p>2. <strong>Collins_rCNV_2022.sliding_window_sumstats.tar.gz</strong>: This compressed directory contains rCNV association summary statistics for 54 phenotypes from genome-wide sliding window meta-analyses. Please refer to the README file included in this compressed directory for more details.</p> <p>3. <strong>Collins_rCNV_2022.gene_association_sumstats.tar.gz</strong>: This compressed directory contains rCNV association summary statistics for 54 phenotypes from exome-wide gene-based meta-analyses. Please refer to the README file included in this compressed directory for more details.&nbsp;</p> <p>4. <strong>Collins_rCNV_2022.gene_features_matrix.tar.gz</strong>: This compressed directory contains gene-level feature annotations for 145 features and 18,641 autosomal protein-coding genes. Please refer to the README file included in this compressed directory for more details.&nbsp;</p> <p>Smaller data files have been provided as supplemental tables alongside the publication online.</p> <p>Please also refer to the original publication for details on data sources, study design, methods, and other analyses.</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)

<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., &amp; Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>.&nbsp;</p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p>&nbsp;</p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Eigen scores for human genome assembly GRCh38 Part 4 (Chr1 - Chr2)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Eigen scores for human genome assembly GRCh38 Part 3 (Chr3 - Chr5)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Virus+ Sequence Masked Human Reference Genome (hg19)

<p>A version of the human genome (hg19) originally masked for ribosomal, plant, animal, fungal and&nbsp;low-entropy sequences&nbsp;by Brian Bushnell (<a href="https://zenodo.org/record/1208052#.X5BuTy9h3UI">Bushnell Masked Human Genome</a>) additionally masked for all possible viral sequences.</p> <p>The following commands were used to generate the additional virus sequence masked reference database:</p> <p><strong>1) Download all RefSeq and Neighbor nucleotide records:</strong></p> <p><a href="https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])">https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])</a></p> <p><strong>2) Shred the downloaded viral genomes using shred.sh from the <a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>shred.sh in=refseq_virus_reformated.fasta out=virus_shred.fasta.gz length=85 minlength=75 overlap=30</p> <p><strong>3) Map shredded virus sequence to the hg19-masked human genome using bbmap.sh&nbsp;from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmap.sh ref=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz in=virus_shred.fasta.gz outm=map_human_all_viruses.sam minid=0.90</p> <p><strong>4) Mask virus sequenced mapped regions from the hg19-masked human genome using bbmask.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmask.sh in=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz out=human_virus_masked.fasta.gz sam=map_human_all_viruses<br> .sam</p> <p><strong>5) Remove all N&#39;s to further reduce file size using <a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></strong><br> seqkit -is replace -p &quot;n&quot; -r &quot;&quot; human_virus_masked.fasta.gz &nbsp;&gt; human_virus_masked.fasta_Ns_removed.gz</p> <p><strong>Additional References:</strong></p> <ol> <li><a href="http://seqanswers.com/forums/showthread.php?t=42552">http://seqanswers.com/forums/showthread.php?t=42552</a> for additional information on the original masking of hg19</li> <li><a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a></li> <li><a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></li> <li><a href="https://www.ncbi.nlm.nih.gov/genome/viruses/">NCBI Virus Genome RefSeq</a></li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Analysis of variant-dependent m6A modifications within the Human genome

<p>Interactive and machine-readable results produced by the&nbsp;<a href="https://github.com/cumbof/m6Ad-SNVs" target="_blank" rel="noopener">m6Ad-SNVs</a> tool to asses if m6A-distal SNVs affect DRACH site accessibility, specifically by evaluating the alteration of base-pairing of nucleotides within segments of the DRACH motif.</p> <p>These results contain the predicted m6Ad-SNV candidates with the length of the reference and m6Ad-SNV-containing alternate sequences limited to 250 base pairs. This constraint has been applied to maintain the reliability of the results predicted by RNAFold (<a href="https://www.tbi.univie.ac.at/RNA/">ViennaRNA</a> package). The sequence composition contains up to 100 base pairs from 3'UTRs, with the remaining base pairs limited to the last two exons.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

GenoNet scores for human genome assembly GRCh38

<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files.&nbsp;</p> <p>Each row represents a genomic region with 131 columns. Please find the header line in &quot;genonet.header.txt&quot;.&nbsp;</p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 &nbsp; &nbsp;10000 &nbsp; &nbsp;10025 &nbsp; &nbsp;chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37&nbsp;<a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover&nbsp;<a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Comprehensive 100-bp resolution genome-wide epigenomic profiling data for the hg38 human reference genome

<p>This is a comprehensive collection of diverse epigenomic profiling data in 100-bp resolution with full genome-wide coverage. The datasets are processed from raw read count data collected from five types of sequencing-based assays collected by the Encyclopedia of DNA Elements (ENCODE, <a href="http://www.encodeproject.org">http://www.encodeproject.org</a>) consortium. A total of 6,305 alignment profiles from various high-throughput sequencing assays available on the ENCODE database were preprocessed and filtered according to ENCODE&rsquo;s data standard</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Salmonella enterica serovar Derby isolated from eggs show genomic and phenotypic traits that may be linked to inability to produce human infection.

<p><em><span>Salmonella enterica</span></em><span> serovar Derby causes foodborne disease (FBD) outbreaks worldwide, mainly from contaminated pork but also from chickens. During a major epidemic of FBD in Uruguay due to <em>S</em>. Enteritidis from poultry, we conducted a large survey of commercially available eggs, where we isolated many <em>S.</em> Enteritidis strains but surprisingly also a much larger number (ratio 5:1) of <em>S</em>. Derby strains. No single case of <em>S</em>. Derby infection was detected in that period, suggesting that the <em>S</em>. Derby egg strains were impaired for human infection. We sequenced fourteen of these egg isolates, as well as fifteen isolates from pork or human infection that were isolated in Uruguay before and after that period, and all sequenced strains had the same sequence type <span>(ST40). Phylogenomic genomic analysis was conducted using more than 3500 genomes from the same sequence type (ST), revealing that Uruguayan isolates clustered into four distantly related lineages. Population structure analysis (BAPS) suggested the division of the analyzed genomes into nine different BAPS1 groups, with Uruguayan strains clustering within four of them. </span>All egg isolates clustered together as a monophyletic group and showed marked differences in gene content with the strains in the other clusters. <span>Differences included the absence of a C-terminal fragment of the <em>speF</em> gene, as well as variations in the composition of mobile genetic elements, such as plasmids, insertion sequences, transposons, and phages, between egg isolates and human/pork isolates.</span></span> <span>Egg isolates showed an acid susceptibility phenotype, reduced ability to reach the intestine after oral inoculation of mice, and reduced induction of SPI-2 <em>ssaG</em> gene, compared to human isolates from other monophyletic groups. Mice challenge experiments showed that mice infected intraperitoneally with human/pork isolates died between 1-7 days p.i., while all animals infected with the egg strain survived the challenge. Altogether, our results suggest that loss of gene functions and the absence of plasmids in egg isolates may explain why these <em>S</em>. Derby were not capable of producing human infection despite being at that time, the main serovar recovered from eggs countrywide.</span></p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A comprehensive catalog of exact short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38

<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of exact repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,233,959 TR intervals, covering 0.783306% (24.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta).&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

G-quadruplexes as pivotal components of cis-regulatory elements in the human genome

<p>This repository stores the scripts for analyzing the relationship between G-quadruplexes (G4s) and <em>cis</em>-regulatory elements (CREs), as well as the data generated directly from the manuscript.</p> <p>Manuscript: <a href="https://doi.org/10.1186/s12915-024-01971-5" target="_blank" rel="noopener">G-quadruplexes as pivotal components of <em>cis</em>-regulatory elements in the human genome</a></p> <p>G4Hunter_w25_s1.5_hg38.txt: All potential G-quadruplexes in the human genome predicted by the G4Hunter software.&nbsp;</p> <ul> <li>Genome assembly: hg38.</li> <li>G4Hunter software parameters were set as follows: score threshold 1.5, window size 25.</li> </ul> <p>G4_cCRE_annotation.txt: Annotation file indicating the presence of G4s in cCREs (candidate CREs; from <a title="SCREEN database" href="https://screen.encodeproject.org/" target="_blank" rel="noopener">SCREEN database</a>).</p> <p>scripts.zip: Source code used for data analysis in this project, based on the R language.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

GenoNet scores for human genome assembly GRCh37

<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files.&nbsp;</p> <p>Each row represents a genomic region with 131 columns. Please find the header line in &quot;genonet.header.txt&quot;.&nbsp;</p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 &nbsp; &nbsp;10000 &nbsp; &nbsp;10025 &nbsp; &nbsp;chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37&nbsp;<a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover&nbsp;<a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Recombination and aneuploidy data for: Insights about variation in meiosis from 31,228 human sperm genomes

<p>This repository holds crossover and aneuploidy data for 31,228 human sperm genomes sequenced with Sperm-seq. Data are described in the preprint/paper &quot;Insights about variation in meiosis from 31,228 human sperm genomes.&quot;&nbsp;Several levels of data are available, e.g., each crossover (allcrossovers_hg38.txt.gz) and the numbers of crossovers detected per sperm cell (numbercrossoverspercell.txt.gz). Each data file is described in its own readme, so&nbsp;README_allgainsdivisionoforigin.txt explains the data present in&nbsp;allgainsdivisionoforigin.txt. All files are either text files or text files compressed via gzip. All data was generated as described in the preprint/paper, using scripts available in the companion repository (most recent version DOI: 10.5281/zenodo.3561080).</p> <p>This new version has updated sex chromosome ploidies after identifying a minor coding error affecting ~24 cells (sexchromosomeploidy.txt.gz) and includes more detail about possible crossover error modes (README_allcrossovers_hg38.txt.gz).</p>

opencc-by-4.0Apr 2019View details →
zenodo44/100

A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects

<p>Unified, inferred tree sequences built from&nbsp;the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects. Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included).&nbsp;Tree sequences were inferred using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.2.1,&nbsp;dated using&nbsp;<a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4&nbsp;and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on&nbsp;<a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using&nbsp;<a href="https://tskit.readthedocs.io/">tskit</a>.&nbsp;</p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with nodes contain&nbsp;the mean and variance of tsdate&#39;s posterior distribution on node time. To access these values, we can use:</p> <pre><code>import json node = ts.node(10000) metadata_dict = json.loads(node.metadata) print("The mean of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["mn"])) print("The variance of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["vr"]))</code></pre> <p>Age estimates for&nbsp;each variant site can be derived from the mean of the age estimates of the&nbsp;upper and lower bounding nodes of the oldest mutation associated with a site. tsdate includes <a href="https://tsdate.readthedocs.io/en/latest/python-api.html?highlight=sites_time_from_ts#tsdate.sites_time_from_ts">a function to find the age estimates of all sites in the tree sequence</a>:</p> <pre><code>import tsdate site_times = tsdate.sites_time_from_ts(ts, node_selection='arithmetic')</code></pre> <p>This returns a numpy array which has a length equal to the number of sites.</p> <p>Accessing variant sites in the tree sequence provides&nbsp;the position and id of variants:</p> <pre><code>site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>)&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Genome-wide characterization of human minisatellite VNTRs: population-specific alleles and gene expression differences

<p>This repository consists of minisatellite VNTR genotypes for 2,800 samples (2,770 individuals). The raw VCF files were produced using <a href="https://github.com/yzhernand/VNTRseek">VNTRseek</a>&nbsp;on xxx data sources: <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000_genomes_project/">30 high coverage WGS datasets</a>&nbsp;from the 1000 Genomes Project phase 3, <a href="https://www.internationalgenome.org/data-portal/data-collection/30x-grch38">2,504 unrelated genomes</a> from New York Genome Center (NYGC), <a href="https://www.internationalgenome.org/data-portal/data-collection/sgdp">253 genomes from Simons Diversity Genome Project</a>&nbsp;(SGDP), <a href="https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/tumor-normal.html">two tumor-normal breast cancer samples</a>&nbsp;from Illumina Basespace, haploid genomes <a href="https://www.ncbi.nlm.nih.gov/sra/SRX652547">CHM1 </a>and <a href="https://www.ncbi.nlm.nih.gov/sra/SRX1009644">CHM13</a>, and seven genomes from the Personal Genome Project from the Genome In A Bottle Consortium (GIAB). Raw VCF files are provided for each data source separately.</p> <p>The raw VCF files were preprocessed (preprocess.sh) to extract genotypes and provided in VNTRseek_preprocessed_data.tar.gz (uncompressed size 10G). The R Markdown code to analyze the preprocessed data and produce figures and tables is also provided (tables_and_figures.Rmd). For more information see the ReadMe file.</p> <p>This work was supported in part by NSF grants IIS-1423022 and DBI-1559829.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Data for: Regularized sequence-context mutational trees capture variation in mutation rates across the human genome

<p>Additional data on output models from Bayer as reported in:</p> <p>Regularized sequence-context mutational trees capture variation in mutation rates across the human genome</p> <p>Adams CJ, Conery M, Auerbach BJ, Jensen ST, Mathieson I, Voight BF. BioRxiv&nbsp;https://doi.org/10.1101/2022.10.14.512160</p> <p>Accepted, PLoS Genetics.&nbsp;</p> <p>Code Available at:&nbsp;https://github.com/bvoightlab/Baymer</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Genomic determinants of pathogenicity in SARS-CoV-2 and other human coronaviruses

<p><strong>Dataset S1.</strong>Complete nucleotide sequence alignment of all human CoV used for region identification.&nbsp;</p> <p><strong>Dataset S2.</strong>Complete nucleotide sequence alignment of all CoV (of human and non-human hosts).</p> <p><strong>Dataset S3.</strong>Distances between leaves (each CoV strain in Dataset S2 was considered), from every reference genome of each of the seven human CoV.</p> <p><strong>Dataset S4.</strong>Alignment of strains used for zoonotic jump analysis.</p>

opencc-by-4.0May 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record