Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

86

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

86 results for “hg38”

Learn how ShareScore rates datasets ↗
zenodo48/100

Archaic introgression for HGDP and 1000genomes in hg38

<p>These files contain the infered positions of introgressed archaic sequence in 1000genomes and HGDP datasets.</p> <p>The segments are identified using hmmix (https://github.com/LauritsSkov/Introgression-detection). Datasets are phased so segments are infered for each haplotype. Each datasets will have two files: a *segments.txt file and a *SNPS.txt file.</p> <p>&nbsp;</p> <p>&gt; The columns in the segments.txt file are:</p> <p><strong>name: </strong>name of individidual<br><strong>haplotype:</strong> either hap1 or hap2 (if a genotype is 0|1 then 0 will be on hap1 and 1 will be on hap2)<br><strong>pop: </strong>population from HGDP or 1000 genomes<br><strong>region:&nbsp;</strong>region from HGDP or 1000 genomes - can be AMERICA, CENTRAL_SOUTH_ASIA, EAST_ASIA, EUROPE, MIDDLE_EAST or OCEANIA<br><strong>chrom:</strong> chromosome in hg38 - X chromosome is not included<br><strong>start</strong>: start coordinate of introgressed segment in hg38<br><strong>end</strong>: end coordinate of introgressed segment in hg38<br><strong>mean_prob:</strong> Mean posterior probability that a segment is archaic according to hmmix (I usually recommend doing a cutoff at 0.8)<br><strong>ND_type:</strong> Which sequenced archaic does the segments share more derived SNPs with. Can be Both, Denisova, Neanderthal or none<br><strong>snps:</strong> Number of derived SNPs on segment NOT seen in Sub saharan Africa<br><strong>admixpopvariants:</strong> How many derived SNPs are shared with a sequenced arhaic genome<br><strong>Altai:</strong> How many derived SNPs are shared with the Altai Neanderthal (Denisova5)<br><strong>Vindija:</strong> How many derived SNPs are shared with Vindija Neanderthal (Vindija33.19)<br><strong>Denisova:</strong> How many derived SNPs are shared with Denisova (Denisova3)<br><strong>Chagyrskaya:</strong> How many derived SNPs are shared with Chagyrskaya Neanderthal (Chagyrskaya8)<br><strong>variants:</strong> List of derived SNPs on segment NOT seen in Sub saharan Africa&nbsp;</p> <p>&nbsp;</p> <p>&gt; The columns in the SNPS.txt file are:</p> <p><strong>chrom:</strong> chromosome in hg38 - X chromosome is not included<br><strong>pos:</strong> position of SNP in hg38 coordinates<br><strong>snptype:</strong> can be shared derived with archaic (DAV), in high LD&nbsp;<br><strong>ancestralbase:</strong> what is the ancestral base<br><strong>derivedbases: </strong>What is the derived base (there can be multiple but &gt;99% are bilallelic)<br><strong>freq_in_dataset:</strong> Frequency of most common derived base (in percent so the number is between 0 and 100)<br><strong>ND:</strong> derived in either Denisovans only (ND01), Neanderthals only (ND10), derived in both Neanderthals and Denisovans (ND11) or none<br><strong>sharedwith: </strong>Which archaic genomes is the derived allele(s) shared with. This does not only include the four high coverage archaics</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Transcription initiation peaks based on FANTOM5 CAGE data on hg38 and mm10

<p><strong>Overview</strong></p> <p>Decomposition-based peak identification (DPI, https://github.com/hkawaji/dpi1) is applied to the re-processed (re-aligned) FANTOM5 data, upon hg38 (GRCh38) and mm10 (GRCm38), obtained from below:</p> <ul> <li>http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v1/basic/</li> <li>http://fantom.gsc.riken.jp/5/datafiles/reprocessed/mm10_v1/basic/</li> </ul> <p>The same parameters to the ones used in the previous paper (Forrest ARR, Kawaji H, Rehli M, et al. Nature 507: 462–470, 2014) was used.</p> <p> </p> <p><strong>Data files</strong></p> <p>Four data files per assembly are prepared as below.</p> <ol> <li>tag cluster in the original definition (*.tc.bed.gz)</li> <li>full set of DPI peaks (*.tc.decompose_smoothing_merged.bed.gz)</li> <li>permissive set of DPI peaks (*.tc.decompose_smoothing_merged.ctssMaxCounts3.bed.gz)</li> <li>robust set of DPI peaks (*.tc.decompose_smoothing_merged.ctssMaxCounts11_ctssMaxTpm1.bed.gz)</li> </ol> <p> </p> <p><strong>Acknowledgement</strong></p> <p>This data set is supported by Research Grant from MEXT to RIKEN Preventive Medicine and Diagnosis Innovation Program, RIKEN Center for Life Science Technologies, and JSPS KAKENHI Grant-in-Aid for Scientific Research No. 16H02902.</p>

opencc-by-4.0Apr 2017View details →
zenodo44/100

Comprehensive 100-bp resolution genome-wide epigenomic profiling data for the hg38 human reference genome

<p>This is a comprehensive collection of diverse epigenomic profiling data in 100-bp resolution with full genome-wide coverage. The datasets are processed from raw read count data collected from five types of sequencing-based assays collected by the Encyclopedia of DNA Elements (ENCODE, <a href="http://www.encodeproject.org">http://www.encodeproject.org</a>) consortium. A total of 6,305 alignment profiles from various high-throughput sequencing assays available on the ENCODE database were preprocessed and filtered according to ENCODE&rsquo;s data standard</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Curated GWAS summary statistics on African ancestry on 19 blood count traits and glycemic traits (hg38)

<p>Genome wide curated summary statistics&nbsp;on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>).&nbsp;GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Annotation of inverted repeats displaying features of pble STIR or IR in the hg38 genome model

<p>Annotation of inverted repeats displaying features of pble STIR or IR in the hg38 genome model. The annotation of <em>pble</em>-like inner inverted repeats was done using Palindrome (EMBOSS package). The output file was then filtered using pal2gff (https://github.com/Leelouh/pal2gff/blob/main/pal2gff.py), using as parameters a repeat size between 5 and 15 nucleotides, a spacer between pairs of inverted repeats (IRs) of 2 to 10 nucleotides, and a number of mismatches within repeats ranging from 0 to 1. These parameters were chosen taking into account those of the inner IRs found at ends of invertebrate pbles.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

mcrpc_wgbs_hg38_chr11

<p>A HDF5-backed RangedSummarizedExperiment for WGBS Data (hg38 CpG sites) for chr11&nbsp;for 100 castration-resistant prostate cancer metastases from the paper&nbsp;&#39;Zhao, Shuang G., et al. &quot;The DNA methylation landscape of advanced prostate cancer.&quot;&nbsp;<em>Nature genetics</em>&nbsp;52.8 (2020): 778-789.&#39;.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Putative human enhancers based on ReMap 2018 (hg38)

<p>To generate a collection of putative enhancer regions, we collected all transcription factor ChIP-seq peaks from ReMap 2018 (ref remap; <a href="http://remap.univ-amu.fr/storage/remap2018/hg38/MACS/remap2018_all_macs2_hg38_v1_2.bed.gz">http://remap.univ-amu.fr/storage/remap2018/hg38/MACS/remap2018_all_macs2_hg38_v1_2.bed.gz</a>). We took the summit of all peaks and extended these 25 bp up- and downstream. Based on this file, we generated a coverage bedGraph using bedtools genomecov. We performed peak calling on this bedGraph file using bdgpeakcall from MACS2 2.2.7.1, with the following settings: l=50 and g=10. We performed the peak calling twice, setting c to 4 and 30, respectively. All peaks from c=30 were combined with all peaks of c=4 that did not overlap with the peaks of c=30. We then removed all regions on chrM and extended the summit of the peaks 100 bp up- and downstream to generate a final collection of 1,268,775 putative enhancers of 200 bp.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Gene expression and splicing counts from 49 tissues from GTEx v8 genome build hg38 - non-strand specific

<p><strong>Dataset description:</strong></p> <p>49 folders, each corresponding to one tissue from GTEx v8 and containing the following files:</p> <ol> <li> <p>geneCounts: gene-level counts&nbsp;</p> </li> <li> <p>k_j: split counts spanning from one exon to another.</p> </li> <li> <p>k_theta: non-split counts covering a splice site</p> </li> <li> <p>n_psi3: total split counts from a given acceptor site</p> </li> <li> <p>n_psi5: total split counts from a given donor site</p> </li> <li> <p>n_theta: total split and non-split counts for a given splice site</p> </li> <li> <p>Sample annotation describing each sample from the dataset</p> </li> <li> <p>Description file with global information from the dataset</p> </li> </ol> <p>The gene counts were originated using the GTF file from&nbsp;<a href="https://www.gencodegenes.org/human/release_29.html">release 29 of GENCODE</a>, and the split and non-split counts contain only the annotated junctions from the same release.&nbsp;Statistics are reported only for GENCODE-annotated introns and splice sites, in compliance with the regulations of the GTEx consortium. For a description of the samples, methods, and protocols, see the GTEx publication specified below.</p> <p><strong>Use:&nbsp;</strong>The count matrices are intended to help researchers that are interested in using RNA-Seq data with the purpose of diagnostics. Researchers can merge their own dataset with the downloaded ones, provided the tissue, genome build, strand, and paired-end specifications match. Afterwards, the&nbsp;<a href="https://github.com/gagneurlab/drop">Detection of RNA outliers Pipeline (DROP)&nbsp;</a>&nbsp;can be used to compute gene expression and splicing outliers.<br> <strong>Organism:</strong>&nbsp;Homo sapiens<br> <strong>Genome assembly:</strong>&nbsp;hg38<br> <strong>Gene annotation:</strong>&nbsp;gencode29<br> <strong>Strand specific:&nbsp;</strong>FALSE<br> <strong>Paired end:&nbsp;</strong>TRUE<br> <strong>Protocol:&nbsp;</strong>poly(A) enrichment</p> <p><strong>Contact:</strong> Vicente A. Yepez, yepez at in.tum.de; Christian Mertes, mertes at in.tum.de; Julien Gagneur, gagneur at in.tum.de</p> <p><strong>Citation:</strong> Write the following in the &quot;Data availability&quot; section of the manuscript or similar replacing the three citations by the ones from the References section below:</p> <blockquote> <p><strong>The count matrices for the GTEx samples &lt;cite GTEx publication,&nbsp;see below&gt;&nbsp;were downloaded from Zenodo (doi: 10.5281/zenodo.6078397) and were generated through DROP &lt;cite DROP, see below&gt;&nbsp;using the release 29 of the GENCODE annotation &lt;cite GENCODE, see below&gt;. </strong></p> </blockquote> <p>Also, write the following in the Acknowledgements section:<br> &nbsp;</p> <blockquote> <p><strong>The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health, and by NCI, NHGRI, NHLBI, NIDA, NIMH, and NINDS. The raw data used for the analyses described in this manuscript were obtained from the GTEx Portal on June 12, 2017, under accession number dbGaP &nbsp;phs000424.v8.p2.</strong></p> </blockquote> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome

<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the&nbsp;GRCh38 (hg38) assembly of the human genome&nbsp;using&nbsp;the&nbsp;<a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing &amp; Analysis Pipeline</a>&nbsp;(details in the documentation on GitHub).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

SpliceMaps for a variety of human tissues (hg38)

<p>SpliceMaps are tissue-specific splice-site annotations quantifying splice-site usage and modeling isoform competition. SpliceMaps can be used to detect aberrant splicing across human tissues with the aberrant splicing prediction model AbSplice. The pre-computed SpliceMaps provided here are under MIT license.</p>

openmit-licenseMar 2022View details →
zenodo40/100

bksnake_reference_genomes_2022-07-15_hg38

<p>Human hg38 reference genomes and annotations files from RefSeq/NCBI and Ensembl. Archived in a tar.gz file.</p> <p>Files are processed and formatted so that they can be used together with the bulk RNASeq pipeline &quot;bksnake&quot;.</p> <p>STAR index files for the genome file are&nbsp;also included.</p> <p>The tar.gz file contains the following data folders</p> <ul> <li>genomes_2022-07-15_hg38/fasta</li> <li>genomes_2022-07-15_hg38/star_2.7.10b</li> <li>genomes_2022-07-15_hg38/gtf/ensembl</li> <li>genomes_2022-07-15_hg38/gtf/refseq</li> </ul> <p>&quot;fasta&quot; contains fasta genomic sequences and index files, as well as ribosomal transcripts &quot;intervals&quot; and chromosome &quot;dictionary&quot; files. &quot;star_2.7.10b&quot; contains STAR aligner index files. &quot;gtf/ensembl&quot; and &quot;gtf/refseq&quot; contain Ensembl and RefSeq gtf files as well as &quot;genePred&quot; and &quot;refFlat&quot; files.</p> <p>The recipe how to generate these files is going to be published in github.</p> <p>The bulk RNASeq pipeline &quot;bksnake&quot; is going to be published in github.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

cpgea_wgbs_hg38

<p>A HDF5-backed RangedSummarizedExperiment for WGBS Data (hg38 CpG sites)&nbsp;for 187 pairs of primary prostate tumours and matching normal prostate samples&nbsp;from the paper &#39;Li, Jing, et al. &quot;A genomic and epigenomic atlas of prostate cancer in Asian populations.&quot;&nbsp;<em>Nature</em>&nbsp;580.7801 (2020): 93-99.&#39;.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

hg38 Jannovar data files

<p>Transcript files for Jannovar (PMID:24677618).</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

FANTOM5 transcribed enhancers in hg38

<p><strong>Overview</strong></p> <p>Transcribed enhancers were identified and their expression was quantified across all human FANTOM5 libraries, following the re-aligned FANTOM5 CAGE data upon hg38 (GRCh38) (obtained from http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v1/basic/), and decomposition-based peak identification (obtained from https://zenodo.org/record/545682#.WPuNy1Pyv2Q) by Kawaji, Hideya.</p> <p> </p> <p><strong>Description</strong></p> <p>Transcribed enhancers were called based on bidirectional balanced RNA signatures as per Andersson et al (2014). Enhancers were only identified distal to known exons (+/-100bp region from boundaries) and transcription start sites (+/-300bp), defined by GENCODE v24 annotation. In total, 63,285 enhancers were identified across 1,829 libraries. The expression was quantified and TPM (tags per million) normalised according to the total number of mapped reads within the full set of TCs. For details regarding the identification of transcribed enhancers from CAGE data, please see Andersson et al (2014) and blog post.</p> <p>Due to varying noise levels across FANTOM5 libraries and the intrinsic low expression levels of transcribed enhancers, library-specific noise levels were estimated to define of robust set of enhancers in each sample. In summary, for each library, expression was quantified in randomly sampled genomic regions distal to assembly gaps, DNase hypersensitive sites (ENCODE), known exons and gene TSSs (GENCODE) to create a genomic background expression distribution. For each library, we called an enhancer active (used) if its expression was above the 99.9th quantile of the library’s genomic background expression distribution. The robust set of enhancers consist of 60,215 over 1,829 libraries, being significantly expressed in at least one library.</p> <p>While this approach ensures less permissive enhancer calling in noisy libraries, for some libraries the noise threshold is zero meaning that a single CAGE tag is sufficient for calling an enhancer active. Furthermore, the possibility of detecting enhancer transcription is affected by sequencing depth, so the number of active enhancers per library might not be biologically meaningful to compare when sequencing depths differ.</p> <p> </p> <p><strong>Data files</strong><br> Each predicted enhancer is described in BED12 format with two blocks denoting the merged regions of transcription initiation on the minus and plus strands. The thickStart and thickEnd columns denote the inferred mid position between blocks of transcription initiation events. Expression and usage matrices are tab delimited and the first row gives the FANTOM5 CNhs IDs and the first column the enhancer ID (same as column 4 in BED file). Usage matrices contain zeroes and ones (0:not used, 1:used).</p> <ul> <li>enhancers (BED12 format)</li> <li>enhancer expression matrix (tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> <li>enhancer expression matrix TPM normalized (tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> <li>binary enhancer usage matrix (0:not used, 1:used, tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> </ul>

opencc-by-4.0Apr 2017View details →
zenodo36/100

Matrix N, TEs vs promoters, 100kb-wide windows, hg38

<p>Regulatory susceptibility matrix N, with rows as hg38 protein coding genes and columns as TE subfamilies. A TE is counted if it is found within 100kb of a gene promoter, strictly out of any promoter and any exon belonging to that gene. Since hg38 contains inflated SVA integrants, most of which small fragments unlikely to be of functional importance, we filtered out SVA integrants shorter than 200bp.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Matrix N, TEs vs promoters, weighted with L = 2.5e5kb, hg38

<p>Regulatory susceptibility matrix N, with rows as hg38 protein coding genes and columns as TE subfamilies. For any given gene, regulatory TEs are strictly out of any promoter and any exon belonging to that gene. The contribution of each single TE is weighted by a gaussian kernel centered on the closest promoter of that gene, with a bandwidth L = 2.5e5kb. Since hg38 contains inflated SVA integrants, most of which small fragments unlikely to be of functional importance, we filtered out SVA integrants shorter than 200bp.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

(1000G phase 3 hg38) Linkage Desequilibrium Matrices by ancestry for RAISS

<p>Linkage desequilibrium matrices generated using 1000G phase 3 data hg38 using plink</p> <p>These matrices can use as input file for the RAISS python packages, which allows for the imputation of GWAS summary statistique.</p> <p>RAISS Python package:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/raiss</p> <p>RAISS paper:</p> <p>https://academic.oup.com/bioinformatics/article/35/22/4837/5512360?login=true&nbsp;</p> <p>Preprint providing RAISS performances by ancestry (Table S3):</p> <p>https://www.biorxiv.org/content/10.1101/2023.06.23.546248v1</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

EBAII - Linux training toy dataset - bed file of hg38 exons

<p>This is a toy dataset for the Linux training session at EBAII (Initiation au traitement des donn&eacute;es de g&eacute;nomique obtenues par s&eacute;quen&ccedil;age &agrave; haut d&eacute;bit)</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

SICILIAN_human_hg38_Refs

<p>SICILIAN pipeline&nbsp;(https://github.com/salzmanlab/SICILIAN)&nbsp;index,&nbsp;reference, and annotation files for Human hg38&nbsp;assembly.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Helperfiles for using hmmix for calling archaic introgression into present day humans (both hg19 and hg38)

<p>These files are:</p> <p>1) Strict callability masks from 1000 genomes project in hg19 and hg38 coordinates</p> <p>2) Outgroup files:</p> <p>hg38_Outgroup_1000g_HGDP.txt: Frequencies of derived alleles in 490 present individuals with Sub-Saharan related ancestry: 426 from 1000genomes project and 64 from HGDP (total=490) Only first two columns are used by hmmix. The remaining columns what the reference base (hg38 refgenome), ancestral base, derived bases and the frequency of the derived bases in HGDP and 1000genomes</p> <p>hg19_Outgroup_1000g.txt: Frequencies of derived alleles in 292 present individuals with Sub-Saharan related ancestry from 1000genomes</p> <p>3) The mutation rate files are based on the outgroup files. They report the mutation rate in 1 Mb window scaled by the genomewide mutation rate&nbsp;</p> <p>4) The reference genome for hg19 and hg38</p> <p>5) The ancestral allele calls for hg19 and hg38</p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record