Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.9.0
Dataset results
86 results for “hg38”
Major and minor allele genomes for hg38 using dbSNP151 with precomputed Bowtie and BWA indexes
<p>Minor and major allele genomes for the hg38 reference genome (GRCh38.p12) built using common SNPs from dbSNP release 151. Processing details are described in https://doi.org/10.1101/2022.04.21.488824.</p> <p>The following files are included:</p> <ul> <li>Fasta files for both the major and minor allele genomes</li> <li>Bowtie indexes for both the major and minor allele genomes</li> <li>BWA indexes for both the major and minor allele genomes</li> </ul> <p>The genomes are also available as Bioconductor BSgenome packages:</p> <ul> <li>https://bioconductor.org/packages/release/data/annotation/html/BSgenome.Hsapiens.UCSC.hg38.dbSNP151.major.html</li> <li>https://bioconductor.org/packages/release/data/annotation/html/BSgenome.Hsapiens.UCSC.hg38.dbSNP151.minor.html</li> </ul> <p> </p>
genomecomb additional files (cadd, minimap2) for reference data for Homo sapiens (hg38) version 0.108.0
Open the record for dataset details and reuse information.
Archaic variants from Altai, Vindija, Chagyrskaya and Denisova (hg38)
<p>This is the variants for Altai, Vindija, Chagyrskaya and Denisova lifted over from hg19 to hg38.</p> <p>Liftover with CrossMap.py</p> <p>CrossMap.py vcf {chain} {vcffile} {refgenome} {outfile} --no-comp-alleles</p>
hg38 reference and annotation files
<p>This repo contains reference and annotation files for hg38. We are following the [TOPMed pipeline](https://github.com/broadinstitute/gtex-pipeline/blob/master/TOPMed_RNAseq_pipeline.md). Reach out to Arushi Varshney at arushiv AT umich DOT edu if you have any questions.</p> <p>Files:</p> <p>1. bwa index = bwa.tar.gz</p> <p>2. star index = star.tar.gz</p> <p>3. ENCODE blacklist = blacklist.tar</p> <p>4. gencode v30 annotations = gencode.tar.gz</p> <p>5. containers with STAR (RNA) and BWA (ATAC) = containers.tar.gz</p> <p>Notes on these files:<br> ### hg38 fasta:<br> I downloaded the TOPMed fasta tar [Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.tar.gz](https://personal.broadinstitute.org/francois/topmed/Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.tar.gz) as use this as-is. The TOPMed GitHub describes that they obtained the Broad institute's GRCh38 reference, removed ALT, HLA and Decoy contigs, and added ERCC spike-in reference annotations. Refer to their [README](https://github.com/broadinstitute/gtex-pipeline/blob/master/TOPMed_RNAseq_pipeline.md) for more details. They don't mention PARs but we checked the reference files and both chrY PARs are hard masked - as [ENCODE](https://www.encodeproject.org/files/GRCh38_no_alt_analysis_set_GCA_000001405.15/) also recommends.<br> ### Gencode v30 gene annotations: gencode.tar.gz<br> I downloaded the file [gencode.v30.annotation.gtf.gz](https://www.gencodegenes.org/human/release_30.html) from the gencode website, and downloaded the file [ERCC92.genes.patched.gtf](https://personal.broadinstitute.org/francois/resources/). I then appended the ERCC patched gtf to the gencode annotation gtf<br> ```<br> gunzip gencode.v30.annotation.gtf.gz<br> cat gencode.v30.annotation.gtf ERCC92.genes.patched.gtf > gencode.v30.annotation.ERCC92.gtf<br> ```<br> ### STAR index: star.tar.gz; container with star in containers.tar.gz<br> A STAR index is shared on the TOPMed GitHub, but it was generated for STAR version STAR_2.6.1d. Since I've been using the version 2.7.3a, I followed their steps to generate the STAR reference again. I used the gencode gtf described above and generated the STAR index.</p> <p>```<br> STAR --runMode genomeGenerate --genomeDir STAR_genome_GRCh38_noALT_noHLA_noDecoy_ERCC_v30_test --genomeFastaFiles Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.fasta --sjdbGTFfile gencode.v30.annotation.ERCC92.gtf --sjdbOverhang 100 --runThreadN 10<br> ```<br> ### BWA index: bwa.tar.gz<br> I generated the BWA index using the fasta above<br> ```</p> <p>ln -s Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.fasta hg38.fa<br> bwa index hg38.fa<br> ```<br> <br> ### ENCODE Blacklist: blacklist.tar<br> I used the blacklist [here](https://theparkerlab.med.umich.edu/data/arushiv/hg38_references_annots/blacklist/) that I obtained from this [Kundaje website](https://sites.google.com/site/anshulkundaje/projects/blacklists).</p>
mcrpc_wgbs_hg38
<p>A HDF5-backed RangedSummarizedExperiment for WGBS Data (hg38 CpG sites) for 100 castration-resistant prostate cancer metastases from the paper 'Zhao, Shuang G., et al. "The DNA methylation landscape of advanced prostate cancer." <em>Nature genetics</em> 52.8 (2020): 778-789.'. </p>
Geographic allele frequency variation in the 1000 Genomes hg38 NYGC dataset
Open the record for dataset details and reuse information.
QDNAseq.hg38: QDNAseq bin annotation for the human genome build hg38
<p><strong>QDNAseq</strong> bin annotations of size 1, 5, 10, 15, 30, 50, 100, 500, and 1000 kbp for the human genome build hg38.</p>
DOHH2 hg38 H3K4me1 ChIP-seq Dataset filtered for Unique Multiread Mappability
<p>DOHH2 hg38 H3K4me1 ChIP-seq Dataset filtered for Unique Multiread Mappability</p>
Database files (hg19 and hg38) and test files for RTpred
<p>Database files (hg19 and hg38) and test files for RTpred</p>
hg38 syntenic ages
<p>### METHOD ###</p> <p>Sequence ages were estimated from hg38 100-way multiZ vertebrate sequence alignments from the UCSC genome browser. Briefly, sequence ages were estimated as the branch length between humans (hg38) and the oldest most recent common ancestor (MRCA) from the 100-way neutral tree. </p> <p> </p> <p>MRCA to taxon mappings are available in the hg38_syn_taxon.bed file. </p> <p> </p> <p># Hg38 100-way vertebrate tree syntenic block age files. </p> <p>Information for hg38 100-way vertebrate multiple sequence alignment can be found here - http://hgdownload.cse.ucsc.edu/goldenPath/hg38/multiz100way/README.txt</p> <p># Scripts generating these files</p> <p>(1) Creating syntenic block .bed files</p> <p>https://github.com/slifong08/enh_ages/blob/54084ff12e521c20cc5288d6cb3590fac93c11ed/age_arch/manuscript_scripts/get_spec_count_msa-hg19.py</p> <p>(2) Assigning most recent common ancestor (MRCA) and patristic distances to syntenic blocks</p> <p>https://github.com/slifong08/enh_ages/blob/54084ff12e521c20cc5288d6cb3590fac93c11ed/age_arch/manuscript_scripts/get_synteny_age_hg19.py</p>
genomecomb reference data for Homo sapiens (hg38) version 0.98.7
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg38) version 0.11.0
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg38) version 0.108.0
Open the record for dataset details and reuse information.
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)
<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs "170630_M00528_0292_000000000-B9JY8" (aka "NC_LIMMS") and "180221_M00528_0334_000000000-B6PJM" (aka "NC_LIMMS2"). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p> </p> <p><em><strong>"170630_M00528_0292_000000000-B9JY8" ("NC_LIMMS") :</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>1 iPSC_control_rep1 4 ACAGAT NNNNNNNN</p> <p>2 iPSC_control_rep2 24 ATCGTG NNNNNNNN</p> <p>3 iPSC_control_rep3 31 CACGAT NNNNNNNN</p> <p>4 S3P1_OK_rep1 36 CACTGA NNNNNNNN</p> <p>5 S3P1_OK_rep2 46 CTGACG NNNNNNNN</p> <p>6 S3P1_OK_rep3 63 GAGTGA NNNNNNNN</p> <p>7 S4P1_OK_rep1 79 GTATAC NNNNNNNN</p> <p>8 S4P1_OK_rep2 92 TCGAGC NNNNNNNN</p> <p>9 S4P1_OK_rep3 9 ACATGA NNNNNNNN</p> <p>10 S4P2_OK_rep1 21 ATCATA NNNNNNNN</p> <p>11 S4P2_OK_rep2 33 CACGTG NNNNNNNN</p> <p>12 S4P2_OK_rep3 45 CGATGA NNNNNNNN</p> <p>13 S1P1_rep1 57 GAGATA NNNNNNNN</p> <p>14 S1P1_rep2 69 GCTCTC NNNNNNNN</p> <p>15 S1P1_rep3 81 GTATGA NNNNNNNN</p> <p>16 S3P1_FAILED_rep1 93 TCGATA NNNNNNNN</p> <p>17 S3P1_FAILED_rep2 11 AGTAGC NNNNNNNN</p> <p>18 S3P1_FAILED_rep3 23 ATCGCA NNNNNNNN</p> <p>19 S4P1_FAILED_rep1 35 CACTCT NNNNNNNN</p> <p>20 S4P1_FAILED_rep2 47 CTGAGC NNNNNNNN</p> <p>21 S4P1_FAILED_rep3 59 GAGCGT NNNNNNNN</p> <p>22 S4P2_FAILED_rep1 71 GCTGCA NNNNNNNN</p> <p>23 S4P2_FAILED_rep2 83 TATAGC NNNNNNNN</p> <p>24 S4P2_FAILED_rep3 95 TCGCGT NNNNNNNN</p> <p> </p> <p><em><strong>"180221_M00528_0334_000000000-B6PJM" ("NC_LIMMS2"):</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>25 PETRI_rep1 04 ACAGAT NNNNNNNN</p> <p>26 PETRI_rep2 24 ATCGTG NNNNNNNN</p> <p>27 PETRI_rep3 31 CACGAT NNNNNNNN</p> <p>28 BIOCHIP_E_rep1 6 CACTGA NNNNNNNN</p> <p>29 BIOCHIP_M_rep1 46 CTGACG NNNNNNNN</p> <p>30 BIOCHIP_S_rep1 63 GAGTGA NNNNNNNN</p> <p>31 BIOCHIP_E_rep2 79 GTATAC NNNNNNNN</p> <p>32 BIOCHIP_M_rep2 92 TCGAGC NNNNNNNN</p> <p>33 BIOCHIP_S_rep2 09 ACATGA NNNNNNNN</p> <p>34 BIOCHIP_E_rep3 21 ATCATA NNNNNNNN</p> <p>35 BIOCHIP_M_rep3 33 CACGTG NNNNNNNN</p> <p>36 BIOCHIP_S_rep3 45 CGATGA NNNNNNNN</p> <p>37 HEPATOCYTES_rep1 57 GAGATA NNNNNNNN</p> <p>38 HEPATOCYTES_rep2 69 GCTCTC NNNNNNNN</p> <p>39 iPSC_control_rep1-2 81 GTATGA NNNNNNNN</p> <p>40 BIOCHIP_E_rep2-2 93 TCGATA NNNNNNNN</p> <p>41 BIOCHIP_M_rep1-2 11 AGTAGC NNNNNNNN</p> <p>42 BIOCHIP_S_rep2-2 23 ATCGCA NNNNNNNN</p> <p> </p>
Human genome annotation file hg38
<p>Human genome annotation file hg38 from ensembl including snoRNA from snoDB database</p>
Curated GWAS summary statistics on European ancestry on 19 blood count traits and glycemic traits (hg38)
<p>Genome wide curated summary statistics on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>). GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p>
hg38 reference panel for RobusTAD
<p>hg38 reference panel for RobusTAD (5kb resolution).</p>
The hg38 genome sequences used in this project
<p>hg38_original.fasta.gz : The original hg38 contig sequences (contigs length < 500kb were excluded ).</p> <p>hg38_sim1.fasta.gz : The simulated heterozygous hg38 genome.</p> <p>hg38_sim2.fasta.gz : The simulated error-included hg38 genome.</p>
Indexed hg38 human genome from NCBI
<p>Raw data obtained from https://www.ncbi.nlm.nih.gov/assembly/88331</p> <p>Processed using</p> <p>- bwa mem for indexing</p>
TRIM28 modulates progesterone and estrogne signaling in endometrium [ATACseq-hg38]
GEO Series GSE205473. Homo sapiens. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.