Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

86

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

86 results for “hg38”

Learn how ShareScore rates datasets ↗
zenodo32/100

Major and minor allele genomes for hg38 using dbSNP151 with precomputed Bowtie and BWA indexes

<p>Minor and major allele genomes for the hg38 reference genome (GRCh38.p12) built using common SNPs from dbSNP release 151. Processing details are described in&nbsp;https://doi.org/10.1101/2022.04.21.488824.</p> <p>The following files are included:</p> <ul> <li>Fasta files for both the major and minor allele genomes</li> <li>Bowtie indexes for both the major and minor allele genomes</li> <li>BWA indexes for both the major and minor allele genomes</li> </ul> <p>The genomes are also available as&nbsp;Bioconductor&nbsp;BSgenome packages:</p> <ul> <li>https://bioconductor.org/packages/release/data/annotation/html/BSgenome.Hsapiens.UCSC.hg38.dbSNP151.major.html</li> <li>https://bioconductor.org/packages/release/data/annotation/html/BSgenome.Hsapiens.UCSC.hg38.dbSNP151.minor.html</li> </ul> <p>&nbsp;</p>

openmit-licenseJul 2022View details →
zenodo32/100

genomecomb additional files (cadd, minimap2) for reference data for Homo sapiens (hg38) version 0.108.0

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Archaic variants from Altai, Vindija, Chagyrskaya and Denisova (hg38)

<p>This is the variants for Altai, Vindija, Chagyrskaya and Denisova lifted over from hg19 to hg38.</p> <p>Liftover with CrossMap.py</p> <p>CrossMap.py vcf {chain} {vcffile} {refgenome} {outfile} --no-comp-alleles</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

hg38 reference and annotation files

<p>This repo contains reference and annotation files for hg38. We are following the [TOPMed pipeline](https://github.com/broadinstitute/gtex-pipeline/blob/master/TOPMed_RNAseq_pipeline.md). Reach out to Arushi Varshney at arushiv AT umich DOT edu if you have any questions.</p> <p>Files:</p> <p>1. bwa index = bwa.tar.gz</p> <p>2. star index = star.tar.gz</p> <p>3. ENCODE blacklist = blacklist.tar</p> <p>4. gencode v30 annotations = gencode.tar.gz</p> <p>5. containers with STAR (RNA) and BWA (ATAC) = containers.tar.gz</p> <p>Notes on these files:<br> ### hg38 fasta:<br> I downloaded the TOPMed fasta tar [Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.tar.gz](https://personal.broadinstitute.org/francois/topmed/Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.tar.gz) as use this as-is. The TOPMed GitHub describes that they obtained the Broad institute&#39;s GRCh38 reference, removed ALT, HLA and Decoy contigs, and added ERCC spike-in reference annotations. Refer to their [README](https://github.com/broadinstitute/gtex-pipeline/blob/master/TOPMed_RNAseq_pipeline.md) for more details. They don&#39;t mention PARs but we checked the reference files and both chrY PARs are hard masked - as [ENCODE](https://www.encodeproject.org/files/GRCh38_no_alt_analysis_set_GCA_000001405.15/) also recommends.<br> ### Gencode v30 gene annotations: gencode.tar.gz<br> &nbsp;I downloaded the file [gencode.v30.annotation.gtf.gz](https://www.gencodegenes.org/human/release_30.html) from the gencode website, and downloaded the file [ERCC92.genes.patched.gtf](https://personal.broadinstitute.org/francois/resources/). I then appended the ERCC patched gtf to the gencode annotation gtf<br> ```<br> gunzip gencode.v30.annotation.gtf.gz<br> cat gencode.v30.annotation.gtf&nbsp; ERCC92.genes.patched.gtf &gt; gencode.v30.annotation.ERCC92.gtf<br> ```<br> ### STAR index: star.tar.gz; container with star in containers.tar.gz<br> A STAR index is shared on the TOPMed GitHub, but it was generated for STAR version STAR_2.6.1d. Since I&#39;ve been using the version 2.7.3a, I followed their steps to generate the STAR reference again. I used the gencode gtf described above and generated the STAR index.</p> <p>```<br> STAR --runMode genomeGenerate&nbsp; --genomeDir STAR_genome_GRCh38_noALT_noHLA_noDecoy_ERCC_v30_test&nbsp; --genomeFastaFiles Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.fasta&nbsp; --sjdbGTFfile gencode.v30.annotation.ERCC92.gtf&nbsp; --sjdbOverhang 100 --runThreadN 10<br> ```<br> ### BWA index: bwa.tar.gz<br> I generated the BWA index using the fasta above<br> ```</p> <p>ln -s Homo_sapiens_assembly38_noALT_noHLA_noDecoy_ERCC.fasta hg38.fa<br> bwa index hg38.fa<br> ```<br> <br> ### ENCODE Blacklist: blacklist.tar<br> I used the blacklist [here](https://theparkerlab.med.umich.edu/data/arushiv/hg38_references_annots/blacklist/) that I obtained from this [Kundaje website](https://sites.google.com/site/anshulkundaje/projects/blacklists).</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

mcrpc_wgbs_hg38

<p>A HDF5-backed RangedSummarizedExperiment for WGBS Data (hg38 CpG sites)&nbsp;for 100 castration-resistant prostate cancer metastases from the paper&nbsp;&#39;Zhao, Shuang G., et al. &quot;The DNA methylation landscape of advanced prostate cancer.&quot;&nbsp;<em>Nature genetics</em>&nbsp;52.8 (2020): 778-789.&#39;.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
dryad32/100

Geographic allele frequency variation in the 1000 Genomes hg38 NYGC dataset

Open the record for dataset details and reuse information.

publicDec 2020View details →
zenodo28/100

QDNAseq.hg38: QDNAseq bin annotation for the human genome build hg38

<p><strong>QDNAseq</strong> bin annotations of size 1, 5, 10, 15, 30, 50, 100, 500, and 1000 kbp for the human genome build hg38.</p>

openother-openNov 2020View details →
zenodo28/100

DOHH2 hg38 H3K4me1 ChIP-seq Dataset filtered for Unique Multiread Mappability

<p>DOHH2 hg38 H3K4me1 ChIP-seq Dataset filtered for Unique Multiread Mappability</p>

opencc-zeroAug 2016View details →
zenodo28/100

Database files (hg19 and hg38) and test files for RTpred

<p>Database files (hg19 and hg38) and test files for RTpred</p>

openmit-licenseNov 2021View details →
zenodo28/100

hg38 syntenic ages

<p>### METHOD ###</p> <p>Sequence ages were estimated from hg38 100-way multiZ vertebrate&nbsp;sequence alignments from the UCSC genome browser. Briefly, sequence ages were estimated as the branch length between humans (hg38) and the oldest most recent common ancestor (MRCA) from the 100-way neutral tree.&nbsp;</p> <p>&nbsp;</p> <p>MRCA to taxon mappings are available in the&nbsp;hg38_syn_taxon.bed file.&nbsp;</p> <p>&nbsp;</p> <p># Hg38&nbsp;100-way vertebrate tree syntenic block age files.&nbsp;</p> <p>Information for hg38&nbsp;100-way vertebrate multiple sequence alignment can be found here -&nbsp;http://hgdownload.cse.ucsc.edu/goldenPath/hg38/multiz100way/README.txt</p> <p># Scripts generating these files</p> <p>(1) Creating syntenic block .bed files</p> <p>https://github.com/slifong08/enh_ages/blob/54084ff12e521c20cc5288d6cb3590fac93c11ed/age_arch/manuscript_scripts/get_spec_count_msa-hg19.py</p> <p>(2) Assigning most recent common ancestor (MRCA) and patristic distances to syntenic blocks</p> <p>https://github.com/slifong08/enh_ages/blob/54084ff12e521c20cc5288d6cb3590fac93c11ed/age_arch/manuscript_scripts/get_synteny_age_hg19.py</p>

opencc-by-4.0Dec 2021View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg38) version 0.98.7

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2018View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg38) version 0.11.0

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2015View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg38) version 0.108.0

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo28/100

Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)

<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs &quot;170630_M00528_0292_000000000-B9JY8&quot; (aka &quot;NC_LIMMS&quot;) and &quot;180221_M00528_0334_000000000-B6PJM&quot; (aka &quot;NC_LIMMS2&quot;). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics&nbsp;2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under&nbsp;the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p>&nbsp;</p> <p><em><strong>&quot;170630_M00528_0292_000000000-B9JY8&quot; (&quot;NC_LIMMS&quot;) :</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>1&nbsp;&nbsp; iPSC_control_rep1&nbsp;&nbsp; 4&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>2&nbsp;&nbsp; iPSC_control_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>3&nbsp;&nbsp; iPSC_control_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>4&nbsp;&nbsp; S3P1_OK_rep1&nbsp;&nbsp; 36&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>5&nbsp;&nbsp; S3P1_OK_rep2&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>6&nbsp;&nbsp; S3P1_OK_rep3&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>7&nbsp;&nbsp; S4P1_OK_rep1&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>8&nbsp;&nbsp; S4P1_OK_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>9&nbsp;&nbsp; S4P1_OK_rep3&nbsp;&nbsp; 9&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>10&nbsp;&nbsp; S4P2_OK_rep1&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>11&nbsp;&nbsp; S4P2_OK_rep2&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>12&nbsp;&nbsp; S4P2_OK_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>13&nbsp;&nbsp; S1P1_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>14&nbsp;&nbsp; S1P1_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>15&nbsp;&nbsp; S1P1_rep3&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>16&nbsp;&nbsp; S3P1_FAILED_rep1&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>17&nbsp;&nbsp; S3P1_FAILED_rep2&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>18&nbsp;&nbsp; S3P1_FAILED_rep3&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>19&nbsp;&nbsp; S4P1_FAILED_rep1&nbsp;&nbsp; 35&nbsp;&nbsp; CACTCT&nbsp;&nbsp; NNNNNNNN</p> <p>20&nbsp;&nbsp; S4P1_FAILED_rep2&nbsp;&nbsp; 47&nbsp;&nbsp; CTGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>21&nbsp;&nbsp; S4P1_FAILED_rep3&nbsp;&nbsp; 59&nbsp;&nbsp; GAGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>22&nbsp;&nbsp; S4P2_FAILED_rep1&nbsp;&nbsp; 71&nbsp;&nbsp; GCTGCA&nbsp;&nbsp; NNNNNNNN</p> <p>23&nbsp;&nbsp; S4P2_FAILED_rep2&nbsp;&nbsp; 83&nbsp;&nbsp; TATAGC&nbsp;&nbsp; NNNNNNNN</p> <p>24&nbsp;&nbsp; S4P2_FAILED_rep3&nbsp;&nbsp; 95&nbsp;&nbsp; TCGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p> <p><em><strong>&quot;180221_M00528_0334_000000000-B6PJM&quot; (&quot;NC_LIMMS2&quot;):</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>25&nbsp;&nbsp; PETRI_rep1&nbsp;&nbsp; 04&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>26&nbsp;&nbsp; PETRI_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>27&nbsp;&nbsp; PETRI_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>28&nbsp;&nbsp; BIOCHIP_E_rep1&nbsp;&nbsp; 6&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>29&nbsp;&nbsp; BIOCHIP_M_rep1&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>30&nbsp;&nbsp; BIOCHIP_S_rep1&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>31&nbsp;&nbsp; BIOCHIP_E_rep2&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>32&nbsp;&nbsp; BIOCHIP_M_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>33&nbsp;&nbsp; BIOCHIP_S_rep2&nbsp;&nbsp; 09&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>34&nbsp;&nbsp; BIOCHIP_E_rep3&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>35&nbsp;&nbsp; BIOCHIP_M_rep3&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>36&nbsp;&nbsp; BIOCHIP_S_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>37&nbsp;&nbsp; HEPATOCYTES_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>38&nbsp;&nbsp; HEPATOCYTES_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>39&nbsp;&nbsp; iPSC_control_rep1-2&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>40&nbsp;&nbsp; BIOCHIP_E_rep2-2&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>41&nbsp;&nbsp; BIOCHIP_M_rep1-2&nbsp;&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>42&nbsp;&nbsp; BIOCHIP_S_rep2-2&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p>

openOct 2017View details →
zenodo28/100

Human genome annotation file hg38

<p>Human genome annotation file hg38 from ensembl including snoRNA from snoDB database</p>

opencc-by-4.0May 2023View details →
zenodo28/100

Curated GWAS summary statistics on European ancestry on 19 blood count traits and glycemic traits (hg38)

<p>Genome wide curated summary statistics&nbsp;on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>).&nbsp;GWAS of glycemic traits come from the&nbsp;<a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a>&nbsp;study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p>

opencc-by-4.0Jun 2023View details →
zenodo28/100

hg38 reference panel for RobusTAD

<p>hg38 reference panel for RobusTAD&nbsp;(5kb resolution).</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

The hg38 genome sequences used in this project

<p>hg38_original.fasta.gz : The&nbsp;original&nbsp;hg38 contig sequences (contigs length &lt; 500kb were excluded&nbsp;).</p> <p>hg38_sim1.fasta.gz : The&nbsp;&nbsp;simulated heterozygous hg38 genome.</p> <p>hg38_sim2.fasta.gz : The&nbsp;&nbsp;simulated error-included&nbsp;hg38 genome.</p>

opencc-by-4.0Sep 2023View details →
zenodo24/100

Indexed hg38 human genome from NCBI

<p>Raw data obtained from https://www.ncbi.nlm.nih.gov/assembly/88331</p> <p>Processed using</p> <p>- bwa mem for indexing</p>

opencc-by-4.0Apr 2020View details →
geo20/100

TRIM28 modulates progesterone and estrogne signaling in endometrium [ATACseq-hg38]

GEO Series GSE205473. Homo sapiens. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record