Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

67

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

67 results for “hg19”

Learn how ShareScore rates datasets ↗
zenodo36/100

GROVER tokenized Human Genome hg19

<p>Data of the Human Genome Hg19, tokenised with byte-pair tokenisation of 600 cycles. Required for the DNA language model GROVER. More information can be found at&nbsp;https://www.biorxiv.org/content/10.1101/2023.07.19.549677v1.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Genome and Transcriptome references based on hg19 from UCSC, 2015

<p>rsem.transcripts.nant2015.fa.gz - bgzipped FASTA reference of transcriptomes</p><p>genome.nant2015.fa.gz - bgzipped FASTA human genome reference, with several viral sequences added.</p><p>refseq.txt.gz - Exact sequence accessions and mapping coordinates for a RefSeq transcriptome based off the UCSC genome browser for hg19.</p><p>Coordinates are BED-style, with one row per transcript, and 1+ transcript per gene.</p><p>Column annotation</p><p>1. RefSeq Accession</p><p>2. Chromosome</p><p>3. Strand</p><p>4. thinStart (gene boundary, including UTR)</p><p>5. thinEnd (gene boundary, including UTR)</p><p>6. thickStart (CDS boundary)</p><p>7. thinStart (CDS boundary)</p><p>8. number of exons</p><p>9. comma separated exon starts</p><p>10. comma separate exon ends</p><p>11. common gene name</p><p>12. refseq gene id</p><p>13. 0 if non-primary transcript, 1 if primary transcript</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

pjhop/DNAmCrosshyb: hg38 and hg19 bisulfite-converted genomes (R .rds files)

<p>Bisulfite-converted genomes as used in the&nbsp;<em>DNAmCrosshyb</em> R package (<a href="http://github.com/pjhop/DNAmCrosshyb">github.com/pjhop/DNAmCrosshyb</a>). Data is saved per chromosome in the &lsquo;DNAString&rsquo; format as implemented in the Biostrings BioConductor package. The data is saved in the R rds file format and can be read using the &lsquo;readRDS()&rsquo; function. Scripts used to generate these data can be found at&nbsp;<a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R</a>&nbsp;and&nbsp;<a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R</a>&nbsp;. See&nbsp;<a href="https://github.com/pjhop/DNAmCrosshyb">https://github.com/pjhop/DNAmCrosshyb</a>&nbsp;for examples on how to map Illumina 450k/EPIC array probes to the bisulfite-converted genomes.</p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

Resource bundle for somatic-conda (hg19)

<p>Resource bundle for the use of somatic-conda workflow (including hg19 only).</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Archaic variants from Altai, Vindija, Chagyrskaya and Denisova (hg19)

<p>This is the variants for Altai, Vindija, Chagyrskaya and Denisova mapped to hg19.</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

TEST_SpliceAI_rocksdb_hg19

<p>TEST&nbsp;SpliceAI RocksDB for chromosomes 1 and 2&nbsp;of hg19&nbsp;as used in the AbSplice publication:&nbsp;<a href="https://www.nature.com/articles/s41588-023-01373-3">https://www.nature.com/articles/s41588-023-01373-3</a></p> <p>Actual precomputed SpliceAI scores for all SNVs and indels up to 4 nucleotides are stored in related entries on Zenodo. To use those&nbsp;databases for fast computation of SpliceAI predictions see:&nbsp;<a href="https://github.com/gagneurlab/spliceai_rocksdb">https://github.com/gagneurlab/spliceai_rocksdb</a></p> <p>This is also implemented in the AbSplice package:&nbsp;<a href="https://github.com/gagneurlab/absplice">https://github.com/gagneurlab/absplice</a></p> <p>This dataset includes SpliceAI scores. The scores are free for academic and not-for-profit use; other use requires a commercial license from Illumina, Inc., see the GitHub repository of SpliceAI:&nbsp;<a href="https://github.com/Illumina/SpliceAI/tree/master">https://github.com/Illumina/SpliceAI/tree/master</a></p>

opencc-by-4.0May 2023View details →
zenodo32/100

Functional matrices N, TEs vs promoters, weighted with L = 2.5e5kb, hg19

<p>Refers to the manuscript: &quot;Statistical learning quantifies transposable element-mediated cis-regulation&quot;, by Pulver et al., 2023.</p> <p>&nbsp;</p> <p>TE subfamilies are split between so-called &quot;functional&quot; and complementary &quot;non-functional&quot; fractions. Functional integrants were defined as such based on overlapping epigenomics data (e.g. transcription factor binding by ChIP-seq, open chromatin by ATAC-seq).</p> <p>&nbsp;</p> <p>The &quot;functional_weighted_mappability&quot; only contain cis-regulatory weights for selected subfamilies, split between low and high mappability (median split over averages per integrants). They should be incorporated into N_functional_weighted matrices for the appropriate experiments, and not be used on their own.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Matrices N, TEs vs promoters, weighted with L in [1e3kB, 1e10kB], hg19

<p>Related to &quot;Statistical learning quantifies transposable element-mediated cis-regulation&quot;, Pulver et al. 2023</p> <p>Regulatory susceptibility matrices N, with rows as hg19 protein coding genes and columns as TE subfamilies. For any given gene, regulatory TEs are strictly out of any promoter and any exon belonging to that gene. The contribution of each single TE is weighted by a gaussian kernel centered on the closest promoter of that gene, with varying bandwiths spanning the range L = 1e3kB to 1e10kB.</p> <p>The &quot;TAD_restricted&quot; N matrices were built by bounding the distances until which TEs were considered as putative cis-regulatory elements for protein-coding genes to TAD boundaries. Pairs of genes - TEs that do not overlap TADs are weighted irrespective of TAD boundaries.</p> <p>The &quot;weighted_mappability&quot; N matrices only contain selected subfamilies and should be incorporated into N_weighted computed with L = 2.5e5kb before usage. Low vs high mappability fractions of subfamilies were separated using a median split on per-integrant average mappability scores.</p>

opencc-by-4.0Jul 2023View details →
dryad32/100

Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)

<p>Analyses of genetic variation in many taxa have established that neutral genetic diversity is shaped by natural selection at linked sites. Whether the mode of selection is primarily the fixation of strongly beneficial alleles (selective sweeps) or purifying selection on deleterious mutations (background selection) remains unknown, however. We address this question in humans by fitting a model of the joint effects of selective sweeps and background selection to autosomal polymorphism data from the 1000 Genomes Project. After controlling for variation in mutation rates along the genome, a model of background selection alone explains ~60% of the variance in diversity levels at the megabase scale. Adding the effects of selective sweeps driven by adaptive substitutions to the model does not improve the fit, and when both modes of selection are considered jointly, selective sweeps are estimated to have had little or no effect on linked neutral diversity. The regions under purifying selection are best predicted by phylogenetic conservation, with ~80% of the deleterious mutations affecting neutral diversity occurring in non-exonic regions. Thus, background selection is the dominant mode of linked selection in humans, with marked effects on diversity levels throughout autosomes.</p>

opencc-zeroAug 2023View details →
dryad32/100

Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)

Open the record for dataset details and reuse information.

publicAug 2023View details →
zenodo28/100

Database files (hg19 and hg38) and test files for RTpred

<p>Database files (hg19 and hg38) and test files for RTpred</p>

openmit-licenseNov 2021View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg19) version 0.11.0

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2017View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg19) version 0.9

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg19) version 0.8.5

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2015View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg19) version 0.98.7

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2018View details →
zenodo28/100

genomecomb reference data for Homo sapiens (hg19) version 0.98.7 cad annotation data

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2018View details →
zenodo28/100

Masked version of hG19 by Brian Bushnell

<p>Masked version of hG19 by Brian Bushnell (http://seqanswers.com/forums/showthread.php?t=42552)</p>

opencc-by-4.0Mar 2018View details →
zenodo28/100

bed12 tables for hg19, dm6 and ce11

<p>tables for tandem exon duplication analysis</p>

opencc-by-4.0Sep 2021View details →
zenodo28/100

Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)

<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs &quot;170630_M00528_0292_000000000-B9JY8&quot; (aka &quot;NC_LIMMS&quot;) and &quot;180221_M00528_0334_000000000-B6PJM&quot; (aka &quot;NC_LIMMS2&quot;). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics&nbsp;2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under&nbsp;the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p>&nbsp;</p> <p><em><strong>&quot;170630_M00528_0292_000000000-B9JY8&quot; (&quot;NC_LIMMS&quot;) :</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>1&nbsp;&nbsp; iPSC_control_rep1&nbsp;&nbsp; 4&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>2&nbsp;&nbsp; iPSC_control_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>3&nbsp;&nbsp; iPSC_control_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>4&nbsp;&nbsp; S3P1_OK_rep1&nbsp;&nbsp; 36&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>5&nbsp;&nbsp; S3P1_OK_rep2&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>6&nbsp;&nbsp; S3P1_OK_rep3&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>7&nbsp;&nbsp; S4P1_OK_rep1&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>8&nbsp;&nbsp; S4P1_OK_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>9&nbsp;&nbsp; S4P1_OK_rep3&nbsp;&nbsp; 9&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>10&nbsp;&nbsp; S4P2_OK_rep1&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>11&nbsp;&nbsp; S4P2_OK_rep2&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>12&nbsp;&nbsp; S4P2_OK_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>13&nbsp;&nbsp; S1P1_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>14&nbsp;&nbsp; S1P1_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>15&nbsp;&nbsp; S1P1_rep3&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>16&nbsp;&nbsp; S3P1_FAILED_rep1&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>17&nbsp;&nbsp; S3P1_FAILED_rep2&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>18&nbsp;&nbsp; S3P1_FAILED_rep3&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>19&nbsp;&nbsp; S4P1_FAILED_rep1&nbsp;&nbsp; 35&nbsp;&nbsp; CACTCT&nbsp;&nbsp; NNNNNNNN</p> <p>20&nbsp;&nbsp; S4P1_FAILED_rep2&nbsp;&nbsp; 47&nbsp;&nbsp; CTGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>21&nbsp;&nbsp; S4P1_FAILED_rep3&nbsp;&nbsp; 59&nbsp;&nbsp; GAGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>22&nbsp;&nbsp; S4P2_FAILED_rep1&nbsp;&nbsp; 71&nbsp;&nbsp; GCTGCA&nbsp;&nbsp; NNNNNNNN</p> <p>23&nbsp;&nbsp; S4P2_FAILED_rep2&nbsp;&nbsp; 83&nbsp;&nbsp; TATAGC&nbsp;&nbsp; NNNNNNNN</p> <p>24&nbsp;&nbsp; S4P2_FAILED_rep3&nbsp;&nbsp; 95&nbsp;&nbsp; TCGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p> <p><em><strong>&quot;180221_M00528_0334_000000000-B6PJM&quot; (&quot;NC_LIMMS2&quot;):</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>25&nbsp;&nbsp; PETRI_rep1&nbsp;&nbsp; 04&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>26&nbsp;&nbsp; PETRI_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>27&nbsp;&nbsp; PETRI_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>28&nbsp;&nbsp; BIOCHIP_E_rep1&nbsp;&nbsp; 6&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>29&nbsp;&nbsp; BIOCHIP_M_rep1&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>30&nbsp;&nbsp; BIOCHIP_S_rep1&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>31&nbsp;&nbsp; BIOCHIP_E_rep2&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>32&nbsp;&nbsp; BIOCHIP_M_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>33&nbsp;&nbsp; BIOCHIP_S_rep2&nbsp;&nbsp; 09&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>34&nbsp;&nbsp; BIOCHIP_E_rep3&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>35&nbsp;&nbsp; BIOCHIP_M_rep3&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>36&nbsp;&nbsp; BIOCHIP_S_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>37&nbsp;&nbsp; HEPATOCYTES_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>38&nbsp;&nbsp; HEPATOCYTES_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>39&nbsp;&nbsp; iPSC_control_rep1-2&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>40&nbsp;&nbsp; BIOCHIP_E_rep2-2&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>41&nbsp;&nbsp; BIOCHIP_M_rep1-2&nbsp;&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>42&nbsp;&nbsp; BIOCHIP_S_rep2-2&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p>

openOct 2017View details →
geo24/100

Targeting KDM4 for treating PAX3-FOXO1-driven alveolar rhabdomyosarcoma [RNAseq_LHCN_hg19]

GEO Series GSE201224. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record