Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “hg19”
GROVER tokenized Human Genome hg19
<p>Data of the Human Genome Hg19, tokenised with byte-pair tokenisation of 600 cycles. Required for the DNA language model GROVER. More information can be found at https://www.biorxiv.org/content/10.1101/2023.07.19.549677v1.</p>
Genome and Transcriptome references based on hg19 from UCSC, 2015
<p>rsem.transcripts.nant2015.fa.gz - bgzipped FASTA reference of transcriptomes</p><p>genome.nant2015.fa.gz - bgzipped FASTA human genome reference, with several viral sequences added.</p><p>refseq.txt.gz - Exact sequence accessions and mapping coordinates for a RefSeq transcriptome based off the UCSC genome browser for hg19.</p><p>Coordinates are BED-style, with one row per transcript, and 1+ transcript per gene.</p><p>Column annotation</p><p>1. RefSeq Accession</p><p>2. Chromosome</p><p>3. Strand</p><p>4. thinStart (gene boundary, including UTR)</p><p>5. thinEnd (gene boundary, including UTR)</p><p>6. thickStart (CDS boundary)</p><p>7. thinStart (CDS boundary)</p><p>8. number of exons</p><p>9. comma separated exon starts</p><p>10. comma separate exon ends</p><p>11. common gene name</p><p>12. refseq gene id</p><p>13. 0 if non-primary transcript, 1 if primary transcript</p>
pjhop/DNAmCrosshyb: hg38 and hg19 bisulfite-converted genomes (R .rds files)
<p>Bisulfite-converted genomes as used in the <em>DNAmCrosshyb</em> R package (<a href="http://github.com/pjhop/DNAmCrosshyb">github.com/pjhop/DNAmCrosshyb</a>). Data is saved per chromosome in the ‘DNAString’ format as implemented in the Biostrings BioConductor package. The data is saved in the R rds file format and can be read using the ‘readRDS()’ function. Scripts used to generate these data can be found at <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R</a> and <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R</a> . See <a href="https://github.com/pjhop/DNAmCrosshyb">https://github.com/pjhop/DNAmCrosshyb</a> for examples on how to map Illumina 450k/EPIC array probes to the bisulfite-converted genomes.</p>
Resource bundle for somatic-conda (hg19)
<p>Resource bundle for the use of somatic-conda workflow (including hg19 only).</p>
Archaic variants from Altai, Vindija, Chagyrskaya and Denisova (hg19)
<p>This is the variants for Altai, Vindija, Chagyrskaya and Denisova mapped to hg19.</p>
TEST_SpliceAI_rocksdb_hg19
<p>TEST SpliceAI RocksDB for chromosomes 1 and 2 of hg19 as used in the AbSplice publication: <a href="https://www.nature.com/articles/s41588-023-01373-3">https://www.nature.com/articles/s41588-023-01373-3</a></p> <p>Actual precomputed SpliceAI scores for all SNVs and indels up to 4 nucleotides are stored in related entries on Zenodo. To use those databases for fast computation of SpliceAI predictions see: <a href="https://github.com/gagneurlab/spliceai_rocksdb">https://github.com/gagneurlab/spliceai_rocksdb</a></p> <p>This is also implemented in the AbSplice package: <a href="https://github.com/gagneurlab/absplice">https://github.com/gagneurlab/absplice</a></p> <p>This dataset includes SpliceAI scores. The scores are free for academic and not-for-profit use; other use requires a commercial license from Illumina, Inc., see the GitHub repository of SpliceAI: <a href="https://github.com/Illumina/SpliceAI/tree/master">https://github.com/Illumina/SpliceAI/tree/master</a></p>
Functional matrices N, TEs vs promoters, weighted with L = 2.5e5kb, hg19
<p>Refers to the manuscript: "Statistical learning quantifies transposable element-mediated cis-regulation", by Pulver et al., 2023.</p> <p> </p> <p>TE subfamilies are split between so-called "functional" and complementary "non-functional" fractions. Functional integrants were defined as such based on overlapping epigenomics data (e.g. transcription factor binding by ChIP-seq, open chromatin by ATAC-seq).</p> <p> </p> <p>The "functional_weighted_mappability" only contain cis-regulatory weights for selected subfamilies, split between low and high mappability (median split over averages per integrants). They should be incorporated into N_functional_weighted matrices for the appropriate experiments, and not be used on their own.</p>
Matrices N, TEs vs promoters, weighted with L in [1e3kB, 1e10kB], hg19
<p>Related to "Statistical learning quantifies transposable element-mediated cis-regulation", Pulver et al. 2023</p> <p>Regulatory susceptibility matrices N, with rows as hg19 protein coding genes and columns as TE subfamilies. For any given gene, regulatory TEs are strictly out of any promoter and any exon belonging to that gene. The contribution of each single TE is weighted by a gaussian kernel centered on the closest promoter of that gene, with varying bandwiths spanning the range L = 1e3kB to 1e10kB.</p> <p>The "TAD_restricted" N matrices were built by bounding the distances until which TEs were considered as putative cis-regulatory elements for protein-coding genes to TAD boundaries. Pairs of genes - TEs that do not overlap TADs are weighted irrespective of TAD boundaries.</p> <p>The "weighted_mappability" N matrices only contain selected subfamilies and should be incorporated into N_weighted computed with L = 2.5e5kb before usage. Low vs high mappability fractions of subfamilies were separated using a median split on per-integrant average mappability scores.</p>
Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)
<p>Analyses of genetic variation in many taxa have established that neutral genetic diversity is shaped by natural selection at linked sites. Whether the mode of selection is primarily the fixation of strongly beneficial alleles (selective sweeps) or purifying selection on deleterious mutations (background selection) remains unknown, however. We address this question in humans by fitting a model of the joint effects of selective sweeps and background selection to autosomal polymorphism data from the 1000 Genomes Project. After controlling for variation in mutation rates along the genome, a model of background selection alone explains ~60% of the variance in diversity levels at the megabase scale. Adding the effects of selective sweeps driven by adaptive substitutions to the model does not improve the fit, and when both modes of selection are considered jointly, selective sweeps are estimated to have had little or no effect on linked neutral diversity. The regions under purifying selection are best predicted by phylogenetic conservation, with ~80% of the deleterious mutations affecting neutral diversity occurring in non-exonic regions. Thus, background selection is the dominant mode of linked selection in humans, with marked effects on diversity levels throughout autosomes.</p>
Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)
Open the record for dataset details and reuse information.
Database files (hg19 and hg38) and test files for RTpred
<p>Database files (hg19 and hg38) and test files for RTpred</p>
genomecomb reference data for Homo sapiens (hg19) version 0.11.0
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg19) version 0.9
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg19) version 0.8.5
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg19) version 0.98.7
Open the record for dataset details and reuse information.
genomecomb reference data for Homo sapiens (hg19) version 0.98.7 cad annotation data
Open the record for dataset details and reuse information.
Masked version of hG19 by Brian Bushnell
<p>Masked version of hG19 by Brian Bushnell (http://seqanswers.com/forums/showthread.php?t=42552)</p>
bed12 tables for hg19, dm6 and ce11
<p>tables for tandem exon duplication analysis</p>
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)
<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs "170630_M00528_0292_000000000-B9JY8" (aka "NC_LIMMS") and "180221_M00528_0334_000000000-B6PJM" (aka "NC_LIMMS2"). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p> </p> <p><em><strong>"170630_M00528_0292_000000000-B9JY8" ("NC_LIMMS") :</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>1 iPSC_control_rep1 4 ACAGAT NNNNNNNN</p> <p>2 iPSC_control_rep2 24 ATCGTG NNNNNNNN</p> <p>3 iPSC_control_rep3 31 CACGAT NNNNNNNN</p> <p>4 S3P1_OK_rep1 36 CACTGA NNNNNNNN</p> <p>5 S3P1_OK_rep2 46 CTGACG NNNNNNNN</p> <p>6 S3P1_OK_rep3 63 GAGTGA NNNNNNNN</p> <p>7 S4P1_OK_rep1 79 GTATAC NNNNNNNN</p> <p>8 S4P1_OK_rep2 92 TCGAGC NNNNNNNN</p> <p>9 S4P1_OK_rep3 9 ACATGA NNNNNNNN</p> <p>10 S4P2_OK_rep1 21 ATCATA NNNNNNNN</p> <p>11 S4P2_OK_rep2 33 CACGTG NNNNNNNN</p> <p>12 S4P2_OK_rep3 45 CGATGA NNNNNNNN</p> <p>13 S1P1_rep1 57 GAGATA NNNNNNNN</p> <p>14 S1P1_rep2 69 GCTCTC NNNNNNNN</p> <p>15 S1P1_rep3 81 GTATGA NNNNNNNN</p> <p>16 S3P1_FAILED_rep1 93 TCGATA NNNNNNNN</p> <p>17 S3P1_FAILED_rep2 11 AGTAGC NNNNNNNN</p> <p>18 S3P1_FAILED_rep3 23 ATCGCA NNNNNNNN</p> <p>19 S4P1_FAILED_rep1 35 CACTCT NNNNNNNN</p> <p>20 S4P1_FAILED_rep2 47 CTGAGC NNNNNNNN</p> <p>21 S4P1_FAILED_rep3 59 GAGCGT NNNNNNNN</p> <p>22 S4P2_FAILED_rep1 71 GCTGCA NNNNNNNN</p> <p>23 S4P2_FAILED_rep2 83 TATAGC NNNNNNNN</p> <p>24 S4P2_FAILED_rep3 95 TCGCGT NNNNNNNN</p> <p> </p> <p><em><strong>"180221_M00528_0334_000000000-B6PJM" ("NC_LIMMS2"):</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>25 PETRI_rep1 04 ACAGAT NNNNNNNN</p> <p>26 PETRI_rep2 24 ATCGTG NNNNNNNN</p> <p>27 PETRI_rep3 31 CACGAT NNNNNNNN</p> <p>28 BIOCHIP_E_rep1 6 CACTGA NNNNNNNN</p> <p>29 BIOCHIP_M_rep1 46 CTGACG NNNNNNNN</p> <p>30 BIOCHIP_S_rep1 63 GAGTGA NNNNNNNN</p> <p>31 BIOCHIP_E_rep2 79 GTATAC NNNNNNNN</p> <p>32 BIOCHIP_M_rep2 92 TCGAGC NNNNNNNN</p> <p>33 BIOCHIP_S_rep2 09 ACATGA NNNNNNNN</p> <p>34 BIOCHIP_E_rep3 21 ATCATA NNNNNNNN</p> <p>35 BIOCHIP_M_rep3 33 CACGTG NNNNNNNN</p> <p>36 BIOCHIP_S_rep3 45 CGATGA NNNNNNNN</p> <p>37 HEPATOCYTES_rep1 57 GAGATA NNNNNNNN</p> <p>38 HEPATOCYTES_rep2 69 GCTCTC NNNNNNNN</p> <p>39 iPSC_control_rep1-2 81 GTATGA NNNNNNNN</p> <p>40 BIOCHIP_E_rep2-2 93 TCGATA NNNNNNNN</p> <p>41 BIOCHIP_M_rep1-2 11 AGTAGC NNNNNNNN</p> <p>42 BIOCHIP_S_rep2-2 23 ATCGCA NNNNNNNN</p> <p> </p>
Targeting KDM4 for treating PAX3-FOXO1-driven alveolar rhabdomyosarcoma [RNAseq_LHCN_hg19]
GEO Series GSE201224. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.