Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “GRCh38”
GREEN-VARAN scores resources (FATHMM-XF GRCh38)
<p>Processed FATHMM-XF non-coding scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for FATHMM-XF v2.3 non-coding annotations.</p> <p>See: <a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use FATHMM-XF score annotations with GREEN-VARAN don't forget to cite also the original FATHMM-XF paper.</p>
GREEN-VARAN scores resources (EIGEN GRCh38)
<p>Processed EIGEN and EIGEN-PC scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for EIGEN v1.1 non-coding annotations, obtained by coordinates liftover.</p> <p>See: <a href="http://www.columbia.edu/~ii2135/eigen.html">http://www.columbia.edu/~ii2135/eigen.html</a></p> <p>If you use EIGEN score annotations with GREEN-VARAN don't forget to cite also the original EIGEN paper.</p>
Eigen scores for human genome assembly GRCh38 Part 4 (Chr1 - Chr2)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh38 Part 3 (Chr3 - Chr5)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
GenoNet scores for human genome assembly GRCh38
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
An updated map of GRCh38 linkage disequilibrium blocks based on European ancestry data
<p>A map of approximately independent linkage disequilibrium (LD) blocks has many uses in statistical genetics. Current publicly available LD block maps are based on sparse recombination maps and are only available for GRCh37 (hg19) and prior genome assemblies. We generated LD blocks in GRCh38 for European (EUR) ancestry populations using a recent recombination map based on more than 115,000 individuals. This new map consists of 1,361 independent LD blocks across the 22 autosomal chromosomes and can be accessed at https://github.com/jmacdon/LDblocks_GRCh38</p>
A comprehensive catalog of exact short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38
<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of exact repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,233,959 TR intervals, covering 0.783306% (24.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta). </p>
ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome
<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the GRCh38 (hg38) assembly of the human genome using the <a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing & Analysis Pipeline</a> (details in the documentation on GitHub).</p>
Mappability tracks for human assemblies (hg19 and GRCh38)
<p>They were created by using the GEM mapper aligner (Derrien et al., 2012) allowing up to two mismatches and considering sliding windows of 100-mer. They are exploited by the EXCAVATOR2 tool for reducing technical biases of Read Count measure in WES/TS experiments.</p>
STAR Reference Files for GRCm39 and GRCh38
<p>This is a <a href="https://github.com/alexdobin/STAR/tree/master">STAR</a> reference index file that uses a recent version (2.7.11b) of STAR to be compatible with STARsolo. The source of these reference files has been taken from the <a href="https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads">CellRanger pre-built references </a>of GRCm39 (2024-A) and GRCh38 (2024-A), respectively.</p> <p>Although the CellRanger pre-built references already contain STAR reference files, they are NOT compatible with STARsolo or the latest version of STAR. Those who need newer STAR references can use this repository. </p> <p>This repository is intended to facilitate the use of the NovaScope pipeline, but can also be used as a general purpose reference for STAR and STARsolo.</p>
A comprehensive catalog of approxiamte short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38
<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of approximate repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,656,159 TR intervals, covering 1.107653% (34.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta). </p>
SVAFotate core GRCh38 input BED including gnomADv4.1
<p>An updated input BED file for use with SVAFotate including the SV data from gnomAD v4.1. This core file is for use with GRCh38 as all included data was derived from GRCh38 alignments (no liftovers included).</p>
Aggregated frequencies of transcription initiations observed in FANTOM5 CAGE data on GRCh38, including alignments with low mapping qualities
<p><strong>Overview</strong></p> <p>Aligned reads of the FANTOM5 CAGE data have been used after filtering (ones with mapping quality less than 20 or percent identity less than 85% were discarded) for general purpose, resulting in the data set consisting of only the reads aligned with confidence. The filtering process made possible to interpret the data without ambiguity, however it also limited interpretation of paralogous or duplicated regions within the genome. Here all of the 5'-ends of the CAGE read alignments, including the ones with low mapping quality, were counted. The counts in the individual profiles were aggregated and summed up. </p> <p> </p> <p><strong>Special usage note</strong></p> <p>As noted above, this data derived from the alignments with low mapping qualities, as well as the ones with high mapping qualities. The result has to be examined very carefully: observations on the genome does not support transcription initiation with confidence, and even absence of such observation does not support silence of transcription with confidence. For example, file size on the forward strand is substantially larger than the one on the reverse strand, which is likely caused by an arbitrary preference of the alignment process. It does not mean transcription happens more frequently on the forward strand. Interpretation has to be made always in comparison with the standard data (BED files under http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v4/basic/ or bigWig files under http://fantom.gsc.riken.jp/5/datahub/hg38/reads/).</p> <p> </p> <p><strong>Data files</strong></p> <p>The resulting data files are formatted as bigWig (https://genome.ucsc.edu/FAQ/FAQformat.html#format6.1). '*.fwd.bw' and '*.rev.bw' represent forward and reverse strand on the genome, respectively. </p> <p> </p> <p><strong>Methods</strong></p> <p>The BAM files under http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v4/basic/ were subjected to 5'-end counting by bedtools v2.27.1 (https://github.com/arq5x/bedtools2), followed by conversion into bigWig with jksrc v357 (http://hgdownload.cse.ucsc.edu/admin/).</p> <p> </p>
GREEN-VARAN scores resources (CADD GRCh38)
<p>Processed CADD scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for CADD v.1.5.</p> <p>See: <a href="https://cadd.gs.washington.edu/">https://cadd.gs.washington.edu/</a></p> <p>If you use CADD score annotations with GREEN-VARAN don't forget to cite also the original CADD paper.</p>
GREEN-VARAN scores resources (FATHMM-MKL GRCh38)
<p>Processed FATHMM-MKL scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for FATHMM-MKL v2.3 non-coding annotations.</p> <p>See: <a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use FATHMM-MKL score annotations with GREEN-VARAN don't forget to cite also the original FATHMM-MKL paper.</p>
Notable genomic regions in T2T-CHM13 and GRCh38
<p>See 00README.md</p>
Ensembl TSS dataset for GRCh38
<p>We used the human genome reference sequence in its GRCh38.p13 version in order to have a reliable source of data in which to carry out our experiments. We chose this version because it is the most recent one available in Ensemble at the moment. However, the DNA sequence by itself is not enough, the specific TSS position of each transcript is needed. In this section, we explain the steps followed to generate the final dataset. These steps are: raw data gathering, positive instances processing, negative instances generation and data splitting by chromosomes.</p> <p>First, we need an interface in order to download the raw data, which is composed by every transcript sequence in the human genome. We used Ensembl release 104 (Howe et al., 2020) and its utility BioMart (Smedley et al., 2009), which allows us to get large amounts of data easily. It also enables us to select a wide variety of interesting fields, including the transcription start and end sites. After filtering instances that present null values in any relevant field, this combination of the sequence and its flanks will form our raw dataset. Once the sequences are available, we find the TSS position (given by Ensembl) and the 2 following bases to treat it as a codon. After that, 700 bases before this codon and 300 bases after it are concatenated, getting the final sequence of 1003 nucleotides that is going to be used in our models. These specific window values have been used in (Bhandari et al., 2021) and we have kept them as we find it interesting for comparison purposes. One of the most sensitive parts of this dataset is the generation of negative instances. We cannot get this kind of data in a straightforward manner, so we need to generate it synthetically. In order to get examples of negative instances, i.e. sequences that do not represent a transcript start site, we select random DNA positions inside the transcripts that do not correspond to a TSS. Once we have selected the specific position, we get 700 bases ahead and 300 bases after it as we did with the positive instances.</p> <p>Regarding the positive to negative ratio, in a similar problem, but studying TIS instead of TSS (Zhang135<br> et al., 2017), a ratio of 10 negative instances to each positive one was found optimal. Following this136<br> idea, we select 10 random positions from the transcript sequence of each positive codon and label them137<br> as negative instances. After this process, we end up with 1,122,113 instances: 102,488 positive and 1,019,625 negative sequences. In order to validate and test our models, we need to split this dataset into three parts: train, validation and test. We have decided to make this differentiation by chromosomes, as it is done in (Perez-Rodriguez et al., 2020). Thus, we use chromosome 16 as validation because it is a good example of a chromosome with average characteristics. Then we selected samples from chromosomes 1, 3, 13, 19 and 21 to be part of the test set and used the rest of them to train our models. Every step of this process can be replicated using the scripts available in https://github.com/JoseBarbero/EnsemblTSSPrediction.</p>
Revised transcript annotations for GRCh38 reference genome and Ensembl v87.
<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh38<br> Ensembl version: 87</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>
2nd version of modified GRCh38 reference
<p>This is the 2nd version of Modified GRCh38 reference that excludes decoys related to FANCD2, DUSP22 and GPRIN2 genes as these are the complex genes.</p>
Eigen scores for human genome assembly GRCh38 Part 1 (Chr12 - Chr22)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.