Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

33

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

33 results for “GRCh38”

Learn how ShareScore rates datasets ↗
zenodo48/100

GREEN-VARAN scores resources (FATHMM-XF GRCh38)

<p>Processed FATHMM-XF&nbsp;non-coding scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38&nbsp;version for FATHMM-XF v2.3 non-coding annotations.</p> <p>See:&nbsp;<a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use&nbsp;FATHMM-XF score&nbsp;annotations with&nbsp;GREEN-VARAN don&#39;t forget to cite also the original FATHMM-XF paper.</p>

opencc-by-4.0Aug 2020View details →
zenodo48/100

GREEN-VARAN scores resources (EIGEN GRCh38)

<p>Processed EIGEN and EIGEN-PC scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38&nbsp;version for EIGEN v1.1 non-coding annotations, obtained by coordinates liftover.</p> <p>See:&nbsp;<a href="http://www.columbia.edu/~ii2135/eigen.html">http://www.columbia.edu/~ii2135/eigen.html</a></p> <p>If you use&nbsp;EIGEN score&nbsp;annotations with&nbsp;GREEN-VARAN don&#39;t forget to cite also the original EIGEN paper.</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Eigen scores for human genome assembly GRCh38 Part 4 (Chr1 - Chr2)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Eigen scores for human genome assembly GRCh38 Part 3 (Chr3 - Chr5)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

GenoNet scores for human genome assembly GRCh38

<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files.&nbsp;</p> <p>Each row represents a genomic region with 131 columns. Please find the header line in &quot;genonet.header.txt&quot;.&nbsp;</p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 &nbsp; &nbsp;10000 &nbsp; &nbsp;10025 &nbsp; &nbsp;chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37&nbsp;<a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover&nbsp;<a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

An updated map of GRCh38 linkage disequilibrium blocks based on European ancestry data

<p>A map of approximately independent linkage disequilibrium (LD) blocks has many uses in statistical genetics. Current publicly available LD block maps are based on sparse recombination maps and are only available for GRCh37 (hg19) and prior genome assemblies. We generated LD blocks in GRCh38 for European (EUR) ancestry populations using a recent recombination map based on more than 115,000 individuals. This new map consists of 1,361 independent LD blocks across the 22 autosomal chromosomes and can be accessed at https://github.com/jmacdon/LDblocks_GRCh38</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

A comprehensive catalog of exact short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38

<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of exact repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,233,959 TR intervals, covering 0.783306% (24.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta).&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome

<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the&nbsp;GRCh38 (hg38) assembly of the human genome&nbsp;using&nbsp;the&nbsp;<a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing &amp; Analysis Pipeline</a>&nbsp;(details in the documentation on GitHub).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Mappability tracks for human assemblies (hg19 and GRCh38)

<p>They were created by using the GEM mapper aligner (Derrien et al., 2012) allowing up to two mismatches and considering sliding windows of 100-mer. They are exploited by the EXCAVATOR2 tool for reducing&nbsp;technical biases&nbsp;of Read Count measure in WES/TS experiments.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

STAR Reference Files for GRCm39 and GRCh38

<p>This is a <a href="https://github.com/alexdobin/STAR/tree/master">STAR</a> reference index file that uses a recent version (2.7.11b) of STAR to be compatible with STARsolo. The source of these reference files has been taken from the <a href="https://www.10xgenomics.com/support/software/cell-ranger/downloads#reference-downloads">CellRanger pre-built references </a>of GRCm39 (2024-A) and GRCh38 (2024-A), respectively.</p> <p>Although the CellRanger pre-built references already contain STAR reference files, they are NOT compatible with STARsolo or the latest version of STAR. Those who need newer STAR references can use this repository.&nbsp;</p> <p>This repository is intended to facilitate the use of the NovaScope pipeline, but can also be used as a general purpose reference for STAR and STARsolo.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

A comprehensive catalog of approxiamte short tandem repeat regions on autosomes and sex chromosomes of the human genome GRCh38

<p>To obtain a general TR catalog across the human genome, we identified genomic intervals with a stretch of approximate repetitions of a DNA motif ranging from 1-6bp on GRCh38 autosomes and sex chromosomes by using STRfinder (v1.0), and each STR region was annotated based on gencode.V38 (https://www.gencodegenes.org/human/release_38.html). To end up, we successfully found 1,656,159 TR intervals, covering 1.107653% (34.2 Mbp) of GRCh38 (https://console.cloud.google.com/storage/browser/_details/genomics-public-data/resources/broad/hg38/v0/Homo_sapiens_assembly38.fasta).&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

SVAFotate core GRCh38 input BED including gnomADv4.1

<p>An updated input BED file for use with SVAFotate including the SV data from gnomAD v4.1. This core file is for use with GRCh38 as all included data was derived from GRCh38 alignments (no liftovers included).</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Aggregated frequencies of transcription initiations observed in FANTOM5 CAGE data on GRCh38, including alignments with low mapping qualities

<p><strong>Overview</strong></p> <p>Aligned reads of the FANTOM5 CAGE data have been used after filtering (ones with&nbsp;mapping quality less than 20 or percent identity less than 85% were discarded) for general purpose, resulting in the data set consisting of only the reads aligned&nbsp;with confidence. The filtering process made possible to interpret the data without ambiguity, however it also limited interpretation of paralogous or duplicated regions within the genome. Here all of the 5&#39;-ends of the CAGE read alignments, including the ones with low mapping quality, were counted. The counts in the individual profiles were aggregated and summed up.&nbsp;</p> <p>&nbsp;</p> <p><strong>Special usage note</strong></p> <p>As noted above, this data derived from the alignments with low mapping qualities, as well as the ones with high mapping qualities. The result has to be examined very carefully: observations on the genome does not support transcription initiation with confidence, and even absence of such observation does not support silence of transcription with confidence. For example, file size&nbsp;on the forward strand is substantially larger than the one on the reverse strand, which is likely caused by an arbitrary preference of the alignment process. It does not mean transcription happens more frequently on the forward strand.&nbsp;Interpretation has to be made always in comparison with the standard data (BED files under http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v4/basic/ or bigWig files under http://fantom.gsc.riken.jp/5/datahub/hg38/reads/).</p> <p>&nbsp;</p> <p><strong>Data files</strong></p> <p>The resulting data files are formatted as bigWig (https://genome.ucsc.edu/FAQ/FAQformat.html#format6.1). &#39;*.fwd.bw&#39; and &#39;*.rev.bw&#39; represent forward and reverse strand on the genome, respectively.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>The BAM files under http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v4/basic/ were subjected to 5&#39;-end counting by bedtools v2.27.1 (https://github.com/arq5x/bedtools2), followed by conversion into bigWig with jksrc v357 (http://hgdownload.cse.ucsc.edu/admin/).</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

GREEN-VARAN scores resources (CADD GRCh38)

<p>Processed CADD&nbsp;scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for CADD v.1.5.</p> <p>See:&nbsp;<a href="https://cadd.gs.washington.edu/">https://cadd.gs.washington.edu/</a></p> <p>If you use&nbsp;CADD score&nbsp;annotations with&nbsp;GREEN-VARAN don&#39;t forget to cite also the original CADD&nbsp;paper.</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

GREEN-VARAN scores resources (FATHMM-MKL GRCh38)

<p>Processed FATHMM-MKL scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38&nbsp;version for FATHMM-MKL v2.3 non-coding annotations.</p> <p>See:&nbsp;<a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use&nbsp;FATHMM-MKL score&nbsp;annotations with&nbsp;GREEN-VARAN don&#39;t forget to cite also the original FATHMM-MKL&nbsp;paper.</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

Notable genomic regions in T2T-CHM13 and GRCh38

<p>See 00README.md</p>

opencc-zeroApr 2024View details →
zenodo36/100

Ensembl TSS dataset for GRCh38

<p>We used the human genome reference sequence in its GRCh38.p13 version in order to have a reliable source of data in which to carry out our experiments. We chose this version because it is the most recent one available in Ensemble at the moment. However, the DNA sequence by itself is not enough, the specific TSS position of each transcript is needed. In this section, we explain the steps followed to generate the final dataset. These steps are: raw data gathering, positive instances processing, negative instances generation and data splitting by chromosomes.</p> <p>First, we need an interface in order to download the raw data, which is composed by every transcript sequence in the human genome. We used Ensembl release 104 (Howe et al., 2020) and its utility BioMart (Smedley et al., 2009), which allows us to get large amounts of data easily. It also enables us to select a wide variety of interesting fields, including the transcription start and end sites. After filtering instances that present null values in any relevant field, this combination of the sequence and its flanks will form&nbsp;our raw dataset. Once the sequences are available, we find the TSS position (given by Ensembl) and the 2 following bases to treat it as a codon. After that, 700 bases before this codon and 300 bases after it are concatenated, getting the final sequence of 1003 nucleotides that is going to be used in our models. These specific window values have been used in (Bhandari et al., 2021) and we have kept them as we find it interesting for comparison purposes. One of the most sensitive parts of this dataset is the generation of negative instances. We cannot get&nbsp;this kind of data in a straightforward manner, so we need to generate it synthetically. In order to get examples of negative instances, i.e. sequences that do not represent a transcript start site, we select random DNA positions inside the transcripts that do not correspond to a TSS. Once we have selected the specific position, we get 700 bases ahead and 300 bases after it as we did with the positive instances.</p> <p>Regarding the positive to negative ratio, in a similar problem, but studying TIS instead of TSS (Zhang135<br> et al., 2017), a ratio of 10 negative instances to each positive one was found optimal. Following this136<br> idea, we select 10 random positions from the transcript sequence of each positive codon and label them137<br> as negative instances. After this process, we end up with 1,122,113 instances: 102,488 positive and 1,019,625 negative sequences. In order to validate and test our models, we need to split this dataset into three parts: train, validation and test. We have decided to make this differentiation by chromosomes, as it is done in (Perez-Rodriguez et al., 2020). Thus, we use chromosome 16 as validation because it is a good example of a chromosome with average characteristics. Then we selected samples from chromosomes 1, 3, 13, 19 and 21 to be part of the test set and used the rest of them to train our models. Every step of this process can be replicated using the scripts available in https://github.com/JoseBarbero/EnsemblTSSPrediction.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Revised transcript annotations for GRCh38 reference genome and Ensembl v87.

<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh38<br> Ensembl version: 87</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

2nd version of modified GRCh38 reference

<p>This is the 2nd version of Modified GRCh38 reference that excludes decoys related to FANCD2, DUSP22 and GPRIN2 genes as these are the complex genes.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Eigen scores for human genome assembly GRCh38 Part 1 (Chr12 - Chr22)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Dec 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record