Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.9.0
Dataset results
33 results for “GRCh38”
Eigen scores for human genome assembly GRCh38 Part 2 (Chr6 - Chr11)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
GREEN-VARAN scores resources (DANN GRCh38)
<p>Processed DANN scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for DANN.</p> <p>See: <a href="https://academic.oup.com/bioinformatics/article/31/5/761/2748191">https://academic.oup.com/bioinformatics/article/31/5/761/2748191</a></p> <p>If you use DANN score annotations with GREEN-VARAN don't forget to cite also the original DANN paper.</p>
GREEN-VARAN scores resources (FIRE GRCh38)
<p>Processed FIRE scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for FIRE.</p> <p>See: <a href="https://sites.google.com/site/fireregulatoryvariation/">https://sites.google.com/site/fireregulatoryvariation/</a></p> <p>If you use FIRE score annotations with GREEN-VARAN don't forget to cite also the original FIRE paper.</p>
GREEN-VARAN scores resources (GRCh38)
<p>Processed non-coding prediction scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for the following scores. When not available from the original source, the GRCh38 coordinates were obtained by liftover.</p> <ul> <li>ReMM v0.3.1 (<a href="https://charite.github.io/software-remm-score.html">https://charite.github.io/software-remm-score.html</a>)</li> <li>NCBoost v.1 (<a href="https://github.com/RausellLab/NCBoost">https://github.com/RausellLab/NCBoost</a>)</li> <li>ExPECTO (<a href="https://hb.flatironinstitute.org/expecto/">https://hb.flatironinstitute.org/expecto/</a>)</li> <li>LinSight (<a href="https://github.com/CshlSiepelLab/LINSIGHT">https://github.com/CshlSiepelLab/LINSIGHT</a>)</li> <li>GWAVA v1.0 (<a href="https://www.sanger.ac.uk/sanger/StatGen_Gwava">https://www.sanger.ac.uk/sanger/StatGen_Gwava</a>)</li> </ul> <p>If you use any of these score annotations with GREEN-VARAN please cite also the corresponding paper.</p>
GRCH38 chr 1,2,3,10
<p>A subsample of GRCH38</p>
Human genome fasta file from ensembl (GRCh38 v109)
<p>Human genome fasta file from ensembl (GRCh38 v109), <span>corresponds to GenBank Assembly ID </span><span>GCA_000001405.28</span></p>
GRCh38 (gencode, release 29) indices for snakePipes 2.2.0 and later versions
<p>A human (GRCh38 from Gencode) tarball containing all of the indices needed for snakePipes.</p>
dbSNP v155 for GRCh37 and GRCh38
<p>Chromosome, position, RSID, reference allele and alternative allele was downloaded from:</p> <p>https://bioconductor.org/packages/release/data/annotation/html/SNPlocs.Hsapiens.dbSNP155.GRCh38.html https://bioconductor.org/packages/release/data/annotation/html/SNPlocs.Hsapiens.dbSNP155.GRCh37.html</p> <p>A description of filters applied are descripbed [here](http://arvidharder.com/tidyGWAS/articles/transforming_dbsnp_to_parquet.html).</p> <p>The locations, alleles and rsIDs were converted a parquet files, partitioned by chromosome. </p>
Leucegene: AML sequencing (GRCh38 reference)
GEO Series GSE232130. Homo sapiens. 691 samples. Type: Expression profiling by high throughput sequencing.
GRCh38 modified reference
<p>This is the modified GRCh38 reference generated by fixing the falsely duplicated regions and falsely collapsed regions. </p>
Molecular Transducers of Human Skeletal Muscle Remodeling under Different Loading States [CDF: GC_M_HTA_3utr_Grch38,binary.cdf]
GEO Series GSE155959. Homo sapiens. 230 samples. Type: Expression profiling by array.
Molecular Transducers of Human Skeletal Muscle Remodeling under Different Loading States [CDF: GC_M_HTA_5utr_Grch38,binary.cdf]
GEO Series GSE155934. Homo sapiens. 230 samples. Type: Expression profiling by array.
Leucegene: AML sequencing (part 1-6, GRCh38 reference)
GEO Series GSE232129. Homo sapiens. 452 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.