Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22
datasets available to search
ShareScore release 0.9.0
Dataset results
22 results for “GRCh37”
GREEN-VARAN scores resources (CADD GRCh37)
<p>Processed CADD scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for CADD v.1.4.</p> <p>See: <a href="https://cadd.gs.washington.edu/">https://cadd.gs.washington.edu/</a></p> <p>If you use CADD score annotations with GREEN-VARAN don't forget to cite also the original CADD paper.</p>
GREEN-VARAN scores resources (DANN GRCh37)
<p>Processed DANN scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for DANN.</p> <p>See: <a href="https://academic.oup.com/bioinformatics/article/31/5/761/2748191">https://academic.oup.com/bioinformatics/article/31/5/761/2748191</a></p> <p>If you use DANN score annotations with GREEN-VARAN don't forget to cite also the original DANN paper.</p>
GenoNet scores for human genome assembly GRCh37
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
GREEN-VARAN scores resources (EIGEN GRCh37)
<p>Processed EIGEN and EIGEN-PC scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for EIGEN v1.1 non-coding annotations.</p> <p>See: <a href="http://www.columbia.edu/~ii2135/eigen.html">http://www.columbia.edu/~ii2135/eigen.html</a></p> <p>If you use EIGEN score annotations with GREEN-VARAN don't forget to cite also the original EIGEN paper.</p>
GREEN-VARAN scores resources (FATHMM-XF GRCh37)
<p>Processed FATHMM-XF non-coding scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for FATHMM-XF v2.3 non-coding annotations.</p> <p>See: <a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use FATHMM-XF score annotations with GREEN-VARAN don't forget to cite also the original FATHMM-XF paper.</p>
GREEN-VARAN scores resources (FATHMM-MKL GRCh37)
<p>Processed FATHMM-MKL scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for FATHMM-MKL v2.3 non-coding annotations.</p> <p>See: <a href="http://fathmm.biocompute.org.uk/">http://fathmm.biocompute.org.uk/</a></p> <p>If you use FATHMM-MKL score annotations with GREEN-VARAN don't forget to cite also the original FATHMM-MKL paper.</p>
GREEN-VARAN scores resources (FIRE GRCh37)
<p>Processed FIRE scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for FIRE.</p> <p>See: <a href="https://sites.google.com/site/fireregulatoryvariation/">https://sites.google.com/site/fireregulatoryvariation/</a></p> <p>If you use FIRE score annotations with GREEN-VARAN don't forget to cite also the original FIRE paper.</p>
Eigen scores for human genome assembly GRCh37 Part 1 (Chr1 - Chr3)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh37 Part 4 (Chr17 - Chr22)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Clinvar database GRCh37
Open the record for dataset details and reuse information.
MACIE scores for human genome assembly GRCh37 Part 1 (Chr1 - Chr3)
<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not damaging protein functional and evolutionarily conserved” (MACIE01); “damaging protein functional and not evolutionarily conserved” (MACIE10); “not damaging protein functional and not evolutionarily conserved” (MACIE00); “both damaging protein functional and evolutionarily conserved” (MACIE11). MACIE_protein is the estimated posterior probability of “damaging protein functional”, which is the sum of MACIE10 and MACIE11; MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “damaging protein functional” or “evolutionarily conserved”, which is the sum of MACIE01, MACIE10, and MACIE11.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not evolutionarily conserved and regulatory functional” (MACIE01); “evolutionarily conserved and not regulatory functional” (MACIE10); “not evolutionarily conserved and not regulatory functional” (MACIE00); “both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of “regulatory functional”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “evolutionarily conserved” or “regulatory functional”, which is the sum of MACIE01, MACIE10, and MACIE11.</p>
MACIE scores for human genome assembly GRCh37 Part 4 (Chr14 - Chr22)
<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not evolutionarily conserved and regulatory functional” (MACIE01); “evolutionarily conserved and not regulatory functional” (MACIE10); “not evolutionarily conserved and not regulatory functional” (MACIE00); “both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of “regulatory functional”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “evolutionarily conserved” or “regulatory functional”, which is the sum of MACIE01, MACIE10, and MACIE11.</p>
MACIE scores for human genome assembly GRCh37 Part 3 (Chr8 - Chr13)
<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not evolutionarily conserved and regulatory functional” (MACIE01); “evolutionarily conserved and not regulatory functional” (MACIE10); “not evolutionarily conserved and not regulatory functional” (MACIE00); “both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of “regulatory functional”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “evolutionarily conserved” or “regulatory functional”, which is the sum of MACIE01, MACIE10, and MACIE11.</p>
Revised transcript annotations for GRCh37 (hg19) reference genome and Ensembl v90.
<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh37<br> Ensembl version: 90</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>
European ancestry: 72 traits spanning multiple clinical domains (GRCh37)
<div> </div> <div> <p><span><span>This collection of 72 traits </span><span>contains</span> <span>a wide variety of traits including </span><span>metabolic traits (</span><span>e.g.</span> <span>lipid level</span><span>s</span><span>, glycemic traits</span><span>…</span><span>), </span><span>immune system</span><span> diseases (</span><span>e.g.</span> <span>inflammatory bowel </span><span>disease</span><span>,</span><span> celiac disease</span><span>)</span><span>, </span><span>cardiovascular outcomes and more. </span><span>This collection allows </span><span>experimenting </span><span>with combination of traits from </span><span>different clinical domains and search</span><span>ing</span><span> for highly pleiotropic variants. </span> <span>More details </span><span>on</span><span> GWAS summary statistics and their curation are available </span></span><a href="https://www.biorxiv.org/content/10.1101/2023.10.27.564319v1" target="_blank" rel="noreferrer noopener"><span><span>here</span></span></a><span><span>.</span></span><span> </span></p> </div>
FUN-LDA scores for human genome assembly GRCh37
<p>FUN-LDA is based on a Latent Dirichlet Allocation (LDA) model for predicting functional effects of non-coding genetic variants in a cell type and tissue-specific way by integrating diverse epigenetic annotations for specific cell types and tissues from large-scale genomics projects such as ENCODE and Roadmap Epigenomics. Using this unsupervised approach, we predict tissue-specific functional effects for every position in the human genome for 127 tissues and cell types in ENCODE and Roadmap Epigenomics. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Format</strong></p> <p>The FUN-LDA scores are stored in the UCSC Genome Browser bigWig Track Format. </p> <p>To extract FUN-LDA scores, the bigWigAverageOverBed utility is required. It can be downloaded from the Genome Browser website at <a href="https://hgdownload.cse.ucsc.edu/admin/exe">https://hgdownload.cse.ucsc.edu/admin/exe</a>.</p> <p>User should prepare a UCSC Genome Browser bed file, a tab-separated four-column file. The first column is the chromosome; the second is the zero-based coordinate of the position of interest; the third is that zero-based coordinate plus one; and the fourth is a unique identifier for the position. The command below is an example of extracting FUN-LDA scores in tissue E007 for the positions defined in a bed file, "input.bed".</p> <pre><code class="language-bash">bigWigAverageOverBed E007.valley9.c89.bigwig input.bed output.tab</code></pre> <p>It produces a six-column file, output.tab. The first column includes the unique position id defined in the input bed file. The last column, average over the covered bases, is the score for this position.</p> <p><strong>Reference</strong></p> <p>Daniel Backenroth, Zihuai He, Krzysztof Kiryluk, Valentina Boeva, Lynn Pethukova, Ekta Khurana, Angela Christiano, Joseph Buxbaum, Iuliana Ionita-Laza. FUN-LDA: A latent Dirichlet allocation model for predicting tissue-specific functional effects of noncoding variation: Methods and applications. American Journal of Human Genetics, 2018.<br> </p> <p> </p>
GREEN-VARAN scores resources (GRCh37)
<p>Processed non-coding prediction scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for the following scores:</p> <ul> <li>ReMM v0.3.1 (<a href="https://charite.github.io/software-remm-score.html">https://charite.github.io/software-remm-score.html</a>)</li> <li>NCBoost v.1 (<a href="https://github.com/RausellLab/NCBoost">https://github.com/RausellLab/NCBoost</a>)</li> <li>ExPECTO (<a href="https://hb.flatironinstitute.org/expecto/">https://hb.flatironinstitute.org/expecto/</a>)</li> <li>LinSight (<a href="https://github.com/CshlSiepelLab/LINSIGHT">https://github.com/CshlSiepelLab/LINSIGHT</a>)</li> <li>GWAVA v1.0 (<a href="https://www.sanger.ac.uk/sanger/StatGen_Gwava">https://www.sanger.ac.uk/sanger/StatGen_Gwava</a>)</li> </ul> <p>If you use any of these score annotations with GREEN-VARAN please cite also the corresponding paper.</p>
Eigen scores for human genome assembly GRCh37 Part 3 (Chr9 - Chr16)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh37 Part 2 (Chr4 - Chr8)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
MACIE scores for human genome assembly GRCh37 Part 2 (Chr4 - Chr7)
<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of “not evolutionarily conserved and regulatory functional” (MACIE01); “evolutionarily conserved and not regulatory functional” (MACIE10); “not evolutionarily conserved and not regulatory functional” (MACIE00); “both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of “evolutionarily conserved”, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of “regulatory functional”, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of “evolutionarily conserved” or “regulatory functional”, which is the sum of MACIE01, MACIE10, and MACIE11.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.