Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “imputation”
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
MIMOSA: A resource consisting of improved methylome imputation models increases power to identify CpG site-phenotype associations
<p>MIMOSA DNA methylation prediction models, set up for MWAS. To run MWAS with this resource, see the tutorial here: <a href="https://github.com/ChongWuLab/MIMOSA">https://github.com/ChongWuLab/MIMOSA</a></p>
Data from: A real data-driven simulation strategy to select an imputation method for mixed-type trait data
Open the record for dataset details and reuse information.
Supplemental material for: Genome-wide association study and fine-mapping using imputed sequences to prioritize candidate genes for 30 complex traits in 50,309 Holstein bulls
Open the record for dataset details and reuse information.
A Flexible, Interpretable, and Accurate Approach for Imputing the Expression of Unmeasured Genes - Data
<p>This file contains the data that was used in the paper titled "A Flexible, Interpretable, and Accurate Approach for Imputing the Expression of Unmeasured Genes", which is submitted for review. The preprint is available at <a href="https://www.biorxiv.org/content/10.1101/2020.03.30.016675v1.abstract">https://www.biorxiv.org/content/10.1101/2020.03.30.016675v1.abstract</a></p>
Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies
Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.
Data from: A comparison of genomic selection models across time in interior spruce (Picea engelmannii × glauca) using unordered SNP imputation methods
Genomic selection (GS) potentially offers an unparalleled advantage over traditional pedigree-based selection (TS) methods by reducing the time commitment required to carry out a single cycle of tree improvement. This quality is particularly appealing to tree breeders, where lengthy improvement cycles are the norm. We explored the prospect of implementing GS for interior spruce (Picea engelmannii × glauca) utilizing a genotyped population of 769 trees belonging to 25 open-pollinated families. A series of repeated tree height measurements through ages 3–40 years permitted the testing of GS methods temporally. The genotyping-by-sequencing (GBS) platform was used for single nucleotide polymorphism (SNP) discovery in conjunction with three unordered imputation methods applied to a data set with 60% missing information. Further, three diverse GS models were evaluated based on predictive accuracy (PA), and their marker effects. Moderate levels of PA (0.31–0.55) were observed and were of sufficient capacity to deliver improved selection response over TS. Additionally, PA varied substantially through time accordingly with spatial competition among trees. As expected, temporal PA was well correlated with age-age genetic correlation (r=0.99), and decreased substantially with increasing difference in age between the training and validation populations (0.04–0.47). Moreover, our imputation comparisons indicate that k-nearest neighbor and singular value decomposition yielded a greater number of SNPs and gave higher predictive accuracies than imputing with the mean. Furthermore, the ridge regression (rrBLUP) and BayesCπ (BCπ) models both yielded equal, and better PA than the generalized ridge regression heteroscedastic effect model for the traits evaluated.
Datasets for MambaCpG: Accurate Imputation of Single-cell DNA Methylation Status Using Mamba
Open the record for dataset details and reuse information.
Code and Data from: An Imputation-Based Approach for Augmenting Sparse Epidemiological Signals
<p>This directory contains R code and required data to run the full data augmentation described in, "An Imputation-Based Approach for Augmenting Sparse Epidemiological Signals." This is the updated code corresponding to the updated medRxiv manuscript. It now includes ILINet data as a predictor in the imputation.</p> <p> </p> <p>"aug_pipeline.R" runs through all component steps and calls individual functions and data files within the directory. "plots_for_pipeline.R" uses data created during the aug_pipeline script to visualize individual steps in the augmentation process.</p>
Imputed multiplexed DNA FISH data from mouse cells
<p>Imputed multiplexed DNA FISH data from the following studies:</p><ol><li>Huang et al 2021 (DOI: 10.1038/s41588-021-00863-6): 5kb resolution chromatin tracing data of mouse embryonic stem cells</li><li>Takei et al 2021 (DOI: 10.1038/s41586-020-03126-2): 5kb and 1Mb resolution DNA seqFISH+ data of mouse embryonic stem cells</li><li>Takei et al 2021 (DOI: 10.1126/science.abj1966): 1Mb resolution DNA seqFISH+ data of mouse brain cells (excitatory neurons not included)</li></ol>
Test data for Imputation Workflow
<p>Test data for Imputation Workflow</p>
ALRA-imputed E10 and E11 mouse embryonic eye region dbit-seq data from Liu et al., Cell 183.6 (2020): 1665-1681
<p>E10 and E11 mouse embryonic eye region dbit-seq data demonstrated in Liu, Yang, et al. “High-spatial-resolution multi-omics sequencing via deterministic barcoding in tissue.” Cell 183.6 (2020): 1665-1681. The data are processed using the original pipeline and then imputed by ALRA.</p>
Data from: Testing hypotheses of marsupial brain size variation using phylogenetic multiple imputations and a Bayesian comparative framework
<p>Considerable controversy exists about which hypotheses and variables best explain mammalian brain size variation. We use a new, high-coverage dataset of marsupial brain and body sizes, and the first phylogenetically imputed full datasets of 16 predictor variables, to model the prevalent hypotheses explaining brain size evolution using phylogenetically corrected Bayesian generalised linear mixed-effects modelling. Despite this comprehensive analysis, litter size emerges as the only significant predictor. Marsupials differ from the more frequently studied placentals in displaying much lower diversity of reproductive traits, which are known to interact extensively with many behavioural and ecological predictors of brain size. Our results therefore suggest that studies of relative brain size evolution in placental mammals may require targeted co-analysis or adjustment of reproductive parameters like litter size, weaning age, or gestation length. This supports suggestions that significant associations between behavioural or ecological variables with relative brain size may be due to a confounding influence of the extensive reproductive diversity of placental mammals.</p>
Masked Conditional Diffusion Model with GNN for Spatial Transcriptomics Data Imputation
Open the record for dataset details and reuse information.
Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data
<p>This repository contains the real and synthetic datasets used in the paper "Benchmarking the Autoencoder Design for Imputing Single-Cell RNA Sequencing Data". The zip file includes three folders:</p> <p>1. overall imputation accuracy: the 12 real scRNA-seq datasets used in the evaluation of overall imputation accuracy.</p> <p>2. cell clustering: the 20 real scRNA-seq datasets with cell type labels used in the evaluation of cell clustering.</p> <p>3. DE gene: the 20 scRNA-seq syntehtic datasets with ground-truth DE genes used in the evaluation of DE gene analysis. These datasets are simulated by simulator scDesign and 20 real datasets. </p> <p> </p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p><p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Imputation of PaO2 From SaO2
ClinicalTrials.gov study NCT02598492. IPD Sharing: Not stated. Countries: 2. Publications: 18.
Hemoderivative Imputable Complications in Initial Uncomplicated Heart Surgery
ClinicalTrials.gov study NCT01457586. IPD Sharing: Not stated. Countries: 1. Publications: 8.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.