Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “imputation”
Merging and imputation of flow cytometry data: a critical assessment
<p>This dataset contains the data necessary to reproduce the analyses by Mocking <em>et al</em>.</p> <p>Scripts for reproducing our analyses are available at:</p> <p>https://github.com/AUMC-HEMA/imputation-manuscript</p>
STGMVA: clustering, imputation, and integration for spatial resolved transcriptomics using spatiotemporal gaussian mixture variational autoencoder
<p> In this study, we present STGMVA, a comprehensive analysis toolkit employs a spatiotemporal gaussian mixture variational autoencoder to tackle these tasks effectively. STGMVA consists of two stages: pretraining the gene expression and spatial location using a gaussian mixture model, and learning the embedding vectors through a variational graph autoencoder. Results demonstrate STGMVA surpasses state-of-the-art approaches on various spatial transcriptomics datasets, exhibiting superior performance across different scales and resolutions. Notably, STGMVA achieves the highest clustering accuracy in human brain, mouse hippocampus, and mouse olfactory bulb tissues. Furthermore, STGMVA enhances and denoises gene expression patterns for gene imputation task. Additionally, STGMVA has the capability to correct batch effects and achieve joint analysis when integrating multiple tissue slices.</p>
Reliable imputation of spatial transcriptome with uncertainty estimation and spatial regularization
<p>Imputation of missing features in spatial transcriptomics is urgently demanded due to technology limitations, while most existing computational methods suffer from moderate accuracy and cannot estimate the reliability of the imputation. <br> To fill the research gaps, we introduce a computational model, TransImp, that imputes the missing feature modality in spatial transcriptomics by mapping it from single-cell reference. Uniquely, we derived a set of attributes that can accurately predict imputation uncertainty, hence enabling us to select reliably imputed genes. Also, we introduced a spatial auto-correlation metric as a regularization to avoid overestimating spatial patterns. Multiple datasets from various platforms have demonstrated that our approach significantly improves the reliability of downstream analyses in detecting spatial variable genes and interacting ligand-receptor pairs. Therefore, TransImp offers a way towards a reliable spatial analysis of missing features for both matched and unseen modalities, e.g., nascent RNAs.</p>
Data of: Imputation-free reconstructions of three-dimensional chromosome architectures in human diploid single-cells using allele-specified contacts
Open the record for dataset details and reuse information.
GWAS summary statistics imputation support data and integration with PrediXcan MASHR
<p># GWAS summary statistics imputation, integration with PrediXcan MASHR-M</p> <p> </p> <p>The file `sample_data.tar` contains all necessary files to perform imputation of GWAS summary statistics to the GTEx v8 QTL data set.</p> <p>It includes 1000 Genomes individuals' genotypes as reference panel.</p> <p>The `.tar` archive, upon uncompression, contains the following folder structure:</p> <p>```</p> <p>data<br> |-- coordinate_map<br> |-- gwas<br> |-- liftover<br> |-- models<br> | |-- eqtl<br> | | `-- mashr<br> | `-- sqtl<br> | `-- mashr<br> |-- reference_panel_1000G<br> `-- ucsc</p> <p>```</p> <p> </p> <p>`data/eur_ld.bed.gz` contains definitions of approximately independent LD-regions in hg38 (Berisa-Pickrell regions, lifted over)</p> <p>`data/gtex_v8_eur_filtered_maf0.01_monoallelic_variants.txt.gz` is a snp annotation file, listing all GTEx v8 variants with MAF>0.01 in europeans.</p> <p>`data/coordinate_map` contains precomputed mapping tables that MetaXcan tools can use to convert GWAS' genomic coordinates in GWAS between genome assemblies.</p> <p>`data/gwas` contains a sample GWAS file for the purposes of a tutorial (data obtained from Nikpay et al (Nat Gen 2016) https://www.ncbi.nlm.nih.gov/pubmed/26343387</p> <p>`data/liftover` contains Liftover chains to map coordinates between human genome assemblies (used by full harmonization tools)</p> <p>`data/models` contains PrediXcan MASHR-M models, and cross-tissue S-MultiXcan LD compilation, from eQTL and sQTL.</p> <p>`data/reference_panel_1000G` contains 1000G hg38 genotypes, in parquet format, to be used by imputation tools.</p> <p>`data/ucsc` contains genomic coordinates of rsids in hg17, hg18 and hg19. You can use these to add chromosome and start position information to a GWAS based on its rsids. (column `end` is not used)</p> <p> </p>
Imputed Multiple Sequence Alignment used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'
<p>Multiple Sequence Alignment of imputed SARS-CoV-2 sequences used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'</p>
Illumina HD genotypes for 3,092 cattle from Burkina Faso, Ghana, Nigeria and Tanzania for: "Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle"
<p>Raw HD data for Riggio et al. 2022: Assessment of genotyping array performance for genome-wide association studies and imputation in African cattle</p> <p>This repository contains the raw Illumina HD genotypes (i.e., 777,962 SNPs) mapped to the bovine UMD3.1 genome assembly for 3,092 animals from four African countries (namely Burkina Faso, Ghana, Nigeria and Tanzania). </p>
A high coverage Mesolithic aurochs genome and effective leveraging of ancient cattle genomes using whole genome imputation. -- VCF file
<p>This is the open-access VCF file that was created in the article: "<strong>A high coverage Mesolithic aurochs genome and effective leveraging of ancient cattle genomes using whole genome imputation."</strong></p> <p><strong>Information about the filtering steps can be found in the method section.</strong></p> <p>Extra information on the sample IDs can be found in the Supplementary tables.</p>
COVID-19 cases in U.S. hospitals with imputed locations and dates
<p>This RDS file contains imputed locations and dates for COVID-19 cases detected in U.S. hospitals. The locations are uniformly sampled within known counties (FIPS), additional covariates of which are included. This file is used for analyses contributing to the paper "Scaling the spatiotemporal Hawkes process to one million COVID-19 cases in U.S. hospitals".</p>
Cross-species and tissue imputation of species-level DNA methylation samples
<p>Imputed dataset of DNA methyation samples representing the predicted mean methylation of a species and tissue type. </p>
Dummy TOPMed Imputed Dataset
<p>This dataset comprises of bgzipped VCF files representative of output from the TOPMed imputation server (TOPMed-r3 imputation panel) October 2024. Each VCF contains 10,000 dummy individuals and "imputed" variants related to autosomal site lists found in the r3 panel. The data in the INFO field for each variant does not represent genotypes provided. All variants have MAF=0.5 with genotypes distributed under HWE. To assist in optimizing compression ratio, the first 2,500 individuals are homozygous for the reference allele, the next 5,000 are heterozygous, and the following 2,500 individuals are homozygous for the ALT allele.</p> <p> </p>
Annotation of a human retina organoid before and after imputation
<p>Two Seurat-objects which include the unimputed and DCA-imputed retina organoid data sets, published by Kim <em>et al. </em>, stored in RDS-format.</p>
Imputed data from for predicting skeletal stature using ancient DNA
<p>These are the three imputed genotype datasets from the following publication. Please consult the paper for details of the imputation approach, metadata for the samples, and the original data sources. These data are freely available but you should cite the original sources, as well as our paper. </p> <p>Predicting skeletal stature using ancient DNA; Cox S, Moots HM, Stock JT, Shbat A, Bitarello B, Nicklisch N, Alt K, Haak W, Rosenstock E, Ruff CB, Mathieson I; American Journal of Biological Anthropology, Jan 2022. <a href="https://doi.org/10.1002/ajpa.24426">https://doi.org/10.1002/ajpa.24426</a></p> <p> </p> <p> </p>
Data from: A real data-driven simulation strategy to select an imputation method for mixed-type trait data
<p>Missing observations in trait datasets pose an obstacle for analyses in myriad biological disciplines. Considering the mixed results of imputation, the wide variety of available methods, and the varied structure of real trait datasets, a framework for selecting a suitable imputation method is advantageous. We invoked a real data-driven simulation strategy to select an imputation method for a given mixed-type (categorical, count, continuous) target dataset. Candidate methods included mean/mode imputation, k-nearest neighbour, random forests, and multivariate imputation by chained equations (MICE). Using a trait dataset of squamates (lizards and amphisbaenians; order: Squamata) as a target dataset, a complete-case dataset consisting of species with nearly completed information was formed for the imputation method selection. Missing data were induced by removing values from this dataset under different missingness mechanisms: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). For each method, combinations with and without phylogenetic information from single gene (nuclear and mitochondrial) or multigene trees were used to impute the missing values for five numerical and two categorical traits. The performances of the methods were evaluated under each missing mechanism by determining the mean squared error and proportion falsely classified rates for numerical and categorical traits, respectively. A random forest method supplemented with a nuclear-derived phylogeny resulted in the lowest error rates for the majority of traits, and this method was used to impute missing values in the original dataset. Data with imputed values better reflected the characteristics and distributions of the original data compared to complete-case data. However, caution should be taken when imputing trait data as phylogeny did not always improve performance for every trait and in every scenario. Ultimately, these results support the use of a real data-driven simulation strategy for selecting a suitable imputation method for a given mixed-type trait dataset.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.