Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “imputation”
Data from: Testing hypotheses of marsupial brain size variation using phylogenetic multiple imputations and a Bayesian comparative framework
Open the record for dataset details and reuse information.
Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies
Open the record for dataset details and reuse information.
Data from: A comparison of genomic selection models across time in interior spruce (Picea engelmannii × glauca) using unordered SNP imputation methods
Open the record for dataset details and reuse information.
Supplementary Methods: Multiple imputation of CVRF trajectories in detail
Open the record for dataset details and reuse information.
Publicly available GWAS summary statistics, harmonized and imputed to GTEx v8' variant reference
<p># harmonized and imputed GWAS summary statistics</p> <p> </p> <p>* `harmonized_imputed_gwas.tar` contains 114 publicly available GWAS traits, harmonized and imputed to GTEx v8 reference</p> <p> </p> <p>* `gwas_metadata.txt` is a table with useful information about each trait, such as:</p> <p>- Tag: trait name (also in the file name)</p> <p>- PUBMED_Paper_Link: PUBMED or publication URL (if available)</p> <p>- Portal: URL to web portal from which data was downloaded</p> <p>- Consortium: GWAS Consortium authoring the data</p> <p>- Sample_Size: number of individuals covered in the study</p> <p>- Population: individuals'ancestry (EUR, EAS, etc)</p> <p>- abbreviation: short name used for figures</p> <p>- new_abbreviation: alternative name for additional figures</p> <p>- Deflation: whether imputed summary statistics exhibited deflation (i.e. association p-values are lower than expected by chance. The summary statistics imputation method is conservative, and in public GWAS with few observed variants (<2M), the distribution of p-values lags towards lower significance spectrums.</p> <p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication "Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits" (https://doi.org/10.1101/814350).</p> <p># Disclaimer</p> <p>The data is provided "as is", and the authors assume no responsibility for errors or omissions. <br> The User assumes the entire risk associated with its use of these data. <br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. <br> The User bears all responsibility in determining whether these data are fit for the User's intended use. </p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. <br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. <br> <br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia's requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>
Data from: Estimations of linkage disequilibrium, effective population size and ROH-based inbreeding coefficients in Spanish Churra sheep using imputed high-density SNP genotypes
In this study, the availability of the Ovine HD SNP BeadChip (HD-chip) and the development of an imputation strategy provided an opportunity to further investigate the extent of linkage disequilibrium (LD) at short distances in the genome of the Spanish Churra dairy sheep breed. A population of 1686 animals, including 16 rams and their half-sib daughters, previously genotyped for the 50K-chip, was imputed to the HD-chip density based on a reference population of 335 individuals. After assessing the imputation accuracy for beagle v4.0 (0.922) and fimpute v2.2 (0.921) using a cross-validation approach, the imputed HD-chip genotypes obtained with beagle were used to update the estimates of LD and effective population size for the studied population. The imputed genotypes were also used to assess the degree of homozygosity by calculating runs of homozygosity and to obtain genomic-based inbreeding coefficients. The updated LD estimations provided evidence that the extent of LD in Churra sheep is even shorter than that reported based on the 50K-chip and is one of the shortest extents compared with other sheep breeds. Through different comparisons we have also assessed the impact of imputation on LD and effective population size estimates. The inbreeding coefficient, considering the total length of the run of homozygosity, showed an average estimate (0.0404) lower than the critical level. Overall, the improved accuracy of the updated LD estimates suggests that the HD-chip, combined with an imputation strategy, offers a powerful tool that will increase the opportunities to identify genuine marker-phenotype associations and to successfully implement genomic selection in Churra sheep.
Data from: The population genomics of archaeological transition in west Iberia: investigation of ancient substructure using imputation and haplotype-based methods
We analyse new genomic data (0.05–2.95x) from 14 ancient individuals from Portugal distributed from the Middle Neolithic (4200–3500 BC) to the Middle Bronze Age (1740–1430 BC) and impute genomewide diploid genotypes in these together with published ancient Eurasians. While discontinuity is evident in the transition to agriculture across the region, sensitive haplotype-based analyses suggest a significant degree of local hunter-gatherer contribution to later Iberian Neolithic populations. A more subtle genetic influx is also apparent in the Bronze Age, detectable from analyses including haplotype sharing with both ancient and modern genomes, D-statistics and Y-chromosome lineages. However, the limited nature of this introgression contrasts with the major Steppe migration turnovers within third Millennium northern Europe and echoes the survival of non-Indo-European language in Iberia. Changes in genomic estimates of individual height across Europe are also associated with these major cultural transitions, and ancestral components continue to correlate with modern differences in stature.
Data from: Using multiple imputation to estimate missing data in meta-regression
1. There is a growing need for scientific synthesis in ecology and evolution. In many cases, meta-analytic techniques can be used to complement such synthesis. However, missing data is a serious problem for any synthetic efforts and can compromise the integrity of meta-analyses in these and other disciplines. Currently, the prevalence of missing data in meta-analytic datasets in ecology and the efficacy of different remedies for this problem have not been adequately quantified. 2. We generated meta-analytic datasets based on literature reviews of experimental and observational data and found that missing data were prevalent in meta-analytic ecological datasets. We then tested the performance of complete case removal (a widely used method when data are missing) and multiple imputation (an alternative method for data recovery) and assessed model bias, precision, and multi-model rankings under a variety of simulated conditions using published meta-regression datasets. 3. We found that complete case removal led to biased and imprecise coefficient estimates and yielded poorly specified models. In contrast, multiple imputation provided unbiased parameter estimates with only a small loss in precision. The performance of multiple imputation, however, was dependent on the type of data missing. It performed best when missing values were weighting variables, but performance was mixed when missing values were predictor variables. Multiple imputation performed poorly when imputing raw data which was then used to calculate effect size and the weighting variable. 4. We conclude that complete case removal should not be used in meta-regression, and that multiple imputation has the potential to be an indispensable tool for meta-regression in ecology and evolution. However, we recommend that users assess the performance of multiple imputation by simulating missing data on a subset of their data before implementing it to recover actual missing data.
Supplementary information to "ScRNA-IMM: Single-cell RNA-Seq Imputation method using Mean/Median Imputation"
Open the record for dataset details and reuse information.
Imputation Datasets
Open the record for dataset details and reuse information.
Supplementary data to BANMF-S: a blockwise accelerated non-negative matrix factorization framework with structural network constraints for single cell RNA-seq data imputation
Open the record for dataset details and reuse information.
Single-Cell ALRA-imputed Rat Lung Data
<p>ALRA-imputed Rat Lung Data from Raredon 2019 (10.1126/sciadv.aaw3851) </p>
R code and data for "Multiple imputation and direct estimation for qPCR data with non-detects"
<p>R code and data to reproduce figures and tables in the manuscript: Multiple imputation and direct estimation for qPCR data with non-detects.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Data from: Assessing among-lineage variability in phylogenetic imputation of functional trait datasets
Open the record for dataset details and reuse information.
Data from: The population genomics of archaeological transition in west Iberia: investigation of ancient substructure using imputation and haplotype-based methods
Open the record for dataset details and reuse information.
Data from: Using multiple imputation to estimate missing data in meta-regression
Open the record for dataset details and reuse information.
Data from: Estimations of linkage disequilibrium, effective population size and ROH-based inbreeding coefficients in Spanish Churra sheep using imputed high-density SNP genotypes
Open the record for dataset details and reuse information.
Data from: Improving the resolution of canine genome-wide association studies using genotype imputation: a study of two breeds
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.