Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
88
datasets available to search
ShareScore release 0.9.0
Dataset results
88 results for “imputation”
Imputed Forest Composition Map for New England Screened by Species Range Boundaries 2001-2006
Initializing forest landscape models (FLMs) to simulate changes in tree species composition requires accurate fine-scale forest attribute information mapped contiguously over large areas. Nearest-neighbor imputation maps have high potential for use as the initial condition within FLMs, but the tendency for field plots to be imputed over large geographical distances results in species frequently mapped outside of their home ranges, which is problematic. We developed an approach for evaluating and selecting field plots for imputation based on their similarity in feature-space, their species composition, and their geographical distance between source and imputation to produce a map that is appropriate for initializing an FLM. We applied this approach to map 13m ha of forest throughout the six New England states (Rhode Island, Connecticut, Massachusetts, New Hampshire, Vermont, and Maine). The map itself is a .img raster file of FIA plot CN numbers. To access FIA data from this map, one has to link the mapcodes in this map to FIA data supplied by USDA FIA database (https://apps.fs.usda.gov/fia/datamart/datamart.html). Due to plot confidentiality and integrity concerns, pixels containing FIA plots were always assigned to some other plot than the actual one found there.
CPTAC TMT protein quantifications imputed with Lupine
<p>TMT proteomics data collected from clinical patient samples as part of the Clinical Tumor Atlas Consortium (CPTAC) project were imputed with Lupine, a deep matrix factorization-based proteomics imputation method. All quantifications have been summarized at the protein level. </p>
Datasets with and without deliberate head movements for detection and imputation of dropout in diffusion MRI
Open the record for dataset details and reuse information.
Imputation panel for low-pass whole genome sequencing (GLIMPSE2 format)
<p>This dataset includes autosomal genotypes from the 1000 Genomes +HGDP project (<a href="https://doi.org/10.1101/2023.01.23.525248" target="_blank" rel="noopener">10.1101/2023.01.23.525248 </a>) as well as X chromosome genotypes from the NY Genome Center (as of yet, a comparable dataset that includes HGDP is not available for the X; see 10.1016/j.cell.2022.08.004). The genotypes were down-sampled so as to be appropriate for low-pass imputation; uncertain phase calls were removed (any PP tags), and individuals deemed to be outliers or relatives (based on autosomal data, as per the first citation) were also removed. Similarly, singleton polymorphisms were also excluded. Hemizygous genotypes on the X were converted into (quasi) diploid genotypes.</p> <p>These data were then converted into a binary imputation panel format using glimpse v2 (https://odelaneau.github.io/GLIMPSE/; using the static binaries provided). The "chunk" size was doubled from the defaults (which considers a minimum number of snps, genetic length and physical length) so as to be more performant.</p>
Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper
<p>Pre-built leave-out-out cross-validation imputation reference panel datasets - LmTag paper (complement for a tutorial at https://github.com/datngu/LmTag)</p> <p>This repo includes chromosome 10 reference panel data constructed for 3 populations:</p> <p>- EAS</p> <p>- EUR</p> <p>- SAS</p>
Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022.
<p>Time series data of COVID-19 cases (rT-PCR-confirmed), hospitalisations (laboratory-confirmed), and hospital-associated deaths (laboratory confirmed) in South Africa, by imputed dates of symptom onset, from the start of the pandemic in March 2020 through April 2022. These data were used to estimate the time-varying reproduction number (R) in South Africa, as described in https://www.medrxiv.org/content/10.1101/2022.07.22.22277932v1.full.</p>
Multi-Temporal Cloud Gap Imputation With HLS Data Across CONUS
<p>This release contains the version 1.0 of the dataset which was used in <a href="https://arxiv.org/abs/2404.19609">Seeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model</a> and is included as one of the tasks in the <a href="https://madewithclay.org/challenge">AI for Earth Challenge 2024</a>.</p>
Processed Datasets - Imputation in Well Log Data: A Benchmark
<p>Imputation of well log data is a common task in the field. However a quick review of the literature reveals a lack of padronization when evaluating methods for the problem. The goal of the benchmark is to introduce a standard evaluation protocol to any imputation method for well log data. </p> <p>In the proposed benchmark, three public datasets are used:</p> <ul> <li><strong>Geolink:</strong> The Geolink Dataset is another public dataset of wells in the Norwegian offshore. The data is provided by the company of the same name, <a href="https://www.geolink-s2.com/" target="_blank" rel="noopener">GEOLINK</a> and follows the NOLD 2.0 license. <br>This dataset contains a total of 223 wells. It also has lithology labels for the wells with a total of 36 lithology classes. [<a href="https://drive.google.com/drive/folders/1EgDN57LDuvlZAwr5-eHWB5CTJ7K9HpDP" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Taranaki Basin:</strong> The Taranaki Basin Dataset is a curated set of wells and a convenient option for experimentation especially due to it is ease of accessibility and use.<br>This collection, under the CDLA-Sharing-1.0 license, contains well logs extracted from the <a href="https://geodata.nzpam.govt.nz/" target="_blank" rel="noopener">New Zealand Petroleum & Minerals Online Exploration Database</a> and <a href="http://pet.gns.cri.nz/" target="_blank" rel="noopener">Petlab</a>.<br>There are a total of 407 wells, of which 289 are onshore and 118 are offshore exploration and production wells. [<a href="https://developer.ibm.com/exchanges/data/all/taranaki-basin-curated-well-logs/" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Teapot Dome:</strong> The Teapot Dome dataset is provided by the Rocky Mountain Oilfield Testing Center (RMOTC) and the US Department of Energy.<br>It contains different types of data related to the Teapot Dome oil field, such as 2D and 3D seismic data, well logs, and GIS data. The data is licensed under the Creative Commons 4.0 license. <br>In total, the dataset has 1,179 wells with available logs. The number of available logs varies across wells. There are only 91 wells with the gamma ray, bulk density, and neutron porosity logs, while only three wells have the complete basic suite. [<a href="http://s3.amazonaws.com/open.source.geoscience/open_data/teapot/rmotc.tar" target="_blank" rel="noopener">direct download</a>]</li> </ul> <p>Here you can download all three datasets already preprocessed to be used with our implementation, found <a href="https://github.com/uai-ufmg/well-log-imputation" target="_blank" rel="noopener">here</a>.</p> <p> </p> <h3>File Description:</h3> <p>There are six files for each fold partition for each dataset.</p> <ul> <li><code><em>datasetname_fold_k_well_log_metadata_train.json </em></code>: JSON file with general information of the slices of <strong>training </strong>partition of the fold <strong>k</strong>. Contains total number of slices and the number of slices per well.<em> </em></li> <li><em><code>datasetname_fold_k_well_log_metadata_val.json</code> </em>: JSON file with general information of the slices of <strong>validation </strong>partition of the fold <strong>k</strong>. Contains total number of slices and the number of slices per well. </li> <li><em><code>datasetname_fold_k_well_log_slices_train.npy</code>: </em>.npy (numpy) file ready to be loaded with the slices for <strong>training </strong>of the fold <strong>k </strong>already processed. When loaded<em> </em>should have shape of<em> (total_slices, 256, number_of_logs)</em></li> <li><em><code>datasetname_fold_k_well_log_slices_val.npy</code> </em>: .npy (numpy) file ready to be loaded with the slices for <strong>validation </strong>of the fold <strong>k </strong>already processed.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_train.json</code> : </em>JSON file with the slices info for all slices in the <strong>training </strong>partition of the fold <strong>k</strong>. For each slice, 7 data points are provided, the last four are discarded (it would contain other information that was not used). The first three are in order the: origin well name, the starting position in that well, and the end position of the slice in that well.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_val.json</code> </em>: JSON file with the slices info for all slices in the <strong>validation </strong>partition of the fold <strong>k</strong>.</li> </ul>
Phenotypic, weather, soil, and imputed genomic data for the apple REFPOP
<p>Supporting datasets for the article "Integrative multi-environmental genomic prediction in apple" by Jung et al. (2024)<em>.</em></p> <p>Pheno_raw.xlsx – Eleven traits were assessed during up to five years from 2018 to 2022 (Year) at up to five locations* (Country). The traits evaluated were floral emergence (Flowering_begin), flowering intensity (Flowering_intensity), harvest date (Harvest_date),<strong> </strong>total fruit weight (Fruit_weight), fruit number (Fruit_number), single fruit weight (Fruit_weight_single), titratable acidity (Acidity), soluble solids content (Sugar), fruit firmness (Firmness), red over color (Color_over), and russet frequency (Russet_freq_all).</p> <p>Weather_raw.xlsx – Hourly measurements from 2018 to 2022 (Date) of temperature (Temperature), relative humidity (Humidity), and global radiation (Radiation) were obtained at five locations* (Location).</p> <p>Soil_raw.xlsx – Soil characteristics (Variable) were measured at five locations* (Group.1) and two soil depths (Group.2) in 2016.</p> <p>SNPs_final_2022.bed, SNPs_final_2022.bim, SNPs_final_2022.fam – imputed genomic dataset of 303,239 biallelic SNPs in the PLINK format.</p> <p>*The locations correspond to Belgium (BEL), Switzerland (CHE), Spain (ESP), France (FRA) and Italy (ITA).</p>
Improved database of public-private partnerships from World Bank with imputed economic, institutional and conflict data.
<p>The <a href="https://ppi.worldbank.org/en/ppi">World Bank's database</a> on private participation in infrastructure (PPI) projects provides detailed information on these initiatives. However, the original dataset includes imputed macro-level data for the countries that is outdated, lacks assigned ISO country codes, and is not linked to other standard country-level variables necessary for proper analysis and control by territory. In the improved version of the database, 10,958 project observations from 1900 to 2021 have been supplemented with ISO 2 and 3 country codes, enabling accurate integration with other databases. Additionally, 49 new variables related to <a href="https://data.worldbank.org/?cid=ECR_GA_worldbank_EN_EXTP_search&s_kwcid=AL!18468!3!704632427243!b!!g!!world%20bank%20projects&gad_source=1">economic</a>, <a href="https://www.worldbank.org/en/publication/worldwide-governance-indicators">institutional</a>, and <a href="https://www.start.umd.edu/gtd/">conflict data</a> are incorporated by country and year. This enhanced database ensures that researchers can retain critical World Bank information that might otherwise be lost in future updates, as it is not always preserved in repositories</p>
Additional file 2 of Methylation data imputation performances under different representations and missingness patterns
<p>Additional file 2 of Methylation data imputation performances under different representations and missingness patterns</p> <p><a href="https://springernature.figshare.com/articles/journal_contribution/Additional_file_2_of_Methylation_data_imputation_performances_under_different_representations_and_missingness_patterns/12585955/1">Additional file 2 of Methylation data imputation performances under different representations and missingness patterns (figshare.com)</a></p>
Additional file 1 of Methylation data imputation performances under different representations and missingness patterns
<p>Additional file 1 Detailed imputation results per dataset.</p> <p><a href="https://springernature.figshare.com/articles/journal_contribution/Additional_file_1_of_Methylation_data_imputation_performances_under_different_representations_and_missingness_patterns/12585952/1">Additional file 1 of Methylation data imputation performances under different representations and missingness patterns (figshare.com)</a></p>
Additional file 3 of Methylation data imputation performances under different representations and missingness patterns
<p>Additional file 3 Performance comparison between complete (450k) and restricted (21k) datasets.</p> <p><a href="https://springernature.figshare.com/articles/journal_contribution/Additional_file_3_of_Methylation_data_imputation_performances_under_different_representations_and_missingness_patterns/12585958/1">Additional file 3 of Methylation data imputation performances under different representations and missingness patterns (figshare.com)</a></p> <p> </p>
Development of a machine learning model to predict non- durable response to anti-TNF therapy in Crohn's disease using transcriptome imputed from genotypes
<p>This is the expression value predicted using PrediXcan version 7 to find a gene feature that can distinguish between patients with and without effect on infliximab.</p> <p>Among the various tissue models provided by PrediXcan v7, three models were selected and used: whole blood, Colon transverse, and terminal ileum of small intestine, and the predicted gene counts of each model were 6,294, 5,612 and 3,107.</p> <p>For each of the three models, predicted gene expression values and phenotype information per sample were submitted.</p>
Data of: Imputation-free reconstructions of three-dimensional chromosome architectures in human diploid single-cells using allele-specified contacts
<p>These files are results obtained in<br><span><span><span><span>Imputation-free reconstructions of three-dimensional chromosome architectures in human diploid single-cells using allele-specified contacts</span></span></span></span><br>by Yoshito Hirata, Arisa H. Oda, Chie Motono, Masanori Shiro & Kunihiro Ohta.</p> <p>There are 33 files for the corresponding each reconstruction of three-dimensional chromosomone structures<br>for each cell.<br>There are 3D structures for 15 GM cells and 18 PBMC cells, which are obtained from the single cell Hi-C data of Tan et al. Science (2018).</p> <p>For each file, there are 6 columns:<br>The first column corresponds to the allele (0: maternal, 1: paternal)<br>The second column corresponds to the chromosome (1-22: chromosome's number, 23: X, 24: Y)<br>The third column corrsponds to the base point.<br>The fourth column, the fifth column and the sixth column correspond to x-, y-, and z-axes of our reconstruction.</p>
The compiled 8-year dataset (2012-2019) consisting of weekly river water quality indicators (CODMn, DO, NH3-N and PH ) in majors 10 sub-basin of Yangtze river based on imputation of machine learning
<p>Water quality is significantly affected by global climate change and human activities, with diverse critical factors shaping its state in rivers and lakes. In the study, we utilized four indicators to characterize water quality: the physical water quality parameters included dissolved oxygen (DO, mg/L) and PH, while the chemical water quality parameters encompassed chemical oxygen demand (CODMn, mg/L) and ammonia nitrogen (NH3-N, mg/L). This study establishes weekly water quality models for typical 10 sub-basins along the Yangtze River using machine learning methods, which incorporate the impacts of hydro-meteorological and anthropogenic factors.These 10 sub-basins represent the principal tributaries of the Yangtze River basin and include Dongting Lake, the upper Han River, the lower Han River, the Jialing River, the Jinsha River, the Li River, the Min River, Poyang Lake, the Xiang River, and the Yuan River. This data collection was performed by National Environmental Monitoring Centre (http://www.cnemc.cn/sssj/szzdjczb/index_1.shtml). The water quality indicators discussed in this study are assessed in accordance with the national standard GB 3838-2002. Please refer to the paper for details.</p>
Ecological and Forest Inventory of Catalonia: Complete Plant Trait Dataset for Imputation Assessment
<p>Plant trait and forest data were retrieved from the Ecological and Forest Inventory of Catalonia (IEFC), carried out between 1988 and 1998 (Gracia et al. 2000‒2004). The subset of the IEFC was limited to 13 study species. Forest structure, lithology and sampling information for each plot were retrieved from the IEFC database. Climate data were obtained from the Climatic Digital Atlas of Catalonia, with a spatial resolution of 180 m (Ninyerola et al. 2000).</p> <p>We selected five plant traits (leaf mass per area, LMA, mg cm<sup>-2</sup>; leaf nitrogen per unit mass, <em>N</em><sub><em>mass</em></sub>, %mass; maximum tree height, <em>H</em><sub><em>max</em></sub><em>,</em><sub><em> </em></sub>m; wood density, WD, gm cm<sup>-3</sup>; leaf biomass to sapwood area ratio, <em>B</em><sub><em>L</em></sub><em>:A</em><sub><em>S</em></sub>, t m<sup>-2</sup>) that are used to describe major plant functional strategies. The auxiliary variables we considered were species identity, a set of climatic variables (mean annual temperature, annual thermal amplitude, both in °C), a set of forest structure variables (total aboveground biomass [T ha<sup>-1</sup>] and stem density [stems ha<sup>-1</sup>]), a set of topographical variables (county, elevation [m.a.s.l.], slope [°] and aspect), lithology (calcareous, non-calcareous or undetermined) and sampling month.</p>
Comparison and Assessment of Family- and Population-based Genotype Imputation Methods in Large Pedigrees Dataset
<p>Here is the data corresponding to the paper "Comparison and Assessment of Family- and Population-based Genotype Imputation Methods in Large Pedigrees" submitted to Genome Research. The data includes the following:<br> <br> Simulated data for 1200 African and European subjects in pedigrees.<br> Lists of subjects selected by each of the 4 subject selection methods examined; Primus, GIGI-Pick, Exome-Picks, and Random selection. <br> <br> The positions of sparse markers for gl_auto for African and European data.<br> <br> The lists of GWAS SNPs for AFR and EUR. </p> <p>Please see the paper for further details of the data generation, this metadata will be updated following publication. </p>
Summary statistics - Imputed gene associations identify replicable trans-acting genes enriched in transcription pathways and complex traits
<p>Summary statistics for all trans-acting/target gene pairs tested in our manuscript.</p> <p>Preprint available: <a href="https://doi.org/10.1101/471748">https://doi.org/10.1101/471748</a></p>
Imputation of missing land carbon sequestration data in the AR6 Scenarios Database
<p>This repository is linked to the following research paper:</p> <ul> <li>Prütz, R., Fuss, S., and Rogelj, J.: Imputation of missing land carbon sequestration data in the AR6 Scenarios Database, Earth Syst. Sci. Data, 2025. <a href="https://doi.org/10.5194/essd-17-221-2025">https://doi.org/10.5194/essd-17-221-2025</a> </li> </ul> <p>This repository includes: </p> <ul> <li>An imputation dataset for missing land carbon sequestation data of the AR6 Scenarios Database for global scenarios and R10 scenario variants</li> <li>Code to test, compare and visualize the performance of regression models to predict missing land removal data</li> <li>Code to compare and visualize available AR6 land removal data and existing AR6 data reanalyses</li> </ul> <p>The following two datasets are required to replicate the analysis:</p> <ul> <li>Byers, E., Krey, V., Kriegler, E., Riahi, K., Schaeffer, R., Kikstra, J., Lamboll, R., Nicholls, Z., Sandstad, M., Smith, C., van der Wijst, K., Al -Khourdajie, A., Lecocq, F., Portugal-Pereira, J., Saheb, Y., Stromman, A., Winkler, H., Auer, C., Brutschin, E., … van Vuuren, D. (2022). AR6 Scenarios Database [Data set]. In Climate Change 2022: Mitigation of Climate Change (1.1). Intergovernmental Panel on Climate Change. <a href="https://doi.org/10.5281/zenodo.7197970">https://doi.org/10.5281/zenodo.7197970</a></li> <li>Gidden, M., Gasser, T., Grassi, G., Forsell, N., Janssens, I., Lamb, W. F., Minx, J., Nicholls, Z., Steinhauser, J., & Riahi, K. (2023). Dataset for Gidden et.al. 2023 Updated AR6 Mitigation Benchmarks using National Emissions Inventories (Version v2) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10158920">https://doi.org/10.5281/zenodo.10158920</a></li> </ul> <p>The variable imputation is based on the dataset by Byers et al. (2022). The dataset by Gidden et al. (2023) is used for variable comparison. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.