Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
101
datasets available to search
ShareScore release 0.9.0
Dataset results
101 results for “Bias correction”
Dataset for "Bias correction and statistical modeling of variable oceanic forcing of Greenland outlet glaciers" by Verjans et al.
<p>Code and data products associated with "Bias correction and statistical modeling of variable oceanic forcing of Greenland outlet glaciers" by Verjans V., Robel A., Thompson A. F., and Seroussi H.</p> <p>Please see readme file for all the information.</p> <p>Contact: vverjans3@gatech.edu</p> <p>Author: Vincent Verjans</p>
UKCP18 RCM precipitation and temperature bias corrected using non-parametric quantile mapping method
<p>The UKCP18 RCM PPE (Met Office Hadley Centre, 2018) projections of precipitation and daily average temperature were bias adjusted using a non-parametric quantile mapping method based on empirical quantiles (Boe et al, 2007, Gudmundsson et al, 2012). The datasets cover the period from December 1980 to November 2080 and are intended for use in climate change impact assessments, where the bias correction helps reduce biases in multiple statistics while <span>maintaining projected climatic changes</span>.</p> <p>-------------------------------------------------</p> <p>Met Office Hadley Centre (2018): UKCP18 Regional Projections on a 12km grid over the UK for 1980-2080. CEDA, <em>8 March 2022</em>. <a href="https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604">https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604</a></p> <p>Boe, J.; Terray, L.; Habets, F. & Martin, E. Statistical and dynamical downscaling of the Seine basin climate for hydro-meteorological studies. International Journal of Climatology, 2007, 27, 1643-1655, doi: 10.1002/joc.1602.</p> <p>Gudmundsson, L.; Bremnes, J. B.; Haugen, J. E. & Engen-Skaugen, T. Technical Note: Downscaling RCM precipitation to the station scale using statistical transformations - a comparison of methods. Hydrology and Earth System Sciences, 2012, 16, 3383-3390, doi:10.5194/hess-16-3383-2012.</p> <p><strong>Paper Citation:</strong><br>We kindly ask users of this dataset to cite the paper that describes the dataset. The paper is published and can be accessed via the following link: <a href="https://doi.org/10.5194/essd-17-2113-2025" target="_new" rel="noopener">https://doi.org/10.5194/essd-17-2113-2025</a>.</p> <p><br>Please reference the paper as:<br>Reyniers, N., Zha, Q., Addor, N., Osborn, T. J., Forstenhäusler, N., and He, Y.: Two sets of bias-corrected regional UK Climate Projections 2018 (UKCP18) of temperature, precipitation and potential evapotranspiration for Great Britain, Earth Syst. Sci. Data, 17, 2113–2133, https://doi.org/10.5194/essd-17-2113-2025, 2025.</p>
Correction of 4sU induced quantification bias of Spt6 data set
<p>This data set contains the GRAND-SLAM output of the Spt6 data set (https://zenodo.org/record/4275956) after correcting the 4sU induced quantification bias using the correction approach described <a href="https://www.biorxiv.org/content/10.1101/2023.04.21.537786v1">here</a> and implemented in the <a href="https://www.nature.com/articles/s41467-023-39163-4">grandR package</a>.</p> <p>The zip file contains the full GRAND-SLAM output, the *.tsv.gz file is the GRAND-SLAM output table. The RData file contains the grandR object as analyzed in the original original Spt6 data set (see respective zenodo repository).</p> <p> </p> <p> </p> <p> </p>
Model results for Statistical bias correction for CESM-simulated PM2.5
<p>This NC file includes CESM-simulated annual mean aerosol concentrations for 100 years. And the model results are used in paper Statistical bias correction for CESM-simulated PM2.5 (doi: 10.1088/2515-7620/acf917)</p>
Data from: Correcting a bias in the computation of behavioral time budgets that are based on supervised learning
Open the record for dataset details and reuse information.
Accounting for imperfect detection in data from museums and herbaria when modeling species distributions: Combining and contrasting data-level versus model-level bias correction
Open the record for dataset details and reuse information.
Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models
Open the record for dataset details and reuse information.
Data from: Evaluation of different bias correction methods for dynamical downscaled future projections of the California Current Upwelling System
Open the record for dataset details and reuse information.
Phylogenetic comparative methods are problematic when applied to gene trees with speciation and duplication nodes: correcting for biases in testing the ortholog conjecture
<p>This repository contains “manuscript_dunn.RData” file, which is reproduced by using the files and scripts of Dunn et al. (Dunn CW, Zapata F, Munro C, Siebert S, Hejnol A (2018) Pairwise comparisons across species are problematic when analyzing functional genomic data. Proc Natl Acad Sci U S A 115: E409–E417. <a href="http://dx.doi.org/10.1073/pnas.1707515115">doi:10.1073/pnas.1707515115</a>).</p> <p>In this repository, we also supplied “Data_TMRR_latest.rda” file, containing the results generated by using our own scripts. Our scripts are available on GitHub: <a href="https://github.com/tbegum/Testing_the_ortholog_conjecture">https://github.com/tbegum/Testing_the_ortholog_conjecture</a>.</p> <p> </p>
Calibration Bias Evaluation and Correction of S-Band Ground-Based Radar Reflectivity Using Ku-Band Space-Borne Radar Observations Along the East Coast of India
Open the record for dataset details and reuse information.
Bias Corrected and Gap Filled Sentinel-1 and University of Arizona Snow Depth Data
Open the record for dataset details and reuse information.
Fire Weather Index for Europe from Downscaled and Bias-Corrected CMIP6 Model Outputs
<p>This dataset contains the Canadian Forest Fire Weather Index (FWI) calculated from six downscaled and bias-corrected CMIP6 model outputs. The models included are:</p> <ul> <li>ACCESS-CM2 (Ziehn et al. 2020)</li> <li>CanESM5 (Swart et al. 2019)</li> <li>CNRM-ESM2-1 (Séférian et al. 2019)</li> <li>EC-EARTH3 (EC-Earth Consortium 2019)</li> <li>MPI-ESM1-2-HR (von Storch et al. 2017)</li> <li>MRI-ESM2-0 (Yukimoto et al. 2019)</li> </ul> <p>The dataset encompasses four Shared Socio-economic Pathway (SSP) projections:</p> <ul> <li>SSP1-2.6</li> <li>SSP2-4.5</li> <li>SSP3-7.0</li> <li>SSP5-8.5</li> </ul> <p>Each model output has been downscaled to a resolution of 0.0703135°, corresponding to approximately 9km×9km grids before the FWI calculation. The data covers Europe spatially and temporally spans from 1950 to 2080, offering comprehensive insights into past, present, and future fire weather conditions.</p> <p>This dataset supports the manuscript titled <strong>"The fire weather in Europe: large-scale trends towards higher danger" </strong>by Hetzer et al., currently under review in ERL. Detailed instructions for accessing the data can be found in the included README file. </p> <p>Note: Downloads are password protected. Please use "FWI_2024" for access. </p> <p>Funding: The authors acknowledge the financial support of the European Union’s Horizon 2020 research and innovation action for the FirEUrisk project under grant agreement ID: 101003890.</p> <p> </p> <p> </p> <p> </p> <p> </p>
Assessing and Correcting Neighborhood Socioeconomic Spatial Sampling Biases in Citizen Science Mosquito Data Collection
<p>Reporting data from the Mosquito Alert citizen science system, active catch basin surveillance, and mosquito trap surveillance used in "Assessing and Correcting Neighborhood Socioeconomic Spatial Sampling Biases in Citizen Science Mosquito Data Collection."</p> <p>The file named mosquito_alert_adult_bite_reports_Barcelona_2014_2023.Rds includes all adult mosquito and mosquito bite reports received from Barcelona Municipality from the start of the Mosqiuto Alert project in 2014 through the end of 2023. The file named mosquito_alert_validated_albopictus_reports_Barcelona_2014_23.Rds includes all expert-validated <em>Ae. albopictus </em>reports received from Barcelona Municipality during the same time period. The data is stored as RDS files and contain the following fields:</p> <ul> <li><strong>year </strong>- the year in which the report was made. Class = dbl.</li> <li><strong>date </strong>- the date om which the report was made. Class = date.</li> <li><strong>type </strong>- the report type, either adult mosquito ("adult") or mosquito breeding site ("site"). Class = chr.</li> <li><strong>lon</strong> - the longitude of the report location. Class = dbl.</li> <li><strong>lat</strong> - the latitude of the report location. Class = dbl.</li> <li><strong>validation_score</strong> - Entolab validation score. Either 1 (possible <em>Ae. albopictus</em>) or 2 (probable <em>Ae. albopictus</em>). This field is present only in the validated reports data. </li> </ul> <p>The file named active_catch_basin_drain_data.Rds includes information about all catch basin drains in Barcelona Municipality in which the Barcelona Public Health Agency (ASPB) detected mosquito activity as part of its continuous monitoring and control of mosquitoes from 2019 through 2023. The data is stored in an RDS file with the following fields:</p> <ul> <li><strong>any_reports </strong>- dummy variable indicating whether any Mosquito Alert adult mosquito or mosquito bite reports were sent through Mosquito Alert from within 200 m of the catch basin drain during the year in which the ASPB detected mosquito activity in hte catch basin drain. Class = lgl.</li> <li><strong>se_expected</strong> - sampling effort for the 0.025 degree lon/lat sampling cell in which the catch basin drain lies during the year in which the ASPB detected mosquito activity in the drain. This value is taken from the SE_expected variable in the sampling_effort_daily_cellres_025.csv.gz file available at https://zenodo.org/records/12602985. Sampling effort is estimated as the expected number of participants sending at least one report from the cell during the day in question given the the number of participants recorded in the cell that day and the amount of time elapsed since each one began participating in the project. Class = dbl.</li> <li><strong>p_singlehh</strong> - proportion of single-member households in the population of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>mean_age </strong>- mean age of the population of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>mean_rent_consumption_unit</strong> - mean income per consumption unit in the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>popd</strong> - population density of the census tract in which the catch basin drain is located. Class = dbl.</li> <li><strong>id_item </strong>- unique identifier given to the catch basin drain. Drain itentifiers appear multiple times in the data when the ASPB detected activity in the drain in multiple years. Class = dbl.</li> </ul> <p>The file named trap_data.Rds includes information on the adult mosquito trap surveillance analyzed in this article. The data is stored in an RDS file with the following fields:</p> <ul> <li><strong>females </strong>- number of Ae. albopictus females found in the trap. Class = dbl.</li> <li><strong>trap_name</strong> - unique identifier for the trap. Class = chr.</li> <li><strong>trapping_effort</strong> - number of days from when the trap was set to when it was checked. Class = dbl.</li> <li><strong>date</strong> - date on which the trap was checked. Class = date.</li> <li><strong>mean_tm30</strong> - mean temperature for the 30 days leading up to the date on which the trap was checked. Class = dbl.</li> <li><strong>mean_rent_consumption_unit </strong>- mean income per consumption unit for the census tract in which the trap was located. Class = dbl.</li> </ul>
Multi-model ensemble bias-corrected precipitation dataset for historical and future climate (1961–2099) in China
<p>本文基于耦合模式比较项目第六阶段(CMIP6)的27个全球气候模式(GCM),采用随机森林(RF)模型和EQM方法整合27个大气监测模型的降水模拟数据,进一步修正中国综合月降水数据。修正后的降水资料在月降水量和极端降水量方面均明显优于原GCM降水资料。数据以 GeoTIFF 格式,其中嵌入了具有 1° 空间分辨率的地理配准信息。LST在GeoTIFF中的单位是mm。压缩文件被命名为历史文件.zip、SSP126.zip、SSP245.zip 和 SSP585.zip。压缩文件中的每个文件都命名为“yyyymm.tif”,其中“yyyy”和“mm”分别表示年份和月份。例如,文件“196101.tif”存储了 1961 年 1 月中国每月降水量。</p>
HablaRÉ: Self-applied Online Training for Correcting Emotional Reasoning Biases in Public Speaking Anxiety.
ClinicalTrials.gov study NCT07051135. IPD Sharing: NO. Countries: 1. Publications: 3.
Bias corrected climate projections from CMIP6 models for Indian sub-continental river basins
<p>Bias-corrected daily projections of precipitation, maximum temperature, and minimum temperature are developed using output from 13 CMIP6-GCMs for the 18 Indian sub-continental river basins. Spatial resolution of bias corrected projections is 0.25 deg. Each zipped basin file contains projections for 13 CMIP6-GCMs for five scenarios (historical, ssp126, ssp245, ssp370, and ssp585). Please see the readme file for further details.</p>
High-resolution gridded climate data for Europe based on bias-corrected EURO-CORDEX: the ECLIPS-2.0 dataset
<p>We developed a new climate dataset for Europe referred to as ECLIPS (European CLimate Index ProjectionS), which contains gridded data for 80 annual, seasonal, and monthly climate variables for two past (1961-1990, 1991-2010) and five future periods (2011-2020, 2021-2140, 2041-2060, 2061-2080, 2081-2100). The future data are based on five Regional Climate Models (RCMs)driven by two greenhouse gas concentration scenarios, RCP 4.5 and 8.5.</p> <p>The ECLIPS dataset has two versions; ECLIPS 1.1 contains data with spatial resolution of 0.11° × 0.11°, which is the resolution of underlying RCMs. ECLIPS1.1 is available at <a href="https://doi.org/10.5281/zenodo.1181780">https://doi.org/10.5281/zenodo.1181780</a>.</p> <p>The ECLIPS 2.0 presented here contains a subset of climate indices of ECLIPS 1.1, downscaled to the resolution of 30 arcsec by means of the delta correction approach. Both ECLIPS versions were evaluated by testing their relationship with independent station data from the European Climate Assessment (ECA) dataset. Correlations of the empirical testing data to ECLIPS 1.1 ranged from 0.63 to 0.78,and to ECLIPS 2.0 from 0.78 to 0.93. suggesting substantial improvement due to downscaling. A large number of climate projections, time periods and indices as well as the availability of these data at two different spatial resolutions can support diverse studies across a range of disciplines and thus extend our understanding of climate-sensitive dynamics of many social-ecological systems</p> <p>The zipfile ECLIPS2.0 contains 5 folders with subfolders</p> <p>File naming system for the subfolders / folder are as follows</p> <p>ECLIPS2.0_196191: past climate 1961-1990: < climate index><period></p> <p>ECLIPS2.0_199110: past climate 1991-2010 < climate index><period></p> <p>ECLIPS2.0_45 : future climate RCP4.5 <subfolder-Model name> < climate index><period></p> <p>ECLIPS2.0_85 : future climate RCP8.5 <subfolder-Model name> < climate index><period></p> <p>Incase zpfile reader 7zip is not available, please install from here: <a href="https://www.7-zip.org/">https://www.7-zip.org/</a></p>
Bias Corrected Climate Projections from CMIP6 Models for South Asia
<p>Bias-corrected data of precipitation, maximum temperature, and minimum temperature are developed for six countries in South Asia. Each zipped country file contains 13 models, and each model includes five scenarios (historical, ssp126, ssp245, ssp370, and ssp585). Inside a scenario folder, a file named PrecipData can be read as the first three columns from the 3rd row contain year month and day numbers. 1st two rows from the 3rd column represent the longitude and latitude.</p> <p>Separate ObservedData.zip file contains daily observed precipitation (mm), maximum temperature (deg C), and minimum temperature (deg C) data in each grid file.</p>
Data from: Short tree, long tree, right tree, wrong tree: new acquisition bias corrections for inferring SNP phylogenies
Single nucleotide polymorphisms (SNPs) are useful markers for phylogenetic studies owing in part to their ubiquity throughout the genome and ease of collection. Restriction site associated DNA sequencing (RADseq) methods are becoming increasingly popular for SNP data collection, but an assessment of the best practises for using these data in phylogenetics is lacking. We use computer simulations, and new double digest RADseq (ddRADseq) data for the lizard family Phrynosomatidae, to investigate the accuracy of RAD loci for phylogenetic inference. We compare the two primary ways RAD loci are used during phylogenetic analysis, including the analysis of full sequences (i.e., SNPs together with invariant sites), or the analysis of SNPs on their own after excluding invariant sites. We find that using full sequences rather than just SNPs is preferable from the perspectives of branch length and topological accuracy, but not of computational time. We introduce two new acquisition bias corrections for dealing with alignments composed exclusively of SNPs, a conditional likelihood method and a reconstituted DNA approach. The conditional likelihood method conditions on the presence of variable characters only (the number of invariant sites that are unsampled but known to exist is not considered), while the reconstituted DNA approach requires the user to specify the exact number of unsampled invariant sites prior to the analysis. Under simulation, branch length biases increase with the amount of missing data for both acquisition bias correction methods, but branch length accuracy is much improved in the reconstituted DNA approach compared to the conditional likelihood approach. Phylogenetic analyses of the empirical data using concatenation or a coalescent-based species tree approach provide strong support for many of the accepted relationships among phrynosomatid lizards, suggesting that RAD loci contain useful phylogenetic signal across a range of divergence times despite the presence of missing data. Phylogenetic analysis of RAD loci requires careful attention to model assumptions, especially if downstream analyses depend on branch lengths.
Data from: Correction for bias in meta-analysis of little-replicated studies
1. Meta-analyses conventionally weight study estimates on the inverse of their error variance, in order to maximize precision. Unbiased variability in the estimates of these study-level error variances increases with the inverse of study-level replication. Here we demonstrate how this variability accumulates asymmetrically across studies in precision-weighted meta-analysis, to cause undervaluation of the meta-level effect size or its error variance (the meta-effect and meta-variance). 2. Small samples, typical of the ecological literature, induce big sampling errors in variance estimation, which substantially bias precision-weighted meta-analysis. Simulations revealed that biases differed little between random- and fixed-effects tests. Meta-estimation of a one-sample mean from 20 studies, with sample sizes of 3 to 20 observations, undervalued the meta-variance by ~20%. Meta-analysis of two-sample designs from 20 studies, with sample sizes of 3 to 10 observations, undervalued the meta-variance by 15-20% for the log response ratio (lnR); it undervalued the meta-effect by ~10% for the standardised mean difference (SMD). 3. For all estimators, biases were eliminated or reduced by a simple adjustment to the weighting on study precision. The study-specific component of error variance prone to sampling error and not parametrically attributable to study-specific replication was replaced by its cross-study mean, on the assumption of random sampling from the same population variance for all studies, and sufficient studies for averaging. Weighting each study by the inverse of this mean-adjusted error variance universally improved accuracy in estimation of both the meta-effect and its significance, regardless of number of studies. For comparison, weighting only on sample size gave the same improvement in accuracy, but could not sensibly estimate significance. 4. For the one-sample mean and two-sample lnR, adjusted weighting also improved estimation of between-study variance by DerSimonian-Laird and REML methods. For random-effects meta-analysis of SMD from little-replicated studies, the most accurate meta-estimates obtained from adjusted weights following conventionally-weighted estimation of between-study variance. 5. We recommend adoption of weighting by inverse adjusted-variance for meta-analyses of well- and little-replicated studies, because it improves accuracy and significance of meta-estimates, and it can extend the scope of the meta-analysis to include some studies without variance estimates.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.