Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

101

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

101 results for “Bias corrected,”

Learn how ShareScore rates datasets ↗
zenodo40/100

FIGURE 3 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 3. The performance of different implementations of the residual diversity estimate (RDE) under different sampling regimes. (3.1) Mean Spearman's rho values of four implementations of the RDE using Formations as a proxy, with values of PFORM, PLOC and PTAPH variable but equal. (3.2) Mean Spearman's rho values of four implementations of the RDE using Localities as a proxy. (3.3) Mean Spearman's rho values of four implementations of the RDE, all using the Smith and McGowan method. (3.4) Mean Spearman's rho values of the taxic and phylogenetic diversity estimate compared to those of the optimum implementation of the RDE. The dashed red line indicates the critical value at p=0.05. Abbreviations as in Table 1.

opencc-by-4.0Nov 2015View details →
zenodo40/100

FIGURE 2 in A simulation-based examination of residual diversity estimates as a method of correcting for sampling bias

FIGURE 2. An illustration of how sampling proxies are generated in this simulation. This schematic illustrates which formations and localities in a single time bin contain fossils of at least one species of the simulated clade after application of the taphonomic filter. Formations and localities are removed at random, representing a lack of sampling. Note that the number of clade-bearing formations and localities does not necessarily equal the number of formations and localities sampled, allowing the generation of four sampling proxies.

opencc-by-4.0Nov 2015View details →
zenodo40/100

Bias-corrected monthly air temperature data over South Siberia for 1979-2020 (CTSS 1.0)

<p>Bias-<strong>C</strong>orrected Air <strong>T</strong>emperature data over <strong>S</strong>outh <strong>S</strong>iberia (<strong>CTSS 1.0</strong>) contains monthly air temperature at 2m for the area within the coordinates 50&ndash;65 N, 60&ndash;120 E for the period from January 1979 to December 2020. CTSS data were combined from monthly total air temperture data from ERA5 reanalysis European Centre for Medium-Range Weather Forecasts (Copernicus Climate Change&hellip;, 2017) and temperature data records from ground weather stations (Bulygina et al., 2014). The ERA5 data were scaled according to the derived correction coefficient. The additive coefficient for each month and weather station were calculated and extrapolated to the study area using the ordinary kriging method. Data spatial resolution is 0.25&deg; in the latitude and 0.25&deg; in the longitude. CTSS reproduces the spatial variability of temperature more precisely than can be done from the weather station observation network. Data provided in NetCDF (Network Common Data Form) format.</p> <p>Copernicus Climate Change Service (C3S), 2017. <em>ERA5: Fifth generation of ECMWF atmospheric reanalyses of the global climate.</em> Copernicus Climate Change Service Climate Data Store (CDS), Available at: <a href="https://cds.climate.copernicus.eu/cdsapp#!/home"><em>https://cds.climate.copernicus.eu/cdsapp#!/home</em></a></p> <p><em>Bulygina O.N., Razuvaev V.N., Trofimenko L.T., Shvets N.V., 2014. &nbsp;Description of the monthly air temperature data at weather stations of Russia.</em> Certificate of state registration of the database No. 2014621485 Available at:&nbsp; <a href="http://meteo.ru/data/156-temperature"><em>http://meteo.ru/data/156-temperature</em></a></p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Bias-corrected data from the preoperational MiKlip system for decadal climate predictions used in the PNRA-IPSODES project

<p>This dataset contains a selection of bias-corrected data from the&nbsp;preoperational MiKlip system for decadal climate predictions (Mueller et al., 2018) used within the Italian research project PNRA18_00199-IPSODES. The adopted method for bias correction is described in the file bias_correction.pdf. Also data from the assimilation run are provided. Nomenclature of variables follows that of the original MiKlip output.</p> <p>Mueller, W., et al. A Higher‐resolution Version of the Max Planck Institute Earth System Model (MPI‐ESM1.2‐HR). J. Adv. Model. Earth Syst. 10, 1383-1413 (2018)</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Bioclimatic outputs for Last Glacial Maximum South America for LPX, bias-corrected to pollen records

<p>LPX model output for South America for the Last Glacial Maximum. We provide outputs for LPX driven by four GCM simulated climates along with an ensemble average:</p> <ul> <li>MIROC.tar.gz: LPX driven by MIROC3.2</li> <li>FGOALS.tar.gz: driven by FGOALS-1.0g</li> <li>HAD.tar.gz: HadCM3M2</li> <li>CNRM.tar.gz: CNRM-CM33</li> <li>Ensemble.tar.gz: The mean of each output variable for the four models.</li> </ul> <p>&nbsp;</p> <p>Driving data comes from the Palaeoclimate Modelling Intercomparison Project Phase II (PMIP2)<sup>1,2</sup>. See Sato et al. <sup>3</sup> for modelling protocol.</p> <p>Each model&rsquo;s directory contains &ldquo;uncorrected&rdquo; and &ldquo;corrected&rdquo; directories. With each of these are bioclimatic maps outputted from LPX and biome information:</p> <ul> <li>fpc.nc: fractional projected cover of all vegetation</li> <li>height.nc: mean height of vegetations</li> <li>gdd.nc: Growing Degree Days base 2</li> <li>tropical.nc: proportion of vegetated areas taken up by tropical trees and c4 grasses</li> <li>temperate.nc: proportion of vegetated areas taken up by temperate trees and c3 grasses</li> <li>evergreen.nc: proportion of tree cover that is composed of evergreen trees</li> <li>biome.nc: the assigned biomes from these data, based on a modified version of Sato et al. <sup>3</sup>. See &ldquo;biomisation&rdquo; below.</li> <li>cluster.nc: In &ldquo;corrected&rdquo; only. The spatial location of kmean clusters of fpc vs height. See <sup>4</sup> for details.</li> </ul> <p>&nbsp;</p> <p>&ldquo;Uncorrected&rdquo; is from Sato et al. <sup>3</sup>. &ldquo;Corrected&rdquo; is bias-corrected to match the 42 pollen-core observations taken from Marchant et al. <sup>5</sup> We do this by shifting the total vegetation cover and composition, height, and growing degree day (GDD) DVM output to the closest boundary of the corresponding biome of the pollen core in that specific location. We then extrapolate this correction between pollen-core locations across the Neotropics. See Kelley et al. <sup>4</sup>&nbsp; for details.</p> <p>&nbsp;</p> <p><strong>Biomeisation</strong></p> <p>&quot;Biomeisation.png&quot; displays the scheme. We primarily split biomes by FPCs of 0.3 and 0.6, with biomes &gt; 0.6 split by a height of 10m. Forests (&gt;0.6 FPC and &gt; 10m) is split by GDD, Evergreen FPC (EG) and Tropical or temperate FPC (TR, TM). Likewise, we split FPCs&gt; 0.6 and heights &lt;10m into savanna, woodland and parkland using EG and TR. We additionally assign Tropical savanna &gt;5m to Woodland/Tropical savanna. We divided desert, dry grassland and (shrub)-tundra by FPC of 0.3 and GDD of 350&deg;C. See Kelley et al. <sup>4</sup> for details.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

UKCP18 RCM precipitation and temperature bias corrected using ISIMIP3BA change-preserving quantile mapping.

<p>We present bias-corrected UK Climate Projections 2018 (UKCP18; Met Office Hadley Centre, 2018) regional datasets for temperature, precipitation, and potential evapotranspiration (1981-2080). All 12 members of the 12 km ensemble were corrected using quantile mapping and a change-preserving variant (Lange, 2019; Lange, 2020). Both methods effectively reduce biases in multiple statistics, while maintaining projected climatic changes. We provide guidance on using the bias-corrected datasets for climate change impact assessment. Please find a detailed description and evaluation in the metadata and accompanying data paper (Reyniers et al., 2025).</p> <p>---</p> <p>Met Office Hadley Centre (2018): UKCP18 Regional Projections on a 12km grid over the UK for 1980-2080. CEDA, <em>8 March 2022</em>. <a href="https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604">https://catalogue.ceda.ac.uk/uuid/589211abeb844070a95d061c8cc7f604</a></p> <p>Lange, S. (2019). Trend-preserving bias adjustment and statistical downscaling with ISIMIP3BASD (v1. 0). <em>GMD,</em> <em>12</em>(7), 3055-3070.</p> <p>Lange, S. (2020). ISIMIP3BASD (2.4.1). Zenodo. https://doi.org/10.5281/zenodo.3898426</p> <p>Reyniers, N., Zha, Q., Addor, N., Osborn, T. J., Forstenh&auml;usler, N., &amp; He, Y. (2025). Two sets of bias-corrected regional UK Climate Projections 2018 (UKCP18) of temperature, precipitation and potential evapotranspiration for Great Britain. <em>Earth System Science Data</em>,&nbsp;<em>2025, 17(5)</em>, 2113&ndash;2133.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Bias corrected era5 skin temperature over the Arctic sea ice – 1981 to 2018 monthly means and climatology

<p>This dataset is generated in the context of the peer-reviewed study of&nbsp;Zampieri et al., 2023. The users can find a detailed description of the bias correction strategy and information on the scientific value of the dataset in the paper. Please, do not hesitate to contact me to obtain further information and suggestions on how to employ this&nbsp;dataset for your specific purpose. &nbsp;</p> <p><strong>References:</strong></p> <p>Zampieri, L.,<strong>&nbsp;</strong>Arduini, G., Holland, M., Keeley, S., Mogensen, K., Shupe, M., Tietsche, S. (2023) A machine learning correction model of the winter clear-sky temperature bias over the Arctic sea ice in atmospheric reanalyses.&nbsp;<em>Monthly Weather Review</em>. DOI:<a href="https://doi-org.cuucar.idm.oclc.org/10.1175/MWR-D-22-0130.1">10.1175/MWR-D-22-0130.1</a></p> <p><strong>Acknowledgments:</strong></p> <p>As part of the Virtual Earth System Research Institute (VESRI), funding for the Multiscale Machine Learning In coupled Earth System Modeling (M2LInES) project was provided to Lorenzo Zampieri&nbsp;by the generosity of Eric and Wendy Schmidt by recommendation of the Schmidt Futures program.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Dataset of trend-preserving bias-corrected daily temperature, precipitation and wind from NEX-GDDP and CMIP5 in the Qinghai-Tibet Plateau——Part Ⅱ

<p>A bias-corrected dataset containing daily meteorological data of the Qinghai-Tibet Plateau has been generated, by using a trend-preserving bias-correction, the Inter-Sectoral Impact Model Intercomparison Project (ISI-MIP) approach together with a high-quality gridded meteorological dataset based on ground observation (CN05.1). The data set contains daily bias-corrected values of maximum/minimum near-surface air temperature, precipitation and mean near-surface wind speed from 15 models from the Fifth Phase of the Coupled Model Intercomparison Project (CMIP5) and their downscaled high-resolution dataset (NEX-GDDP) in the Qinghai-Tibet Plateau (QTP) during 1986-2095. This dataset can provide important reference for the study on future climate change and its impacts in the Qinghai-Tibet Plateau region.</p> <p><strong>Note: For Tmin in historical periods, the values larger than 2606 refer to no data. Set them to NaN before using, for example (Matlab): Tmin(Tmin&gt;2606)=nan;</strong></p> <p>More details about this dataset can be found in the article: S. Chen, T. Ye, W. Liu, A. Wang and P. Shi. Evaluation and bias correction of the historical and future near-surface climate forcing in NEX-GDDP and CMIP5 over the Qinghai-Tibet plateau[J], Plateau Meteorology (in Chinese), 2020, DOI: 10.7522/j.issn.1000-0534. 2020. 00019.</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Dataset of trend-preserving bias-corrected daily temperature, precipitation and wind from NEX-GDDP and CMIP5 in the Qinghai-Tibet Plateau——Part Ⅰ

<p>A bias-corrected dataset containing daily meteorological data of the Qinghai-Tibet Plateau has been generated, by using a trend-preserving bias-correction, the Inter-Sectoral Impact Model Intercomparison Project (ISI-MIP) approach together with a high-quality gridded meteorological dataset based on ground observation (CN05.1). The data set contains daily bias-corrected values of maximum/minimum near-surface air temperature, precipitation and mean near-surface wind speed from 15 models from the Fifth Phase of the Coupled Model Intercomparison Project (CMIP5) and their downscaled high-resolution dataset (NEX-GDDP) in the Qinghai-Tibet Plateau (QTP) during 1986-2095. This dataset can provide important reference for the study on future climate change and its impacts in the Qinghai-Tibet Plateau region.</p> <p><strong>Note: For Tmax in historical periods, the values larger than 2606 refer to no data. Set them to NaN before using, for example (Matlab): Tmax(Tmax&gt;2606)=nan;</strong></p> <p>More details about this dataset can be found in the article: S. Chen, T. Ye, W. Liu, A. Wang and P. Shi. Evaluation and bias correction of the historical and future near-surface climate forcing in NEX-GDDP and CMIP5 over the Qinghai-Tibet plateau[J], Plateau Meteorology (in Chinese), 2020, DOI: 10.7522/j.issn.1000-0534. 2020. 00019.</p>

opencc-by-4.0Apr 2020View details →
dryad36/100

Accounting for imperfect detection in data from museums and herbaria when modeling species distributions: Combining and contrasting data-level versus model-level bias correction

The digitization of museum collections as well as an explosion in citizen science initiatives has resulted in a wealth of data that can be useful for understanding the global distribution of biodiversity, provided that the well-documented biases inherent in unstructured opportunistic data are accounted for. While traditionally used to model imperfect detection using structured data from systematic surveys of wildlife, occupancy models provide a framework for modelling the imperfect collection process that results in digital specimen data. In this study, we explore methods for adapting occupancy models for use with biased opportunistic occurrence data from museum specimens and citizen science platforms using 7 species of Anacardiaceae in Florida as a case study. We explored two methods of incorporating information about collection effort to inform our uncertainty around species presence: (1) filtering the data to exclude collectors unlikely to collect the focal species and (2) incorporating collection covariates (collection type, time of collection, and history of previous detections) into a model of collection probability. We found that the best models incorporated both the background data filtration step as well as collector covariates. Month, method of collection and whether a collector had previously collected the focal species were important predictors of collection probability. Efforts to standardize meta-data associated with data collection will improve efforts for modeling the spatial distribution of a variety of species.

opencc-zeroJun 2021View details →
zenodo36/100

ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data. Additional Datasets.

<p>Data supporting the main figures of the "ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data" manuscript and a copy of the ParaMask software and scripts for analysis, and SV calls from longreads. README files are included.</p>

opengpl-3.0-or-laterApr 2024View details →
dryad36/100

Data from: Evaluation of different bias correction methods for dynamical downscaled future projections of the California Current Upwelling System

<p class="Abstract">Biases in global Earth System Models (ESMs) are an important source of errors when used to obtain boundary conditions for regional models. Here we examine historical and future conditions in the California Current System (CCS) using three different methods to force the regional model: (1) interpolation of ESM output to the regional grid with no bias correction; (2) a "seasonally-varying" delta method that obtains a season-dependent mean climate change signal from the ESM for a 30-year future period; and (3) a "time-varying" delta method that includes the interannual variability of the ESM over the 1980–2100 period. To compare these methods, we use a high-resolution (0.1˚) physical-biogeochemical regional model to dynamically downscale an ESM projection under the RCP8.5 emission scenario. Using different downscaling methods, the sign of future changes agrees for most of the physical and ecosystem variables, but the spatial patterns and magnitudes of these changes differ, with the seasonal- and time-varying delta simulations showing more similar changes. Not correcting the ESM forcing leads to amplification of biases in some ecosystem variables as well as misrepresentation of the California Undercurrent and CCS source waters. In the non-bias corrected and time-varying delta simulations, most of the ecosystem variables inherit trends and decadal variability from the ESM, while in the seasonally-varying delta simulation, the future variability reflects the observed historical variability (1980–2010). Our results demonstrate that bias correcting the forcing prior to downscaling improves historical simulations and that the bias correction method may impact the spatial and temporal variability of future projections. </p>

opencc-zeroNov 2023View details →
dryad36/100

Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models

<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Replication Data for the Paper "Is there a secular decline in disruptive patents? Correcting for measurement bias"

<p><strong>Working Paper Title: The Illusive Slump of Disruptive Patents</strong></p> <p>The repository contains replication data and scripts for the paper:&nbsp;</p> <p>Jeffrey T. Macher, Christian Rutzer, Rolf Weder,<br>Is there a secular decline in disruptive patents? Correcting for measurement bias,<br>Research Policy,<br>Volume 53, Issue 5,<br>2024,<br>104992,<br>ISSN 0048-7333,<br>https://doi.org/10.1016/j.respol.2024.104992</p> <p>The core component of the repository is the file `replication_file.R` which is an R script to replicate all figures and tables of the paper.</p> <p>The script relies on multiple data files. The datasets contain CD-values for granted USPTO utility patents for the years 1976-2016. All data is available in two formats, `.fst' (from the R package fst) and `.csv'.</p> <p>&nbsp;</p> <p><strong>Dataset Descriptions</strong></p> <p><strong>Main Datasets:</strong></p> <p>1. `dat_cd5_no_trunc`: This dataset contains the CD5 index of patents based on a method without truncation.&nbsp;</p> <p>2. `dat_cd5_trunc_1975`: This dataset contains the CD5 index of patents based on a truncation method as in Park et al. (Nature, 2023).</p> <p>3. `dat_cd5_no_trunc_app_adj`: This dataset contains the CD5 index of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>&nbsp;</p> <p><strong>Additional Datasets:</strong></p> <p>4. `dat_cd5_trunc_1985`: This dataset contains the CD5 index of patents based on a truncation of all backward citations to patents published before 1985.</p> <p>5. `dat_cd5_trunc_1995`: This dataset contains the CD5 index of patents based on a truncation of all backward citations to patents published before 1995.</p> <p>6. `dat_cd5_no_trunc_app`: This dataset contains the CD5 index of patents based on a method without truncation and including citations to patent applications granted by 2021, as well as those not yet granted.&nbsp;</p> <p>7. `age_bwc_untrunc`: This dataset contains the age of backward citations using untruncated data.</p> <p>8. age_bwc_trunc: This dataset contains the age of backward citations using truncated data as in Park et al. (Nature, 2023).</p> <p>9. `age_bwc_untrunc_app_adj`: This dataset contains the age of backward citations using untruncated data and considering citations of patent applications granted until 2021.</p> <p>10. `dat_cd10_trunc_1975`: This dataset contains the CD10 index of patents based on a truncation method as in Park et al. (Nature, 2023).&nbsp;</p> <p>11. `dat_cd10_no_trunc_app_adj`: This dataset contains the CD10 index of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>12. `dat_cd2021_trunc_1975`: This dataset contains the CD index as of 2021 of patents based on a truncation method as in Park et al. (Nature, 2023).&nbsp;</p> <p>13. `dat_cd2021_no_trunc_app_adj`: This dataset contains the CD index as of 2021 of patents based on a method without truncation and including citations to patent applications granted by 2021.&nbsp;</p> <p>&nbsp;</p> <p>To successfully run the `replication_file.R' script, make sure all data files are in the directory and the `mainDir1' variable at the beginning of the script is set to the correct path to where the data is stored. In addition, set the `mainDir2' variable to the folder where you want to store the figures created by the script.&nbsp;</p> <p>If you have any questions, please contact christian.rutzer@unibas.ch</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Bias-corrected EURO-CORDEX daily temperature and precipitation dataset for Hungary

<p>These datasets contain bias-corrected regional climate model outputs for daily precipitation, near surface mean-, minimum- and maximum temperature for the historical period 1993-2005 on a regular 0.11&deg;x0.11&deg; lon/lat grid (between latitudes 45.6825&deg;N and 48.6525&deg;N, and longitudes 15.9075&deg;E and 22.9475&deg;E).</p> <p>&nbsp;</p> <p>The datasets contain daily outputs of the following high-resolution (0.11&deg;) regional climate models from the framework of EURO-CORDEX (Jacob et al., 2014):</p> <p>-CCLM</p> <p>-HIRHAM</p> <p>-RACMO</p> <p>-RCA</p> <p>-REMO</p> <p>&nbsp;</p> <p>The reference dataset is HUCLIM, which covers Hungary for the period 1971-2022 (downloaded in 2023) produced by the HungaroMet Hungarian Meteorological Service.</p> <p>File format: NetCDF</p> <p>All bias-corrected data produced by the use of HUCLIM have been created following the work of Mezghani et al. (2017).</p> <p>&nbsp;</p> <p>References:</p> <p>Jacob, D., Petersen, J., Eggert, B., Alias, A., Christensen, O.B., Bouwer, L.M., Braun, A., Colette, A., D&eacute;qu&eacute;, M., Georgievski, G., Georgopoulou, E., Gobiet, A., Menut, L., Nikulin, G., Haensler, A., Hempelmann, N., Jones, C., Keuler, K., Kovats, S., Kr&ouml;ner, N., Kotlarski, S., Kriegsmann, A., Martin, E., van Meijgaard, E., Moseley, C., Pfeifer, S., Preuschmann, S., Radermacher, C., Radtke, K., Rechid, D., Rounsevel, M., Samuelsson, P., Somot, S., Soussana, J.-F., Teichmann, C., Valentini, R., Vautard, R., Weber, B. and Yiou, P. (2014) EURO-CORDEX New high resolution climate change projections for European impact research. Reg. Environ. Change, 14, 563&ndash;578.&nbsp;<a href="https://doi.org/10.1007/s10113-013-0499-2" target="_blank" rel="noopener">https://doi.org/10.1007/s10113-013-0499-2</a></p> <p>Mezghani, A., Dobler, A., Haugen, J.E., Benestad, R.E., Parding, K.M., Piniewski, M., Kardel, I. and Kundzewicz, Z.W. (2017) CHASE-PL Climate Projection dataset over Poland &ndash; bias adjustment of EURO-CORDEX simulations. Earth Syst. Sci. Data, 9, 905&ndash;925.&nbsp;<a href="https://doi.org/10.5194/essd-9-905-2017" target="_blank" rel="noopener">https://doi.org/10.5194/essd-9-905-2017</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
dryad36/100

Data from: Correcting a bias in the computation of behavioral time budgets that are based on supervised learning

<p>Supervised learning of behavioral modes from body-acceleration data has become a widely used research tool in Behavioral Ecology over the past decade. One of the primary usages of this tool is to estimate behavioral time budgets from the distribution of behaviors as predicted by the model. These serve as the key parameters to test predictions about the variation in animal behavior. In this paper we show that the widespread computation of behavioral time budgets is biased, due to ignoring the classification model confusion probabilities. Next, we introduce <em>the confusion matrix correction for time budgets</em> -- a simple correction method for adjusting the computed time budgets based on the model's confusion matrix. Finally, we show that the proposed correction is able to eliminate the bias, both theoretically and empirically in a series of data simulations on body acceleration data of a fossorial rodent species (Damaraland mole-rat, <em>Fukomys damarensis</em>). Our paper provides a simple implementation of <em>the confusion matrix correction for time budgets</em>, and we encourage researchers to use it to improve accuracy of behavioral time budget calculations.</p>

opencc-zeroMar 2022View details →
zenodo36/100

JAG Bias Corrected SPP

<p>This dataset belongs to the paper: Comparison of Bias-Corrected MultiSatellite Precipitation Products by Deep Learning Framework</p> <p>Daily dataset from 01/01/2016 to 12/2019</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Data supplementing the article "Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring" V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal

<p>These data supplement the article &quot;Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring&quot; V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal</p> <p>The directory contains the following files:</p> <p>1<strong>5&nbsp;fastq files raw reads (5 mock communities, 3 replicates)</strong><strong>.rar </strong>- contains the 15&nbsp;fastq files provided by the sequencing platform with demultiplexed DNA reads (raw data prior any bioinformatics treatments).</p> <p><strong>15 fastq files information.xlsx</strong>&nbsp;:</p> <p>- contains the information relative to the 15 fastq files corresponding to the PGM raw data of the 5 mock communities (sequenced with 3 replicates), including:&nbsp;the ID of the fastq files, the mock community name,&nbsp;the replicate number, the final sample Id and the number of raw reads per fastq file.</p> <p>- contains the information of the proportion of the 8 diatoms species (%) used to create the 5 mock communities (estimated from microscopy).</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Dataset of the simulations in the GōMartini 3: From large conformational changes in proteins to environmental bias corrections

<p>This section describes the data files for the nanomechanics of SARS-CoV-2 RBD in complex with a potent nanobody (H11-H4) by GōMartini approach.</p> <p>(1) The <strong>data_4iu3.tar.xz</strong> file has all the files necessary to reproduce the pulling simulations of PDB ID 4IU3 using GōMartini methodology. The files are described below.</p> <p>---------------------------------------------------------------------------------------------------------------------------------------------</p> <p>https://www.rcsb.org/structure/4iu3</p> <p>Cohesin-dockerin -X domain complex from Ruminococcus flavefacience</p> <p>&nbsp;</p> <p>https://github.com/rams-research/OVrCSU</p> <p>http://pomalab.ippt.pan.pl/GoContactMap/</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------</p> <p>4iu3.pdb&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : All-atoms PDB file</p> <p>4iu3.map&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : All-atoms OVrCSU contac map file</p> <p>A.pdb,B.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;: All-atoms Structures of Cohesin and Dockering, later referenced as molA and molB</p> <p>index.ndx &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : base index file</p> <p>martini_v3*.itp &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;: Martini3 force field definitions</p> <p>&nbsp;</p> <p>GoMartini, topology files:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; molA.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; molA_exclusions_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; molB.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; molB_exclusions_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; go-table_VirtGoSites_interface.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; go-table_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; topol.top</p> <p>&nbsp;</p> <p>GoMartini, GROMACS MDP equilibration and final GRO files:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Pre-equilibration:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; min.mdp, min.gro</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; npt.mdp, npt,gro</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; nvt.mdp, nvt.gro</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Equilibration MD run:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; md-corr.log</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; md-corr.mdp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; md-corr.edr</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; md-corr.gro</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; md-corr.tpr</p> <p>&nbsp;</p> <p>Topology and force field definitions:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/BB-part-def_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/go-table_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/index.ndx</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/martini_v3.0.4.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/martini_v3.0_ions.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/martini_v3.0_phospholipids.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/martini_v3.0_solvents.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/molA_exclusions_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/molA.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/molB_exclusions_VirtGoSites.itp</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/molB.itp</p> <p>&nbsp;</p> <p>Pulling MD runs from 100 runs ({1-100}), GROMACS files:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}-trajectory-protein.pdb : Trajectory in PDB format</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}_pullf.xvg&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : force profile</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}_pullx.xvg&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;: distance profile</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}.log&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : log files</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}.edr&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;: EDR files</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; GGG_pull/md_pull_{1-100}.tpr&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;: TPR files</p> <p>&nbsp;</p> <p>2) The&nbsp;<strong>data_6zh9_1.tar.xz</strong>, <strong>data_6zh9_2.tar.xz</strong>,<strong> &nbsp;data_6ZH9_GM2.tar.xz</strong>,<strong> </strong>and<strong> data_6ZH9_GM3.tar.xz </strong>files contain</p> <p>The dataset comprises a total of 50 replicas, each with a total of 751 snapshots, and the mdp and tpr files, which are located in the CG_Pull/Trajectories folder, for the RBD-H11-H4 complex. The gro, mdp and tpr files for NVT and NPT equilibrations are included.</p> <p>The force and displacement xvg files for each replica are included in the GC_Pull folder (pullf*.xvg and pullx*.xvg, respectively). The output files from GōMartini are included.</p> <p>The system consists of a single-domain antibody (i.e. nanobody) and the receptor-binding domain (RBD) portion of the SARS-CoV-2 spike protein. In this regard, the nanobody named H11-H4 and the RBD form a mechanostable protein complex with PDB ID: 6ZH9. The entire system was modeled by the Martini 3 force field. The same protocol as for the XMod-Doc:Coh complex was employed. The GōMartini model requires a total of 715 contacts divided into 404 and 285 for the RBD and the H11-H4 respectively and the protein complex interface was represented by 26 contacts. The value of the dissociation energy of the LJ potentials was set to a value of &epsilon; = 15.0 kJ/mol. The dimensions of the water box for the RBD-Nb complex were 16x12x90 nm3 with a 0.15 M NaCl. There were 135,669 CG water beads representing 542,676 water molecules.</p> <p>To conduct the nanomechanical studies, the positions of RBD residues GLY-526, PRO-527, and&nbsp; LYS-528 were kept frozen in the z-axis and the coordinates of residues SER-126, SER-127, and LYS-128 in H11-H4 were kept frozen in x- and y-axis, and were chosen for SMD simulation at constant speed.</p> <p>Two different protocols according to all-atom SMD simulation were considered. i) in&nbsp;<strong>data_6zh9_1.tar.xz: </strong>Restraints were applied to the backbone atoms of the RBD with a spring contact of kb =1000 kJ/(mol&middot;nm2) to avoid large deformation and the residue in the SMD protocol with velocity equal to 5x10-3 nm/ps was coupled to another harmonic potential with a spring constant equal to 600 kJ/(mol&middot;nm2) and ii)&nbsp;<strong>data_6zh9_2.tar.xz: </strong>Pulling velocity was equal to 1x10-4 nm/ps was coupled to another harmonic potential with a spring constant equal to 600 kJ/(mol&middot;nm2).&nbsp;</p> <p>The supplementary files <strong>data_6ZH9_GM2.tar.xz</strong> and <strong>data_6ZH9_GM3.tar.xz</strong> followed a similar protocol as <strong>data_6ZH9_2.tar.xz</strong>, with the following differences: the box size was set to 10 &times; 10 &times; 60 nm, and the spring constant was adjusted to 37.6 kJ/(mol&middot;nm&sup2;). For the GōMartini 2 model, the default value of &epsilon; = 9.414 kJ/mol was used.</p> <p>3) The <strong>IDP_GoMartini3.tar.gz </strong>file has the following folders:</p> <p>a) Atomistic</p> <p>The _atomistic_ directory contains subfolders for each of the IDPs tested. Each subdirectory contains the topology for the protein and an initial starting structure. There is also a folder of the mdp files used. In each simulation, the correct temperature as described in the manuscript was applied.</p> <p>b) Martini</p> <p>The _martini_ directory contains subfolders for each of the IDPs tested. Each subdirectory contains topology for the corresponding protein model tested, and the starting structure used. The TEMP variable in the mdp files was replaced with the experimentally corresponding temperature in each simulation.</p> <p>c) Condensate</p> <p>The _condensate_ directory contains two systems. The artifical IDP (_aIDP_) condensate system of Dzuricky et al. and the short peptides of Abbas et al. (_peptides_)</p> <p>_aIDP_ has two systems comparing the native and optimised topologies for the phase separation of the aIDP.</p> <p>_peptides_ contains a single system with topologies for protonated and deprotonated synthetic peptides, as well as the starting structures, and topology files for the differently rescaled protein-water interactions used.</p> <p>4) The WALPDATA.tar.gz has the following folders:</p> <p>a) ProteinITPs has all GROMACS itp files for the WALP &alpha;-helices &ndash; 16, 19, 23, and 27 denoted as WALP16, WALP29, WALP23 and WALP27 as well as the GōMartini inputs</p> <p>b) mdps and the script for preparing the mebrane for the WALP peptides PrepareMembranePeptideSim.ipynb</p> <p>5) The <strong>BetaSheet_GoMartini.tar.gz</strong> file has the following folders:</p> <p>molecule_0.itp &nbsp; is the itp file for beta strand 1.<br>molecule_1.itp &nbsp; is the itp file for beta strand 2.</p> <p>martini.itp &nbsp; &nbsp; &nbsp;is the Martini forcefield itp used.</p> <p>rCSU.map &nbsp; &nbsp; &nbsp; &nbsp; is the rCSU map generated.<br>OV.map &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; is the OV map generated.</p> <p><br>GoMartini files:</p> <p>&nbsp; &nbsp; molecule_0_exclusions_VirtGoSites.itp&nbsp;<br>&nbsp; &nbsp; molecule_0_BB-part-def_VirtGoSites.itp<br>&nbsp; &nbsp; molecule_0_go-table_IDPsolubility.itp &nbsp;<br>&nbsp; &nbsp; molecule_0_go-table_VirtGoSites.itp&nbsp;<br>&nbsp; &nbsp; molecule_0_go4view_harm.itp &nbsp;<br>&nbsp; &nbsp; BB-part-def_VirtGoSites.itp&nbsp;<br>&nbsp; &nbsp; go-table_VirtGoSites.itp&nbsp;</p> <p><br>mdp files used:</p> <p>&nbsp; &nbsp; min.mdp&nbsp;<br>&nbsp; &nbsp; rel_310.mdp&nbsp;<br>&nbsp; &nbsp; prod_310.mdp&nbsp;</p> <p>&nbsp;</p> <p>rada_sheet_canonical.pdb &nbsp; &nbsp; is the PDB file of the AA RADA16 structure.<br>rada_cg.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;is the PDB file of the Martini RADA16 structure.</p> <p>&nbsp;</p> <p>Inside folder Negative_Control is the topology file for<br>the system with no structural or interaction bias.</p> <p><br>Inside folder GoVirt_Bias is the topology file for<br>the system with a GoMartini model applied between the<br>two strands.</p> <p><br>Inside folder Scaled_Water is the topology file for<br>the system with GoVirt water interaction scaling.</p> <p>6) The <strong>md_inputs_go_IDPs_PETER.tar.gz</strong> file has the following folders:</p> <p>Directories "a1", "hst5", and "pnt" contain the input files for proteins with Go potentials.&nbsp;And in the "controls" directory, input files for the same proteins, without Go potentials.</p> <p>To run the simulations, the user can run the following commands:</p> <p>gmx grompp -f md.mdp -p topol.top -c pre_md.gro -o md<br>gmx mdrun -s md.tpr -deffnm md</p> <p>Additionally, atomistic structures are available in .gro format - {protein}_atomistic.gro</p> <p>7) The <strong>repository_Hafez.zip&nbsp;</strong></p> <p>In the folders 1AOH, 1TIT, 1UBQ, 3W0K, AQP1, and Ist2, you will find the necessary files (ITP, MDP, and GRO) to repeat the simulations.</p> <p>In the Martini3FF folder, you can find the MARTINI3 force field files.</p> <p>8) The <strong>Hairpin-M3-15.tar.xz</strong> and <strong>Hairpin-M2.tar.xz</strong> and <strong>Hairpin-M2-15.tar.xz</strong>&nbsp;</p> <p>contains the information mdp, itp and trajectories generated for the study of folding simulation with GoMartini implemented in 2017 and the latest implementation on top of the Martini 3 force field.</p> <p>The Hairpin-M2, Hairpin-M2-15, and Hairpin-M3-15 folders contain all the necessary files to reproduce the pulling simulations of PDB ID 1GB1 (a novel, highly stable fold of the immunoglobulin binding domain of streptococcal protein G) using the GōMartini 2 and GōMartini 3 methodologies, respectively.</p> <p>Each folder contains:</p> <p>A.pdb: All-atom PDB file of the unfolded structure.<br>B.pdb: All-atom PDB file of the folded structure.<br>A-CG.pdb: Martini coarse-grained representation of the unfolded structure.<br>B-CG.pdb: Martini coarse-grained representation of the folded structure.<br>A-GC.gro: GROMACS format file of the unfolded structure.<br>B-GC.gro: GROMACS format file of the folded structure.<br>system.top: Topology file.</p> <p><br>Force Fields and Topologies:<br>GōMartini 2:<br>martini_v2.2.itp: Martini 2 force field for proteins.<br>martiniv2.2_ions.itp: Martini 2 force field for ions.<br>Protein_A.itp: GōMartini 2 topology for the folded structure.<br>Protein_A_unfolded.itp: GōMartini 2 topology for the unfolded structure.</p> <p><br>GōMartini 3:<br>martini_v3.0.0.itp: Martini 3 force field for proteins.<br>martini_v3.0.0_ions_v1.itp: Martini 3 force field for ions.<br>martini_v3.0.0_solvents_v1.itp: Martini 3 force field for solvents.<br>go_martini.itp: GōMartini 3 topology for the folded structure.<br>go_martini_unfold.itp: GōMartini 3 topology for the unfolded structure.<br>go_molecule1.itp: Additional GōMartini 3 topology file.<br>A.map: Contact map for the unfolded structure.<br>B.map: Contact map for the folded structure.</p> <p>GROMACS Files for Equilibration and MD Runs:</p> <p>Equilibration:<br>minimization.mdp, minimization.gro: Files for energy minimization.<br>nvt.mdp, nvt.gro, nvt.xtc, nvt.tpr: Files for NVT equilibration.<br>npt.mdp, npt.gro, npt.xtc, npt.tpr, npt_dry.tpr: Files for NPT equilibration (dry simulations without waters or ions).</p> <p>Folding MD Runs:<br>dynamic.mdp: File for folding molecular dynamics runs.<br>Folding MD Runs (1-100):<br>MD/hairpin-M2_fit.xtc*: Trajectories (without waters or ions, &epsilon; = 9.414 kJ/mol).<br>MD/hairpin-M2-15_fit*.xtc*: Trajectories (without waters or ions, &epsilon; = 15.0 kJ/mol).</p> <p>&nbsp;</p> <p>MD/hairpin-CG-M3-_fit*.xtc: Trajectories (without waters or ions, &epsilon; = 15.0 kJ/mol).</p> <p>Data Analysis Scripts:<br>MD/Distances_native_WT.ipynb: Jupyter notebook for backbone distance analysis.<br>MD/all_native_graphs.ipynb: Jupyter notebook for native contact analysis.<br>MD/total_contacts_column.ipynb: Python script for contact statistics.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Review data for: SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping

<p>Climatology of snow water equivalent of Switzerland between winters 1962 and 2021. Obtained using quartile mapping between a model using data assimilation since 1998 and a model without data assimilation. This version of the dataset corresponds to the publication revision time. The publication has been submitted to GMD Copernicus journal as: <em>SnowQM 1.0: A fast R Package for bias-correcting spatial fields of snow water equivalent using quantile mapping</em></p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record