Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
648
datasets available to search
ShareScore release 0.7.1
Dataset results
648 results for “uncertainty”
Dataset and program scripts for the reproducibility of the hierarchical data structure file. Related to the manuscript entitled: Hierarchical Representation of Measurement Data, Metrological Uncertainty and Metadata for Calibrated Battery Tests
<p>We present an interoperable hierarchical data representation for battery tests, leading to improved scalability of data transmission and enhanced data accessibility and comprehensibility for both human interpretation and machine processing. The hierarchical data format includes the raw trace electrical measurement data, the metrological calibration and uncertainty data, the metadata such as experimental settings, instruments and software versions, as well as post-processed data such as electrochemical model fit parameters. This data representation allows repetition of the battery test under the exact same conditions such that identical results are achieved within defined error bounds. This is in line with the general F.A.I.R. data approach and provides repeatability and traceability in the battery value chain. As an application of the hierarchical data representation, we show the classification of cells as pass/fail being performed with quantitative confidence levels. We demonstrate the complete workflow of establishing the hierarchical data structure for electrochemical impedance spectroscopy (EIS), starting from metrological traceability of the calibration and uncertainty analysis towards the storage of the structured data as a single integrated file that preserves the hierarchical data format.</p>
Uncertainty Analysis of Digital Elevation Models by Spatial Inference From Stable Terrain – Dataset
<p><strong>Dataset of <a href="https://doi.org/10.1109/jstars.2022.3188922">Hugonnet et al. (2022), Uncertainty Analysis of Digital Elevation Models by Spatial Inference From Stable Terrain</a>.</strong></p> <p>The data is composed of:</p> <ul> <li><strong>For the Mont-Blanc case study: </strong>the Pléiades reference DEM, the SPOT-6 DEM, the Pléiades–SPOT-6 elevation difference, and the forest mask generated from the ESA CCI landcover (delainey polygonization);</li> <li><strong>For the Northern Patagonian Icefield case study: </strong>the ASTER reference DEM, the SPOT-5 DEM, the ASTER–SPOT-5 elevation difference, and the quality of stereo-correlation of the ASTER DEM from MicMac.</li> </ul> <p>The filenames correspond to those used in the <strong>associated GitHub repository</strong>: <a href="https://github.com/rhugonnet/dem_error_study">https://github.com/rhugonnet/dem_error_study</a>. The shapefiles used for masking glaciers are available directly from the <strong>Randolph Glacier Inventory 6.0</strong> at <a href="https://www.glims.org/RGI/">https://www.glims.org/RGI/</a>.</p> <p>The date of the DEMs is in their original format: <strong>year-month-day for all but ASTER</strong> that has the original naming of <a href="https://lpdaac.usgs.gov/products/ast_l1av003/">AST L1A products</a>. <strong>Units are meters</strong> for the DEMs and elevation differences, <strong>and percentages</strong> for the quality of stereo-correlation.</p>
Uncertainty in Migration Scenarios. QuantMig Project Deliverable D9.2 Data Description
<p>This open data deposit contains the data and code accompanying used in the report: Barker and Bijak (2021), Uncertainty in Migration Scenarios, QuantMig Project Deliverable D9.2. The cover note should be read in conjunction with the report, available via www.quantmig.eu, and with the individual readme files in the data folders that can be found within this Zenodo repository (DOI: 10.5281/zenodo.7709443).</p>
Quantifying Both Socioeconomic and Climate Uncertainty in Coupled Human-Earth Systems Analysis
<p>This data repository is associated with the paper:</p> <p>Morris,J., A. Sokolov, J. Reilly, A. Libardoni, C. Forest, S. Paltsev, A Schlosser, R. Prinn and H. Jacoby (2025). Quantifying Both Socioeconomic and Climate Uncertainty in Coupled Human-Earth Systems Analysis. <em>Nature Communications </em><strong>16</strong>, 2703. https://doi.org/10.1038/s41467-025-57897-1</p> <p>This paper quantifies key socio-economic and climate uncertainties using the MIT Integrated Global System Model. </p>
Centre frequencies and uncertainties for "Evidence for a kilometre-scale seismically slow layer atop the core-mantle boundary from normal modes"
<p>A table containing the centre frequencies and uncertainties used for the study presented in "Evidence for a kilometre-scale seismically slow layer atop the core-mantle boundary from normal modes". This table is the same as is contained in the supplementary materials of that paper.</p> <p>Russell, S., Irving, J. C. E., Jagt, L., & Cottaar, S. (2023). Evidence for a kilometer-scale seismically slow layer atop the core-mantle boundary from normal modes. Geophysical Research Letters, 50, e2023GL105684. <a href="https://doi.org/10.1029/2023GL105684">https://doi.org/10.1029/2023GL105684</a></p>
Model simulation data used in "Exploring the uncertainties in the aviation soot-cirrus effect" (Righi et al., Atmos. Chem. Phys., 2021)
<p>This dataset contains the output of the EMAC global model simulations analysed and discussed in Righi et al. (<i>Atmos. Chem. Phys.</i>, 2021). For details see the README.md file and Table 1 in the paper.</p>
Data from systematic audit for paper: Insights into the quantification and reporting of model-related uncertainty across different disciplines
<p>This upload contains 7 data files (each contains cleaned and compiled data for a given scientific field) and 2 R scripts. These files support the paper: Insights into the quantification and reporting of model-related uncertainty across different disciplines.</p> <p> </p> <p><strong>Description of the data</strong></p> <p>Compiled data files for each field contain all reviewers audit answers for eligible papers. All papers that met exclusion criteria have been removed.</p> <p>Data checks have been performed and formatting errors corrected either in R or manually, following steps detailed in the STAR methods.</p> <p>Column names and description:</p> <ul> <li>Number: number of question from 1 to 9</li> <li>Questions: question text – question to be answered by the reviewer</li> <li>QuestionCode: shortened code for each question</li> <li>Paper: paper code - first author surname/initial and surname and year</li> <li>Initials: initials of reviewer</li> <li>Answer: answer to the question</li> <li>Details: extra details to support the answer</li> <li>Location: where in the text the uncertainty was presented</li> <li>Presentation: how the uncertainty was presented</li> <li>ModelType: type of model (focal model)</li> <li>Comments: any other comments from the reviewer</li> <li>Checks: checks of whether NA or no have been included in correct places e.g. if answers to questions 1:4 are no then question 9 is NA, if question 7 is no then 8 is NA</li> <li>Check 1 = when Answer = No, Location is NA</li> <li>Check 2 = when Answer to Number 1-4, 6 or 8-9 is Yes that Details are not NA</li> <li>Check 3 = when Answer = No, Presentation = NA</li> <li>Check 4 = when Location is not NA, presentation is not NA</li> <li>Check 5 = if the Answer to 5 or 7 is "No" then Answer to 6 and 8 = "NA"</li> <li>Check 6 = if Answer for 1-4 is "No", then Answer for 9 = "NA"</li> </ul> <p><strong>Code description</strong></p> <p>Two scripts are included, the first is theme_script.R, this includes code to set up a ggplot theme for the figures. The second is Figure_code.R, this script contains all code to plot and save the three figures from the paper.</p>
Pan-European exposure maps and uncertainty estimates from HANZE v2.0 model, 1870-2020
<p>This dataset provides all output data generated in the standard settings of HANZE v2.0 model. The 100-m pan-European maps (GeoTIFF) provide gridded totals of five variables for years 1870-2020 for 42 countries. The rasters are group in five ZIP files:</p> <p>- CLC: land cover/use (Corine Land Cover classification; legend files are included in a separate ZIP)</p> <p>- Pop: population</p> <p>- GDP: gross domestic product (2020 euros)</p> <p>- FA: fixed asset value (2020 euros)</p> <p>- imp: imperviousness density (%)</p> <p>Two additional CSV files contain uncertainty estimates of population, GDP and fixed asset value per NUTS3 region and flood hazard zone. The files provide 5th, 20th, 50th, 80th and 95th percentile for all timesteps, separately for coastal and riverine floods.</p> <p>Two further Excel files contain subnational and national-level statistical data on population, land use and economic variables.</p> <p>For detailed description of the files, see the documentation provided with the code.</p> <p>This version replaces the airport list, which was previously incorrectly taken from HANZE v1, and adds land cover/use legend files for ArcGIS and QGIS.</p>
MEWL: Few-shot multimodal word learning with referential uncertainty
<p><strong>Dataset Release for <a href="https://arxiv.org/abs/2306.00503">MEWL: Few-shot multimodal word learning with referential uncertainty (ICML 2023) </a></strong></p> <p><strong>GitHub:</strong> <a href="https://github.com/jianggy/MEWL">https://github.com/jianggy/MEWL</a></p> <p><strong>Abstract: </strong>Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human's core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children's core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents' performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines.</p>
Dataset for Multidisciplinary Uncertainty Mining - ver1
<p>This dataset contains sentences extracted from articles in various disciplines and annotated with respect to uncertainty in science. It has been produced as part of the <a href="https://project-inscim.github.io/">ANR InSciM (Modelling Uncertainty in Science) project</a>. </p> <p>The dataset is drawn from reputable scientific articles from a variety of disciplines. It consists of two distinct samples of sentences, each annotated using a different method. The first sample is obtained through uncertainty cue mapping, while the second sample is derived from manual annotation of randomly selected articles. To ensure comprehensive annotation, both samples were manually annotated using our multidimensional annotation framework.</p> <p>For a more comprehensive understanding of the construction of the dataset, including the selection of journals, sampling procedure, and the annotation methodology, see (Ningrum and Atanassova, 2023).</p> <p>This dataset provides valuable insights into the representation of uncertainty within scientific literature across different domains. Researchers and practitioners can utilize this dataset to study and analyze the different dimensions of uncertainty in scientific discourse.</p> <p>The dataset is presented as a CSV table where colons ( are used as delimiters. The columns of the table are as follows :</p> <ul> <li>source : 'db' or 'manual' referring to the method used to identify and extract the sentence;</li> <li>article_id : internal id of the article from which the sentence was extracted;</li> <li>sen_id : internal unique id of the sentence;</li> <li>cue : uncertainty cue present in the sentence;</li> <li>text : sentence text;</li> <li>journal_id : short name of the journal;</li> <li>check : 'Y' if the sentence expresses uncertainty and 'N' otherwise;</li> <li>ref, nature, context, timeline, expression : annotations of the type of uncertainty according to the annotation framework proposed by (Ningrum and Atanassova, 2023).</li> </ul> <p>It is essential to highlight the presence of duplicate data in the dataset. These duplicates arise from the detection of multiple cues in sentences during the cue mapping procedure. While one might consider omitting these duplicates, we deliberately chose to retain them. This decision allows for a more comprehensive understanding of how the cues manifest within the sentences. By analyzing the duplicate instances, we can gain valuable insights into the various ways in which the cues are expressed.</p> <p> </p> <p><strong>Bibliography</strong></p> <p>Ningrum, P. K., Atanassova, I. (2023) "Scientific Uncertainty: an Annotation Framework and Corpus Study in Different Disciplines" In 19th International Conference of the International Society for Scientometrics and Informetrics (ISSI 2023), Bloomington, Indiana, US.</p>
Data for paper "Magnetohydrodynamic Equilibrium Reconstruction with Consistent Uncertainties"
<p>Data and scripts for the conference paper "Magnetohydrodynamic Equilibrium Reconstruction with Consistent Uncertainties" for the 42nd International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering.</p> <p><strong>Abstract</strong>: We report on progress towards a probabilistic framework for consistent uncertainty quantification and propagation in analysis and numerical modeling of physics in magnetically confined plasmas in the stellarator configuration. A frequent starting point in this process is the calculation of a magnetohydrodynamic equilibrium from plasma profiles. Profiles and therefore the equilibrium are typically reconstructed from experimental data. What sets equilibrium reconstruction apart from usual inverse problems is that profiles are given as functions over a magnetic flux derived from the magnetic field, rather than spatial coordinates. This makes it a fixed-point problem that is traditionally left inconsistent or solved iteratively in a least-squares sense[1–3]. The aim here is towards a straightforward and transparent process to quantify and propagate uncertainties and their correlations for function-valued fields and profiles in this setting. We propose a framework that utilizes a low dimensional prior distribution of equilibria, constructed with principal component analysis. A surrogate of the forward model[4] is trained to enable faster sampling.</p> <p><strong>Funding</strong>: The present contribution is supported by the Helmholtz Association of German Research Centers under the joint research school HIDSS-0006 'Munich School for Data Science - MUDS'. This work has been carried out within the framework of the EUROfusion Consortium, funded by the European Union via the Euratom Research and Training Programme (Grant Agreement No 101052200 - EUROfusion). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Commission. Neither the European Union nor the European Commission can be held responsible for them.</p>
Soil organic carbon and associated uncertainty at 90 m resolution for peninsular Spain
Soil organic carbon (SOC) must be quantified and monitored to assess soil management practices, adapt policies, and evaluate environmental impacts. However, due to SOC spatial variability, soil surveys become a very challenging task because of the high costs of acquiring data, operational complexity, and updating. Digital soil mapping based on machine learning approaches in combination with remote sensing techniques have enabled soil carbon spatial distribution to be significantly improved, even with limited soil samples. A legacy soil database of 8,361 georeferenced profiles and a selection of environmental data-driven covariates intimately related to soil-forming factors (e.g., biota, climate, parent material) were used to generate SOC maps. Modeling of data was based on three supervised learning approaches: quantile regression forest, ensemble machine learning and auto-machine learning. For the final SOC spatial distribution maps, each pixel was assigned the prediction from the most accurate model, i.e., lowest uncertainty. We applied this modeling technique to generate cost-effective, high-resolution maps (90 m pixel resolution) of SOC distribution, and its associated spatially explicit uncertainty, in peninsular Spain. These maps showed 15.7 g.kg-1 mean SOC concentration at 0-30 cm and 3.6 g.kg-1 at 30-100 cm depth. The total SOC stock at its effective depth was 3.8 Pg C, storing the 74% in the upper 30 cm (2.82 Pg C). The correlation between SOC observed and predictions final values showed R2=0.68 for SOCc and R2=0.54 for SOCs at the upper 30cm. The methodology proposed in this study aims to improve benchmark SOC estimates in support of the National GHG Emissions Inventory Report
Data used to create figures in the ACP Letters manuscipt "The value of remote marine aerosol measurements for constraining radiative forcing uncertainty" by Regayre et al. (2020)
<p>This dataset was created from perturbed parameter ensembles (PPEs) using the HadGEM-UKCA atmospheric composition climate model. All data needed to reproduce figures in the Regayre et al. (2020) ACP Letters article "The value of remote marine aerosol measurements for constraining radiative forcing uncertainty" are included. Other output from the PPEs can be obtained by contacting the lead author.</p> <p>The following data are included here:</p> <ul> <li>CCN measurement data degraded to match the model-measurement comparison resolution.</li> <li>Unconstrained and constrained CCN<sub>0.2</sub> output from the PPE used to make Figure 1. These compressed files contain 48 .dat files. Each .dat file contains the PPE mean, variance and 95% creidble interval data. Files are named consecutively, containing data from 90<sup>o</sup>S to 90<sup>o</sup>N at 0<sup>o</sup>E, then continuing Eastward. When combined, these files provide data for each latitude/longitude pair at the N48 spatial resolution.</li> <li>A zip file of an netcdf file containing 26-dimensional data for parameter values, used to create the sample of 1 million model variants from our statistical emulators of model output.</li> <li>A zip file containing a folder of files made of one million ones and zeros that indicate the retention/rejection criteria from applying our constraint methodology for various constraint combination scenarios, for each model variant. A value of 1 indicates the model variant was retained. Data in these files is in the same order as the unconstrained sample file of parameter values.</li> <li>Compressed files containing global, annual mean RF<sub>aci</sub> and ERF<sub>aci</sub> values for the unconstrained set of one million model variants. The compressed netcdf files contain RF (ERF), RF<sub>aci</sub> (ERF<sub>aci</sub>) and RF<sub>ari</sub> (ERF<sub>ari</sub>) values.</li> </ul>
Historical uncertainty in Gregory of Tours's History of the Franks (book 7)
<p>Our goal was to create a research dataset based on geographical and chronological uncertainties in the work of Gregory of Tours's *History of the Franks* (book 7). We used and modified a topology of geographical and chronological uncertainty based on a rudimentary schema that would be universal when analysing an historical source :</p> <p>Chronological :<br> * uncertain dating<br> * uncertain method of dating<br> * lack of dating<br> * precise dating</p> <p>Geographical: <br> * uncertain location<br> * general location (region, country)<br> * lack of location <br> * precise location</p> <p>After working on book 7 for a while, that schema was reworked as those 9 types of uncertainty : </p> <p>Chronological :<br> * uncertain_dating<br> * uncertain_method_dating<br> * event_dating_null<br> * precise_dating</p> <p>Geographical: <br> * uncertain_location<br> * general_location <br> * event_location_null<br> * uncertain_method_location<br> * precise_location</p> <p><br> The geographical and chronological focus makes it possible to identify where and when, in a source, the historical uncertainty is higher. </p> <p>Using python, that dataset was then automatically cleaned and enhanced with bounding box based on geo-mapping information for the entries of geographical uncertainty. Those were classified as either precise_location or general_location. </p> <p>For example, anything relating to a city general area (like "in the Rouen area") creates a general_location bounding box encompassing the *current* geographical space occupied by the municipality of Rouen (in the format 'LongMin', 'LongMax', 'LatMin', 'LatMax' in a single column "bbox"). Anything described as a unique point in space (like "in Paris") creates a precise_location and its corresponding lat/long system of coordinates. </p> <p>This is an arbitrary way to translate slightly undefined geographical concepts of uncertainty into formal data, but at least it can be fully explained explicitly.<br> </p> <p>Translation used: Tours G. <em>et alii</em>, <em>The history of the franks</em>, Penguin Books Limited, 1974, <a href="https://books.google.ch/books?id=4Lx-M2RHGgoC">https://books.google.ch/books?id=4Lx-M2RHGgoC</a>.</p>
L4A - Biomass map of the Brazilian Amazon and uncertainty
<p>The AGB final map (further referred as EBA - Estimativa de Biomassa para a Amazônia - map) presented a maximum AGB value of 518 Mg ha-1, a mean AGB of 174 Mg ha-1, and a standard deviation of 102 Mg ha-1. The map is provided in TIF format, projected using EPSG 4236. The uncertainty map is provided in TIF format, projected using EPSG 4236. The information is offered in Mg ha-1.</p>
QPE uncertainty assessment over Iowa
<p>This repository has several Hydrologic discharge simulation runs made using two QPE products (MRMS and IFC) altered by a multiplicative factor oscillating between 0.1 and 5. The simulations are recorded at the USGS gauges with observations in Iowa between 2015 and 2022. Additionally, we included the performance metrics and the code to generate the analysis developed for the manuscript titled: "Assessing the Impact of Radar-Rainfall Uncertainty on Streamflow Prediction."</p> <p>The data on this repo contains three compressed files:</p> <ul> <li><strong>hlm_runs.rar</strong>: Holds the hydrologic model runs using both QPEs and the multiplicative factor.</li> <li><strong>performance.rar: </strong>It has a summary of the performance metrics<strong> </strong>and a folder with the raw performance at each gauge.</li> <li><strong>rainfall.rar: </strong>Has the rainfall computed at each watershed using COOP gauges (folder), as well as the results for MRMS and IFCA?</li> </ul> <p>All the data was stored in <strong>gzip </strong>compressed <strong>parquet </strong>format using <strong>Pandas</strong>. </p>
GrainLearning: A Bayesian uncertainty quantification toolbox for discrete and continuum numerical models of granular materials
GrainLearning is a Bayesian uncertainty quantification and propagation toolbox for computer simulations of granular materials. The software is primarily used to infer and quantify parameter uncertainties in computational models of granular materials from observation data, also known as inverse analyses or data assimilation. Implemented in Python, GrainLearning can be loaded into a Python environment to process the simulation and observation data, or alternatively, as an independent tool where simulation runs are done separately, e.g., via a shell script.
Supplementary data for "Effect of Uncertainty in Water Vapor Continuum Absorption on CO2 Forcing, Longwave Feedback, and Climate Sensitivity"
<h3>This dataset is supplementary to the article "Effect of Uncertainty in Water Vapor Continuum Absorption on CO2 Forcing, Longwave Feedback, and Climate Sensitivity".</h3> <h3>spectral_olr.nc</h3> <p>This file contains the spectral outgoing longwave radiation (OLR) calculated using the line-by-line radiative transfer model ARTS and the radiative-convective equilibrium model konrad. It contains spectral OLR for surface temperatures from 270K to 330K for different strengths of the water vapor continuum absorption.</p> <h3>opacity_emission_level.py</h3> <p>This file also contains the spectrally resolved optical depth and the emission level of outgoing longwave radiation for the considered absorption species (H2O lines, H2O continuum, H2O self continuum, H2O foreign continuum, CO2, N2, and O2).</p> <h3>continuum_reference_conditions.nc</h3> <p>This file contains the reference continuum absorption coefficients that were used to calculate the adjustment to the foreign continuum for the single-constraint experiment.</p> <h3>continuum_all_profiles.nc</h3> <p>This file contains the reference continuum absorption coefficients that were used to calculate the adjustment to the foreign continuum for the general-constraint experiment.</p> <h3>modified_continuum_input_files_single_constraint.zip and modified_continuum_input_files_general_constraint.zip</h3> <p>These files contain the modified continuum data files used for the implementation of the MT_CKD continuum model in the line-by-line model ARTS for the single-constraint and general-constraint experiments, respectively.</p> <h3>tau_column.nc and tau_profile.nc</h3> <p>These files contain separately for each absorption species the vertically integrated opacity spectra, and the opacity profiles at two selected wavenumbers.</p> <p> </p>
Annotated Dataset for Uncertainty Mining : Gold Standard
<p> </p> <h1>Description of the dataset</h1> <p>In order to study the expression of uncertainty in scientific articles, we have put together an interdisciplinary corpus of journals in the fields of Science, Technology and Medicine (STM) and the Humanities and Social Sciences (SHS). The selection of journals in our corpus is based on the Scimago Journal and Country Rank (SJR) classification, which is based on Scopus, the largest academic database available online. We have selected journals covering various disciplines, such as medicine, biochemistry, genetics and molecular biology, computer science, social sciences, environmental sciences, psychology, arts and humanities. For each discipline, we selected the five highest-ranked journals. In addition, we have included the journals PLoS ONE and Nature, both of which are interdisciplinary and highly ranked.</p> <p>Based on the corpus of articles from different disciplines described above, we created a set of annotated sentences as follows:</p> <ul> <li>593 were pre-selected automatically, by studying the occurrences of the lists of uncertainty indices proposed by Bongelli et al. (2019), Chen et al. (2018) and Hyland (1996).</li> <li>The remaining sentences were extracted from a subset of articles, consisting of two randomly selected articles per journal. These articles were examined by two human annotators to identify sentences containing uncertainty and to annotate them.</li> <li>600 sentences not expressing scientific uncertainty were manually identified and reviewed by two annotators<br><br></li> </ul> <p>The sentences were annotated by two independent annotators following the annotation guide proposed by Ningrum and Atanassova (2024). The annotators were trained on the basis of an annotation guide and previously annotated sentences in order to guarantee the consistency of the annotations. <br>Each sentence was annotated as expressing or not expressing uncertainty (<strong>Uncertainty</strong> and <strong>No Uncertainty)</strong>.<br>Sentences expressing uncertainty were then annotated along five dimensions: Reference , Nature, Context , Timeline and Expression. <br>The annotators reached an average agreement score of 0.414 according to Cohen's Kappa test, which shows the difficulty of the task of annotating scientific uncertainty.<br>Finally, conflicting annotations were resolved by a third independent annotator.</p> <p><br>Our final corpus thus consists of a total of 1 840 sentences from 496 articles in 21 English-language journals from 8 different disciplines.<br>The columns of the table are as follows:</p> <ol> <li><strong>journal</strong>: name of the journal from where the article originates</li> <li><strong>article_title</strong>: title of the article from where the sentence is extracted</li> <li><strong>publication_year</strong>: year of publication of the article</li> <li><strong>sentence_text</strong>: text of the sentence expressing or not expressing uncertainty</li> <li><strong>uncertainty</strong>: 1 if the sentence expresses uncertainty and 0 otherwise;</li> <li><strong>ref, nature, context, timeline, expression</strong>: annotations of the type of uncertainty according to the annotation framework proposed by Ningrum and Atanassova (2023). The annotation of each dimension in this dataset are in numeric format rather than textual. The mapping betwen textual and numeric labels is presented in the Table below.</li> </ol> <table> <tbody> <tr> <td>Dimension</td> <td>1</td> <td>2</td> <td>3</td> <td>4</td> <td>5</td> </tr> <tr> <td>Reference</td> <td>Author</td> <td>Former</td> <td>Both</td> <td> </td> <td> </td> </tr> <tr> <td>Nature</td> <td>Epistemic</td> <td>Aleatory</td> <td>Both</td> <td> </td> <td> </td> </tr> <tr> <td>Context</td> <td>Background</td> <td>Methods</td> <td>Res&Disc</td> <td>Conclusion</td> <td>Others</td> </tr> <tr> <td>Timeline</td> <td>Past</td> <td>Present</td> <td>Future</td> <td> </td> <td> </td> </tr> <tr> <td>Expression</td> <td>Quantified</td> <td>Unquantified</td> <td> </td> <td> </td> <td> </td> </tr> </tbody> </table> <p><br>This gold standard has been produced as part of the <a href="https://project-inscim.github.io/">ANR InSciM (Modelling Uncertainty in Science) project.</a> </p> <h1>References</h1> <p><br>Bongelli, R., Riccioni, I., Burro, R., & Zuczkowski, A. (2019). Writers’ uncertainty in scientific and popular biomedical articles. A comparative analysis of the British Medical Journal and Discover Magazine [Publisher: Public Library of Science]. PLoS ONE, 14 (9). <a href="https://doi.org/10.1371/journal.pone.0221933">https://doi.org/10.1371/journal.pone.0221933</a></p> <p>Chen, C., Song, M., & Heo, G. E. (2018). A scalable and adaptive method for finding semantically equivalent cue words of uncertainty. Journal of Informetrics, 12 (1), 158–180. <a href="https://doi.org/10.1016/j.joi.2017.12.004">https://doi.org/10.1016/j.joi.2017.12.004</a></p> <p><br>Hyland, K. E. (1996). Talking to the academy forms of hedging in science research articles [Publisher: SAGE Publications Inc.]. Written Communication, 13 (2), 251–281. <a href="https://doi.org/10.1177/0741088396013002004">https://doi.org/10.1177/0741088396013002004</a></p> <p>Ningrum, P. K., & Atanassova, I. (2023). Scientific Uncertainty: An Annotation Framework and Corpus Study in Different Disciplines. 19th International Conference of the International Society for Scientometrics and Informetrics (ISSI 2023). <a href="https://doi.org/10.5281/zenodo.8306035">https://doi.org/10.5281/zenodo.8306035</a></p> <p>Ningrum, P. K., & Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. Scientometrics. <a href="https://doi.org/10.1007/s11192-024-05009-z">https://doi.org/10.1007/s11192-024-05009-z</a></p>
Supplement to "Proof of concept for Bayesian inference of dynamic rating curve uncertainty" (v3)
<div>This deposit contains part of the updated supplement to “Proof of concept for Bayesian inference of dynamic rating curve uncertainty” (<a href="https://www.tandfonline.com/doi/full/10.1080/02626667.2024.2401094" target="_blank" rel="noopener">Cornelio et al. 2024, HSJ</a>). This version, in particular, contains two files in which the following changes were made from the earlier version (v2.0.1):</div> <div> <ul> <li><strong><em>250117_Lbn_RC_new.R</em></strong> is the updated R code. The argument for the random number generator (RNG) kind is defined for the set.seed() functions used in the script. </li> <li><strong><em>Lbn-DMs-csv0.csv</em></strong> is the updated input file containing the stage-discharge gaugings. The column for the stage values has been renamed to "H_rec" (instead of "H_m" as in the original CSV) to be consistent with the attribute name used throughout the R code.</li> </ul> <p>Except for the above files, all the input and output files in <a href="https://zenodo.org/records/12792513" target="_blank" rel="noopener">v2.0.1</a><span> remain unchanged. </span></p> </div> <p><u> </u></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.