Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.9.0
Dataset results
50 results for “Uncertainty Estimation”
Pan-European exposure maps and uncertainty estimates from HANZE v2.0 model, 1870-2020
<p>This dataset provides all output data generated in the standard settings of HANZE v2.0 model. The 100-m pan-European maps (GeoTIFF) provide gridded totals of five variables for years 1870-2020 for 42 countries. The rasters are group in five ZIP files:</p> <p>- CLC: land cover/use (Corine Land Cover classification; legend files are included in a separate ZIP)</p> <p>- Pop: population</p> <p>- GDP: gross domestic product (2020 euros)</p> <p>- FA: fixed asset value (2020 euros)</p> <p>- imp: imperviousness density (%)</p> <p>Two additional CSV files contain uncertainty estimates of population, GDP and fixed asset value per NUTS3 region and flood hazard zone. The files provide 5th, 20th, 50th, 80th and 95th percentile for all timesteps, separately for coastal and riverine floods.</p> <p>Two further Excel files contain subnational and national-level statistical data on population, land use and economic variables.</p> <p>For detailed description of the files, see the documentation provided with the code.</p> <p>This version replaces the airport list, which was previously incorrectly taken from HANZE v1, and adds land cover/use legend files for ArcGIS and QGIS.</p>
Photometric Redshifts for Cosmology: Improving accuracy and uncertainty estimates using Bayesian Neural Networks
<p><strong>This data consists of 286,401 with broad-band g,r,i,z,y photometry from the HSC DR2 survey and spectroscopic redshifts. The majority of galaxies in our sample lies between redshift of 0.01 and 2.5</strong></p>
Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"
<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories. </p>
Supplemental material to the paper: Incorporating epistemic uncertainty in the assessment of an existing masonry building through a Point Estimate Method
<p>This repository contains the meterials used to study numerically the effect of model uncertainties in the seismic assessment of a case-study stone masonry building. The simulations were run in the research version of the software <a href="http://www.tremuri.com/">Tremuri</a>. The repository contains input files, model results and the matlab code to reproduce the figures presented in the paper:</p> <blockquote> <p>Vanin F., Beyer K., "Incorporating epistemic uncertainty in the assessment of an existing masonry building through a Point Estimate Method", submitted for publication (2019)</p> </blockquote>
Winter Precipitation-Type Models for "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications"
<p>This contains trained model weights, scalers, and evaluation metrics for the winter precipitation-type models trained as part of the paper "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications". </p>
Reliable imputation of spatial transcriptome with uncertainty estimation and spatial regularization
<p>Imputation of missing features in spatial transcriptomics is urgently demanded due to technology limitations, while most existing computational methods suffer from moderate accuracy and cannot estimate the reliability of the imputation. <br> To fill the research gaps, we introduce a computational model, TransImp, that imputes the missing feature modality in spatial transcriptomics by mapping it from single-cell reference. Uniquely, we derived a set of attributes that can accurately predict imputation uncertainty, hence enabling us to select reliably imputed genes. Also, we introduced a spatial auto-correlation metric as a regularization to avoid overestimating spatial patterns. Multiple datasets from various platforms have demonstrated that our approach significantly improves the reliability of downstream analyses in detecting spatial variable genes and interacting ligand-receptor pairs. Therefore, TransImp offers a way towards a reliable spatial analysis of missing features for both matched and unseen modalities, e.g., nascent RNAs.</p>
A database for benchmarking organ dose estimates and uncertainties in CT
<p>This database includes patient images and associated verified Monte Carlo based estimates of organ doses that may be used for benchmarking different organ dose estimation techniques against a reference standard.</p>
Uncertainties associated with microwave link rainfall estimates in an urban environment
<p>These dataset were collected from a dedicated microwave link setup between Mt View Reservoir (T) and 33 Lakeside Burwood (R). There were two OTT1 disdrometers installed at both ends of the microwave link complemented by 3 tipping bucket rain gauges. </p>
Taxonomic Uncertainty on Range Size and Niche Estimation in a Southern Ocean Cryptic Species Complex
<p>R code and data related to assessing the effect of taxonomic uncertainty on range size and environmental niche estimates for a Southern Ocean invertebrate. </p> <p>Clarke, D.A., Wilson, N.G. and McGeoch, M.A. (2025) ‘Effects of Taxonomic Uncertainty on Range Size and Niche Estimation in a Southern Ocean Cryptic Species Complex’, Journal of Biogeography, n/a(n/a), p. e15182. Available at: https://doi.org/10.1111/jbi.15182.</p>
Estimation of Unfactorizable Systematic Uncertainties
<p>Dataset for the "Efficient Estimation of Unfactorizable Systematic Uncertainties" paper. </p> <p>The dataset consists of 30,000 simulated three-jet events. </p> <p>Keys and datasets:</p> <ul> <li>j1_threeM: Dataset of shape (30000, 3) containing the three-momenta of the hardest jets in the events.</li> <li>j2_threeM: Dataset of shape (30000, 3) containing the three-momenta of the second hardest jets in the events.</li> <li>j3_threeM: Dataset of shape (30000, 3) containing the three-momenta of the softest jets in the events.</li> </ul>
Uncertainty sensitivity estimates for EXIOBASE 3.8.2 footprints
<p>This data set provides uncertainty sensitivity estimates for EXIOBASE 3.8.2 footprints. Sensitivity levels of footprints of all EXIOBASE regions/extensions were calculated using Linear Error propagation (assuming a 0.1 relative standard deviation) such as that each of the 49 EXIOBASE regions has 124 footprint sensitivity estimates and a total of 6076 sensitivity estimates. This dataset was calculated as part of the paper "<em>Uncertainty propagation in EE-MRIO footprint estimates</em>" Badr & Stadler (2024).</p> <p>We reccoment interpreting footprint sensititivity levels as: 0-0.02: low sensitivity, 0.02-0.03: medium sensitivity, 0.03-0.04: Highly sensitive, 0.04 or more: Extremely sensitive. </p> <p>The paper repository can be found on: gitlab.com/hitea/variance-in-uncertainties-in-mrios</p>
In-situ observations of nitrate loss factor for "Estimation method is the primary source of uncertainty in cropland nitrate leaching estimate in China"
<p>This database includes In-situ observations of nitrate loss factor. Details can be found in paper named "Estimation method is the primary source of uncertainty in cropland nitrate leaching estimate in China".</p>
Optimal policy for uncertainty estimation concurrent with decision making
<p>Dataset for "Optimal policy for uncertainty estimation concurrent with decision making".</p> <p>The dataset should be merged into the code folder, thus the program can work.</p>
Datasets used in "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications"
<p>The precipitation type (p-type) dataset (ptype.parquet) comprises observational weather reports sourced from the Meteorological Phenomena Identification Near the Ground (mPING) project, combined with corresponding numerical weather prediction data from the NOAA Rapid Refresh (RAP) model. These crowd-sourced mPING reports offer precipitation type labels (rain, snow, sleet, and freezing rain) across North America, while the RAP model provides atmospheric data, including temperature, humidity, and wind profiles, on pressure levels.</p> <p> </p> <p>The RAP data covers the contiguous United States (CONUS) from 2015 to 2022 on an hourly 13km grid. The mPING observations are matched to the nearest RAP grid cell and hour, allowing the two data sources to be merged into a labeled dataset suitable for classification tasks. </p> <p> </p> <p>The surface layer flux dataset (surface_layer.csv) contains high-frequency meteorological observations spanning from 2013 to 2015, collected at the Cabauw Experimental Site in the Netherlands. It includes measurements of various variables such as temperature, humidity, wind, radiation, and soil moisture, recorded every 10 minutes. The target output encompasses friction velocity, sensible heat, and latent heat.</p> <p><br> The code used for processing the datasets and training neural network models is available in the Miles-Guess repository (<a href="https://github.com/ai2es/miles-guess">https://github.com/ai2es/miles-guess</a>).</p>
Spatial uncertainty in herbarium data: Simulated displacement but not error distance alters estimates of phenological sensitivity to climate in a widespread California wildflower
Open the record for dataset details and reuse information.
Estimating drivers and identifying uncertainties in smallmouth bass population dynamics in an invaded river network
Open the record for dataset details and reuse information.
Data from: Natural language processing systems for pathology parsing in limited data environments with uncertainty estimation
Open the record for dataset details and reuse information.
Data from: Use of hidden Markov capture-recapture models to estimate abundance in presence of uncertainty: application to estimating the prevalence of hybrids in animal populations
Estimating the relative abundance (prevalence) of different population segments is a key step in addressing fundamental research questions in ecology, evolution, and conservation. The raw percentage of individuals in the sample (naive prevalence) is generally used for this purpose, but it is likely to be subject to two main sources of bias. First, the detectability of individuals is ignored; second, classification errors may occur due to some inherent limits of the diagnostic methods. We developed a hidden Markov (also known as multievent) capture–recapture model to estimate prevalence in free‐ranging populations accounting for imperfect detectability and uncertainty in individual's classification. We carried out a simulation study to compare naive and model‐based estimates of prevalence and assess the performance of our model under different sampling scenarios. We then illustrate our method with a real‐world case study of estimating the prevalence of wolf (Canis lupus) and dog (Canis lupus familiaris) hybrids in a wolf population in northern Italy. We showed that the prevalence of hybrids could be estimated while accounting for both detectability and classification uncertainty. Model‐based prevalence consistently had better performance than naive prevalence in the presence of differential detectability and assignment probability and was unbiased for sampling scenarios with high detectability. We also showed that ignoring detectability and uncertainty in the wolf case study would lead to underestimating the prevalence of hybrids. Our results underline the importance of a model‐based approach to obtain unbiased estimates of prevalence of different population segments. Our model can be adapted to any taxa, and it can be used to estimate absolute abundance and prevalence in a variety of cases involving imperfect detection and uncertainty in classification of individuals (e.g., sex ratio, proportion of breeders, and prevalence of infected individuals).
Data from: Accounting for uncertainty in gene tree estimation: summary-coalescent species tree inference in a challenging radiation of Australian lizards
Accurate gene tree inference is an important aspect of species tree estimation in a summary-coalescent framework. Yet, in empirical studies, inferred gene trees differ in accuracy due to stochastic variation in phylogenetic signal between targeted loci. Empiricists should, therefore, examine the consistency of species tree inference, while accounting for the observed heterogeneity in gene tree resolution of phylogenomic data sets. Here, we assess the impact of gene tree estimation error on summary-coalescent species tree inference by screening ${\sim}2000$ exonic loci based on gene tree resolution prior to phylogenetic inference. We focus on a phylogenetically challenging radiation of Australian lizards (genus Cryptoblepharus, Scincidae) and explore effects on topology and support. We identify a well-supported topology based on all loci and find that a relatively small number of high-resolution gene trees can be sufficient to converge on the same topology. Adding gene trees with decreasing resolution produced a generally consistent topology, and increased confidence for specific bipartitions that were poorly supported when using a small number of informative loci. This corroborates coalescent-based simulation studies that have highlighted the need for a large number of loci to confidently resolve challenging relationships and refutes the notion that low-resolution gene trees introduce phylogenetic noise. Further, our study also highlights the value of quantifying changes in nodal support across locus subsets of increasing size (but decreasing gene tree resolution). Such detailed analyses can reveal anomalous fluctuations in support at some nodes, suggesting the possibility of model violation. By characterizing the heterogeneity in phylogenetic signal among loci, we can account for uncertainty in gene tree estimation and assess its effect on the consistency of the species tree estimate. We suggest that the evaluation of gene tree resolution should be incorporated in the analysis of empirical phylogenomic data sets. This will ultimately increase our confidence in species tree estimation using summary-coalescent methods and enable us to exploit genomic data for phylogenetic inference.
Data from: Estimating correlated rates of trait evolution with uncertainty
Correlated evolution among traits, which can happen due to genetic constraints, ontogeny, and selection, can have an important impact on the trajectory of phenotypic evolution. For example, shifts in the pattern of evolutionary integration may allow the exploration of novel regions of the morphospace by lineages. Here we use phylogenetic trees to study the pace of evolution of several traits and their pattern of evolutionary correlation across clades and over time. We use regimes mapped to the branches of the phylogeny to test for shifts in evolutionary integration while incorporating the uncertainty related to trait evolution and ancestral regimes with joint estimation of all parameters of the model using Bayesian Markov chain Monte Carlo. We implemented the use of summary statistics to test for regime shifts based on a series of attributes of the model that can be directly relevant to biological hypotheses. In addition, we extend Felsenstein's pruning algorithm to the case of multivariate Brownian motion models with multiple rate regimes. We performed extensive simulations to explore the performance of the method under a series of scenarios. Finally, we provide two test cases; the evolution of a novel buccal morphology in fishes of the family Centrarchidae and a shift in the trajectory of evolution of traits during the radiation of anole lizards to and from the Caribbean islands.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.