Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

935

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

935 results for “probability”

Learn how ShareScore rates datasets ↗
edi60/100

Detection Probability of Red Wood Ants in Friedenweiler, Germany 2015

Estimation of population sizes and species ranges is central to population and conservation biology. It is widely appreciated that imperfect detection of mobile animals must be accounted for when estimating population size from presence-absence data. Sessile organisms also are imperfectly detected, but correction for detection probability in estimating their population sizes is rare. We illustrate challenges of detection probability and population estimation of sessile organisms using censuses of red wood ant (Formica rufa-group) nests as a case study. These ants, widespread in the northern hemisphere, can make large (up to 2m tall), highly visible nests. Using data from a two-day mapping campaign by eight individuals of 147 ant nests spread across sixteen 3600-m2 plots in the Black Forest region of southwest Germany, we developed a Bayesian model for quantifying detection probability of sessile organisms. Detection probabilities by individual observers of red wood ant nests ranged from 0.31 – 0.56, and depended on experience of the observers, size and density of nests, and habitat characteristics. Robust estimation of population density of sessile organisms—even highly apparent ones such as red wood ant nests—requires unbiased estimation of detection probability, just as it does when estimating population density of rare or cryptic species.

openCC0Dec 2023View details →
zenodo52/100

Predicted occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution

<p>The dataset contains predictions of occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution + all covariate layers used for modeling. Over seven million electronic health records (EHRs), among which 11,741 EHRs reported tick attachment, were used to evaluate climate, environmental and animal host factors affecting the risk of tick attachment in cats and dogs in Great Britain (GB). The tick presence/absence EHRs for dogs and cats were further overlaid with spatiotemporal time-series of climatic, vegetation, human influence, hydrological and terrain variables (slope, wetness index) to produce a spatiotemporal regression matrix; an Ensemble Machine Learning framework was used to fine-tune hyperparameters for Random Forest (classif.ranger), Gradient boosting (classif.xgboost) and GLM-net (classif.glmnet) algorithms, which were then used to produce a final ensemble meta-learner that predicts the probability of occurrence of ticks across GB with monthly intervals.</p> <ul> <li>gb1km_covariates.zip contains ALL covariate layers as GeoTIFFs (time-series) used for modeling ticks dynamics;</li> <li>data_1km_2014_M01.rds = contains all covariates for January 2014 prepared as SpatialGridDataFrame (R data object);</li> </ul> <p>Codes of files indicate e.g.:</p> <ul> <li>&quot;monthly.tick.prob_savsnet.mar_p_1km_s_2014_2021&quot; = monthly occurrence probability for January based on the training data from 2014 to 2021;</li> <li>&quot;monthly.tick.prob_savsnet.oct_md_1km_s_20211001_20211031&quot; = monthly prediction (model) error derived as the standard deviation from multiple base learners;</li> </ul> <p>The dataset is described in detail in the following publication:</p> <ul> <li>Arsevska, E., Hengl, T., Singelton, D. et al. (2023?) <strong>Risk factors for tick attachment in companion animals in Great Britain: a spatiotemporal analysis covering 2014&ndash;2021</strong>. Submitted to Parasites &amp; Vectors (in review).</li> </ul> <p>The model summary shows:</p> <pre><code>Call: stats::glm(formula = f, family = "binomial", data = getTaskData(.task, .subset), weights = .weights, model = FALSE) Deviance Residuals: Min 1Q Median 3Q Max -1.4749 -0.0557 -0.0471 -0.0430 3.7611 Coefficients: Estimate Std. Error z value Pr(&gt;|z|) (Intercept) -7.64495 0.02095 -364.957 &lt; 2e-16 *** classif.ranger 4.95061 0.63615 7.782 7.13e-15 *** classif.xgboost 189.75543 5.53109 34.307 &lt; 2e-16 *** classif.glmnet 140.24208 5.05375 27.750 &lt; 2e-16 *** --- Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1 (Dispersion parameter for binomial family taken to be 1) Null deviance: 170604 on 7303013 degrees of freedom Residual deviance: 162571 on 7303010 degrees of freedom AIC: 162579 Number of Fisher Scoring iterations: 9</code></pre> <p><em>Acknowledgements</em>: We are grateful to data providers in veterinary practice (VetSolutions, Teleos, CVS, and other practitioners). We are grateful to the INRAE MIGALE bioinformatics facility (MIGALE, INRAE, 2020. Migale Bioinformatics Facility, doi: <a href="https://entrepot.recherche.data.gouv.fr/dataverse/migale">10.15454/1.5572390655343293E12</a>) for providing computing resources. We are also grateful for<br> the help and support provided by <a href="https://www.liverpool.ac.uk/savsnet/">SAVSNET team members</a> Bethaney Brant, Susan Bolan and Steven Smyth.<br> This study was funded mainly by a grant from the <strong>Biotechnology and Biological Sciences Research Council</strong>,<br> BB/NO19547/1 and <strong>British Small Animal Veterinary Association</strong> (BSAVA). The research was partly funded by the National Institute for <strong>Health Research Health Protection Research Unit</strong> (NIHR HPRU) in Emerging and Zoonotic Infections at the <strong>University of Liverpool</strong> in partnership with <strong>Public Health England</strong> (PHE) and <strong>Liverpool School of Tropical Medicine</strong> (LSTM). This work has been partially funded by the <em>&ldquo;Monitoring outbreak events for disease surveillance in a data science context&quot;</em> (MOOD) project from the European Union&rsquo;s Horizon 2020 research and innovation program under grant agreement No. 874850 (<a href="https://mood-h2020.eu/">https://mood-h2020.eu/</a>). The views expressed are those of the authors and not necessarily those of the NHS, the NIHR, the Department of Health or Public Health England.</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Iterative Mapping of Probabilities

<p>This repository contains data and scripts for implementing the Iterative Mapping of Probabilities (IMP) algorithm proposed in the preprint submitted to the International Journal of Applied Earth Observation and Geoinformation (JAG). The framework aims to improve the accuracy of land cover mapping by iteratively refining class maps to match independent area statistics. The experiment focuses on generating classification maps for five countries (Belgium, Czechia, Germany, Luxembourg, Netherlands) based on input probability rasters.</p> <h2>Usage</h2> <ol> <li><strong>Create a project folder</strong> where you'll store the files.</li> <li><strong>Download all the files</strong> to the project folder.</li> <li><strong>Extract the countries data</strong> into the project folder (be.zip=Belgium, cz=Czechia, de.zip=Germany, lu.zip=Luxembourg, nl.zip=Netherlands).<br>(Note: ensure the folder structure matches the "Data Description" section provided below)&nbsp;</li> <li><strong>Install Dependencies</strong> by navigating to the project folder in your terminal and install the necessary dependencies by running:<br>(Note: make sure you have Python installed on your system)<br><code>pip install -r ./requirements.txt</code></li> <li><strong>Run the script</strong> using the following command in the terminal:<br><code>python ./main.py</code></li> </ol> <h2>Data Description</h2> <p>After downloading and decompressing the files, the data must have the following structure.</p> <ul> <li><strong>area_estimates.csv:&nbsp;</strong>The area estimates for each land cover provided by Eurostat.</li> <li><strong>[country_code]/</strong><br> <ul> <li><strong>[model]/:</strong><br> <ul> <li><strong>classified_highest_likelihood/: </strong>Contains classification maps generated using the maximum likelihood mapping algorithm.</li> <li><strong>classified_proportional/: </strong>Stores classification maps produced using the Iterative Mapping of Probabilities algorithm.</li> <li><strong>iterations/: </strong>Stores images representing the iteration number in which each pixel was classified using the iterative proportional algorithm.</li> <li><strong>probabilities/: </strong>Contains input probability rasters for both mapping algorithms.</li> </ul> </li> </ul> </li> <li><strong>main.py</strong>: Python script implementing the Iterative Mapping of Probabilities framework.</li> <li><strong>requirements.txt</strong>: List of required libraries to run the script.</li> <li><strong>graphical_abstracl.pdf (optional)</strong>: Illustration on the IMP algorithm.</li> </ul> <h2>Script Explanation:</h2> <p>The script <strong>main.py</strong> implements IMP algorithm and process land use and land cover classification maps from probability rasters. These probabilities were generated for different countries and used two diffrent models (local and general). Please, refer to the paper for more details on how these models were trained.</p> <h3>Script Steps:</h3> <ol> <li><strong>Data Preparation:</strong><br>Loads <code>area_estimate.csv</code> file containing area estimates for different land cover classes in various countries and years.</li> <li><strong>Parameter Setup:</strong><br>Sets up parameters for each country, year, and model combination.<br>Each parameter set includes the country code, year, model type (local or general), and a list of land cover class codes.<br>(Note: you can change this section to set up parametersto process just some countries)</li> <li><strong>Processing Maps:</strong><br>Iterates over each parameter set and:<br> <ol> <li>Loads reference proportions of land cover classes for the specified country, year, and model.</li> <li>Loads probabilities from raster images.</li> <li>Runs the Iterative Mapping of Probabilities algorithm using the loaded probabilities and reference proportions.</li> <li>Saves the resulting land use and land cover classification map as an output.</li> </ol> </li> </ol> <h3>Script Inputs:</h3> <p>- CSV file containing area estimates for land cover classes (<code>./area_estimates.csv</code>).<br>- Probability raster images generated by classification models stored in <code>./[country_code]/[model]/probabilities/</code> folders.</p> <h3>Script Outputs:</h3> <p>Land use and land cover classification maps obtained by running the Iterative Mapping of Probabilities algorithm for each parameter set. The outputs are stored in <code>./[country_code]/[model]/classified_proportional_user/</code> folders.</p> <h2>Citation</h2> <p>If you use this code or data in your research, please cite the corresponding paper:</p> <p><em>Witjes, M., Herold, M., &amp; de Bruin, S. (2024). Iterative Mapping of Probabilities: A data fusion framework for generating accurate land cover maps that match area statistics. Journal of Applied Earth Observation and Geoinformation (JAG), in review.</em></p> <h2>License</h2> <p>The code in this repository is licensed under the MIT License.</p> <p>The data provided in this repository is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Probability of wildfire containment

<p>Raster layer depicting the probatility of containing a fire according to the landscape configuration: accessibility, aerial means, relief complexity and vegetation density.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Predicted USDA soil orders at 250 m (probabilities)

<p>Distribution of the USDA orders (12) based on machine learning predictions of great groups (<a href="https://doi.org/10.5281/zenodo.1476844">https://doi.org/10.5281/zenodo.1476844</a>) from global compilation of soil profiles. To learn more about soil orders and great groups please refer to the&nbsp;<a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail&nbsp;<strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil/tree/master/soil">here</a></strong>. Antartica is not included.</p> <p>To access and visualize maps use:&nbsp;<a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code:&nbsp;<a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a>&nbsp;</li> <li>General questions and comments:&nbsp;<a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using &quot;COMPRESS=DEFLATE&quot; creation&nbsp;option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>order = variable: USDA order,</li> <li>usda.histosols = determination method: USDA soil taxonomy class Histosols,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>

opencc-by-sa-4.0May 2019View details →
zenodo48/100

Predicted USDA soil suborders at 250 m (probabilities)

<p>Distribution of the USDA suborders&nbsp;based on machine learning predictions of great groups (<a href="https://doi.org/10.5281/zenodo.1476844">https://doi.org/10.5281/zenodo.1476844</a>) from global compilation of soil profiles. To learn more about soil suborders and great groups please refer to the&nbsp;<a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail&nbsp;<strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil">here</a></strong>. Antartica is not included.</p> <p>To access and visualize maps use:&nbsp;&nbsp;<a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code:&nbsp;<a href="https://gitlab.com/openlandmap/global-layers/-/issues">https://gitlab.com/openlandmap/global-layers/-/issues</a>&nbsp;</li> <li>General questions and comments:&nbsp;<a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using &quot;COMPRESS=DEFLATE&quot; creation&nbsp;option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>suborder = variable: USDA suborder,</li> <li>usda.ustolls = determination method: USDA soil taxonomy class Ustolls,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.1 = version number: 0.1,</li> </ul>

opencc-by-sa-4.0May 2019View details →
zenodo48/100

Predicted USDA soil great groups at 250 m (probabilities)

<p>Distribution of the USDA soil great groups based on machine learning&nbsp;predictions from global compilation of soil profiles (&gt;350,000 training points). To learn more about soil great groups please refer to the <a href="https://www.nrcs.usda.gov/wps/PA_NRCSConsumption/download/?cid=stelprdb1247203.pdf">Illustrated Guide to Soil Taxonomy - NRCS - USDA</a>. Processing steps are described in detail <strong><a href="https://gitlab.com/openlandmap/global-layers/tree/master/soil">here</a></strong>. Antarctica is not included.</p> <p>To access and visualize maps use:&nbsp;<a href="http://www.openlandmap.org/">OpenLandMap.org</a></p> <p>A back-up copy of all predictions (&gt;65GB) can be downloaded from: <a href="http://gofile.me/6J25n/mQ3cHOOMr">http://gofile.me/6J25n/mQ3cHOOMr</a></p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code:&nbsp;<a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a>&nbsp;</li> <li>General questions and comments:&nbsp;<a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using &quot;COMPRESS=DEFLATE&quot; creation&nbsp;option in GDAL. File naming convention:</p> <ul> <li>sol = theme: soil,</li> <li>grtgroup = variable: USDA great group,</li> <li>usda.argiustolls = determination method: USDA soil taxonomy class Argiustolls,</li> <li>p = probability,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: soil surface,</li> <li>1950..2017 = time reference: period 1950-2017,</li> <li>v0.2&nbsp;= version number: 0.2,</li> </ul>

opencc-by-sa-4.0Nov 2018View details →
zenodo48/100

Experimental determination of the gadolinium L subshells fluorescence yields and Coster-Kronig transition probabilities

<p>Data associated with the two main tables of the publication &quot;Experimental determination of the gadolinium L subshells fluorescence yields and Coster-Kronig transition probabilities&quot;. There are two .txt files containing tabulator-separated&nbsp;values:</p> <p><em>tabl01_ck.txt: </em>Data associated with Table 1 of the publication. This file contains the L subshell Coster-Kronig (CK) factors and their respective uncertainties.</p> <p><em>tabl02_fy.txt: </em>Data associated with Table 2 of the publication. This file contains the experimentally determined Gd L subshell fluorescence yields in comparison to available literature sources and their respective uncertainties.</p> <p>For more details see the original Open Access publication:</p> <p>Kayser, Y.,&nbsp;H&ouml;nicke, P.,&nbsp;Wansleben, M.,&nbsp;W&auml;hlisch, A.,&nbsp;Beckhoff, B.,&nbsp;<em>X-Ray Spectrom</em>&nbsp;2022,&nbsp;1.&nbsp;<a href="https://doi.org/10.1002/xrs.3313">https://doi.org/10.1002/xrs.3313</a></p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Prevalent trends in realized probability of occurrence of main European forest tree species for 2000–2020

<p>High resolution maps resulting from a trend analysis conducted for the period 2000&ndash;2020 on the probability of occurrence maps prepared by <a href="https://doi.org/10.7717/peerj.13728">Bonannella et al. (2022)</a>. For this analysis we selected the realized distribution time series layers at 30m spatial resolution for 6 out of 16 species described in the mentioned publication:</p> <ul> <li>Silver fir (<em>Abies alba </em>Mill.)</li> <li>European beech (<em>Fagus sylvatica </em>L.)</li> <li>Norway spruce (<em>Picea abies </em>L.)</li> <li>Black pine (<em>Pinus nigra </em>J. F. Arnold)</li> <li>Scots pine (<em>Pinus sylvestris </em>L.)</li> <li>Common oak (<em>Quercus robur </em>L.)</li> </ul> <p>The trend analysis was conducted per pixel on each of these species individually. We fitted simple OLS regression models with the probability of occurrence as the dependent variable and time as the independent variable. After the model fitting, we also calculated the t-test statistics to determine the presence of an increasing (positive) or decreasing (negative) trend or no trend at all.</p> <p>By combining the regression slope coefficient (<em>&beta;</em>) and the <em>p</em>-value from the t-test statistics we assigned each pixel to one of three classes:</p> <ul> <li><em>positive</em>: <em>&beta;</em> &gt; 0.25 AND <em>p</em>-value &lt; 0.05</li> <li><em>negative</em>: <em>&beta;</em> &lt; &minus;0.25 AND <em>p</em>-value &lt; 0.05</li> <li><em>no trend / stable</em>: &minus;0.25 &le; <em>&beta;</em> &ge; 0.25 OR <em>p</em>-value &gt; 0.05</li> </ul> <p>We then aggregated the resulting classes at 1km resolution maps to capture the prevalent trend in probability of occurrence over a certain area. Files are named according to the following naming convention, e.g.:</p> <ul> <li>veg_abies.alba_slope_30m_0..0cm_epsg3035_v1.0</li> </ul> <p>with the following fields:</p> <ul> <li>theme: e.g. <strong>veg</strong>,</li> <li>species code: e.g. <strong>abies.alba</strong>,</li> <li>variable name: e.g. <strong>slope</strong>,</li> <li>resolution in meters e.g. <strong>30m</strong>,</li> <li>reference depths (vertical dimension): e.g. <strong>0..0cm</strong>,</li> <li>coordinate system: e.g. <strong>epsg3035</strong>,</li> <li>data set version: e.g. <strong>v1.0</strong>.</li> </ul> <p>For each species here we provide the following layers:</p> <ul> <li>veg_abies.alba_<strong>slope</strong>:<strong> </strong>slope coefficient (scaling factor: 10000)</li> <li>veg_abies.alba_<strong>pvalue</strong>:<strong> </strong><em>p</em>-value (scaling factor: 1000)</li> <li>veg_abies.alba_<strong>pos.trends_30m</strong>: pixels classified as <em>positive </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>pos.trends_1km</strong>: proportion of pixels of the <em>positive </em>class over a 1&times;1 km area (range 0&ndash;100)</li> <li>veg_abies.alba_<strong>neg.trends_30m</strong>: pixels classified as <em>negative </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>neg.trends_1km</strong>: proportion of pixels of the <em>negative </em>class over a 1&times;1 km area (range 0&ndash;100)</li> <li>veg_abies.alba_<strong>no.trends_30m</strong>: (pixels classified as <em>no trend / stable </em>on the original maps at 30m resolution (boolean layer with range 0&ndash;100, only the two extremes values are present)</li> <li>veg_abies.alba_<strong>no.trends_1km</strong>:<strong> </strong>proportion of pixels of the <em>no trend / stable </em>class over a 1&times;1 km area (range 0&ndash;100)</li> </ul> <p>Files are provided as GeoTIFFs and projected in the Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035). Styling files are provided in <em>QML</em> format</p> <p>A publication describing, in detail, all processing steps is currently in review. See at:<br> <br> Bonannella, C., Parente, L., de Bruin, S. and Herold, M. (2023). Multi-decadal trend analysis and forest disturbance assessment of European tree species: concerning signs of a subtle shift, PREPRINT (Version 1) available at Research Square [<a href="https://doi.org/10.21203/rs.3.rs-3288937/v1">https://doi.org/10.21203/rs.3.rs-3288937/v1</a>]</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

List of human genes and their probabilities of being intolerant to heterozygous protein truncating variants.

<p>Recalculation of the Supplementary Table 2 (doi:10.1371/journal.pcbi.1004647.s002) of the journal article "The Characteristics of Heterozygous Protein Truncating Variants in the Human Genome" by Bartha and Rausell published in PLoS Computational Biology (http://dx.doi.org/10.1371/journal.pcbi.1004647). Probabilities in this dataset were computed using human variation data from the Exome Aggregation Consortium (http://exac.broadinstitute.org/).</p> <p>Methods described in that article is relevant for this dataset. All author and affiliation information in that article is relevant for this dataset.</p> <p>Credit for the original human variation data is for the Exome Aggregation Consortium (http://exac.broadinstitute.org/, doi:10.1038/nature19057).</p>

opencc-by-4.0Jan 2017View details →
zenodo44/100

Data accompanying the manuscript "Assessing the Probability of Extremely Low Wind Energy Production in Europe at Sub-seasonal to Seasonal Time Scales"

<p>This dataset contains time series of wind energy production aggregated over France and Europe, obtained from a 1000-year climate simulation from the CESM model (version 1.2.2, Hurrel et al. 2013), coupled to a simple energy model to compute grid-point capacity factor from surface wind. Wind power is then computed by multiplying the capacity factor by the installed capacity, taken from 5 e-Highway scenarios (X5, X7, X10, X13 and X16), and integrated over the regions of interest. More details about the climate simulation, wind energy model and installed capacity scenarios can be found in the associated manuscript, "Assessing the Probability of Extremely Low Wind Energy Production in Europe at Sub-seasonal to Seasonal Time Scales" (Cozian et al. 2023).</p><p>The data is organized into 10 files for France and 10 files for Europe. In each case, the 10 files correspond to 10 batches of 100 years each, with 3-hourly output. Each file contains 5 time series corresponding to the 5 installed capacity scenarios.</p><h4>References</h4><ul><li>Hurrell J W, Holland M M, Gent P R, Ghan S, Kay J E, Kushner P J, Lamarque J F, Large W G, Lawrence D, Lindsay K, Lipscomb W H, Long M C, Mahowald N, Marsh D R, Neale R B, Rasch P, Vavrus S, Vertenstein M, Bader D, Collins W D, Hack J J, Kiehl J and Marshall S (2013). The community earth system model: A framework for collaborative research. Bulletin of the American Meteorological Society, 94, 1339–1360. <a href="https://doi.org/10.1175/BAMS-D-12-00121.1">https://doi.org/10.1175/BAMS-D-12-00121.1</a></li><li>e-Highway 2050 (2015). Europe's future secure and sustainable electricity infrastructure. <a href="https://docs.entsoe.eu/baltic-conf/bites/www.e-highway2050.eu/results">https://docs.entsoe.eu/baltic-conf/bites/www.e-highway2050.eu/results</a></li><li>Cozian B, Herbert C and Bouchet F (2023). Assessing the Probability of Extremely Low Wind Energy Production in Europe at Sub-seasonal to Seasonal Time Scales. <a href="https://doi.org/10.48550/arXiv.2311.13526">https://doi.org/10.48550/arXiv.2311.13526</a></li></ul>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Damage Absolute Probability map

<p>Damage Absolute Probability map is a layer in support to Area of Interest (AOI) definition for an earthquake event. It provides the spatial distribution, in the examined area, of the absolute values of probability of damage derived considering the most severe damage classes provided by loss assessment data</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Original single session datasets from "Slowly evolving dopaminergic activity modulates the moment-to-moment probability of reward-related self-timed movements."

<p>This archive contains the original&nbsp;single-session recording datasets associated with the paper &quot;Slowly evolving dopaminergic activity modulates the moment-to-moment probability of reward-related self-timed movements&quot; by Allison E Hamilos, Giulia Spedicato, Ye Hong, Fangmiao Sun, Yulong Li, and John A Assad (https://doi.org/10.1101/2020.05.13.094904). Files can be loaded and collated with code from our GitHub repository to reproduce all analyses (https://www.github.com/harvardschoolofmouse).</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Input data set for the statistical analsysis of rockfall reach probabilities

<p>These files contain reach probability values extracted from 3D rockfall simulations for field-mapped block deposits as well as a series of attributes characterising the deposits. They served for the statistical analysis of the reach probability values as a function of site, forest and rockfall characteristics. The results of the analysis are published in Dorren et al. 2022: Delimiting rockfall runout zones using reach probability values simulated with a Monte-Carlo based 3D trajectory model. Natural Hazards and Earth System Scienses.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

German Student Responses to Probability Theory and Statistics Bachelor Course (WuS24): Evaluated with Rubrics

<p><strong>Description:</strong></p> <p>This dataset contains questions and answers from an introductory computer science bachelor course on statistics and probability theory at Hochschule Bonn-Rhein-Sieg. The dataset includes three questions and a total of 90 answers, each evaluated using binary rubrics (yes/no) associated with specific scores.</p> <p>&nbsp;</p> <p><strong>Dataset Components:</strong></p> <ol> <li><em>questions.csv</em>: Contains the details of the three questions. <ul> <li>Columns: <ul> <li><em>question_id</em>: Unique identifier for each question</li> <li><em>question</em>: The text of the question</li> <li><em>solution</em>: The reference answer for the question</li> <li><em>max_score</em>: The maximum score for this question</li> </ul> </li> </ul> </li> <li><em>rubrics.csv</em>: Contains the grading rubrics for each question. <ul> <li>Columns: <ul> <li><em>question_id</em>: Unique identifier for each question</li> <li><em>rubric_id</em>: Unique identifier for each rubric within a question</li> <li><em>rubric</em>: The rubric phrased as a question</li> <li><em>score</em>: The score associated with fulfilling the rubric</li> </ul> </li> </ul> </li> <li><em>answers.csv</em>: Contains 90 student answers to the questions. <ul> <li>Columns: <ul> <li><em>answer_id</em>: Unique identifier for each answer</li> <li><em>question_id</em>: Unique identifier of the question that is answered</li> <li><em>answer</em>: The text of the student's answer</li> <li><em>score</em>: The score associated with fulfilling the rubric</li> </ul> </li> </ul> </li> <li><em>answer_rubrics.csv</em>: Contains the evaluations of rubrics for each answer.<br> <ul> <li>Columns: <ul> <li><em>answer_id</em>: The identifier of the answer.</li> <li><em>question_id</em>: The identifier of the question.</li> <li><em>rubric_id</em>: The identifier of the rubric for that question.</li> <li><em>label</em>: Indicates if the rubric crierion is fulfilled for the specific answer (true / false).</li> </ul> </li> </ul> </li> </ol> <pre><strong><br>Working with the Dataset:</strong> The easiest way to work with this dataset is using the class `RubricsDataset` defined in the file `dataloader.py`. Example:<br><br></pre> <pre><code>from dataloader import RubricsDataset<br><br>dataset = RubricsDataset.from_directory("data")<br> dataset.get_question(1) # Get a dictionary containing info about the first question, including rubrics dataset.get_answers(1) # Get all the answers for the first question as a list. Each answer is a dictionary with answer, score, rubrics.</code></pre>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling: Datasets.

<p>Datasets related to the publication [1].<br> Including:</p> <ul> <li>KRAS G12X mutations derived from COSMIC v.79 [http://cancer.sanger.ac.uk/cosmic/] (KRAS_G12X_mut_COSMICv79..xlsx)</li> <li>RMSFs (300-2000ns) of GDP-systems (300_2000rmsf_GDP_systems_RAW_AVG_SE.xlsx)</li> <li>RMSFs (300-2000ns) of GTP-systems (300_2000RMSF_GTP_systems_RAW_AVG_SE.xlsx)</li> <li>PyInteraph analysis data for salt-bridges and hydrophobic clusters (.dat files for each system in the PyInteraph_data.zip-file)</li> <li>Backbone&nbsp;trajectories for each system (residues 4-164; frames for every 1ns). Last number (e.g. _1) refers to the replica of the&nbsp;simulated system.</li> <li>backbone_4-164.gro/.pdb/.tpr -files (resid 4-164)&nbsp;&nbsp;</li> </ul> <p><br> [1] Pantsar T et al.&nbsp;Assessment of mutation probabilities of KRAS G12 missense mutants and their long-time scale dynamics by atomistic molecular simulations and Markov state modeling. <em>PLoS Comput Biol Submitted</em>&nbsp;(2018)</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Circular seal in a copper alloy engraved with a standing female figure, probably Lakṣmī; inscription at one side.

<p>Circular seal in a copper alloy engraved with a standing female figure, probably Lakṣmī; inscription at one side. British Museum 1897,0528.4.</p>

opencc-by-4.0Aug 2019View details →
zenodo44/100

R code and data for "Flake selection and scraper retouch probability: an alternative model for explaining Middle Paleolithic assemblage retouch variability"

<p>R code and data used for &quot;Flake selection and scraper retouch probability: an alternative model for explaining Middle Paleolithic assemblage retouch variability&quot; (Archaeological and Anthropological Sciences,&nbsp;Volume 10, Issue 7, pp 1791&ndash;1806)</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Sentinel-2: Cloud Probability in Earth Engine

<p>Links:</p> <ul> <li><a href="https://developers.google.com/earth-engine/datasets/catalog/COPERNICUS_S2_CLOUD_PROBABILITY">Sentinel-2: Cloud Probability</a>&nbsp;in Earth Engine&#39;s Public Data Catalog</li> <li><a href="https://radiantearth.github.io/stac-browser/#/external/storage.googleapis.com/earthengine-stac/catalog/COPERNICUS/COPERNICUS_S2_CLOUD_PROBABILITY.json">Sentinel-2: Cloud Probability</a>&nbsp;in Earth Engine STAC viewed with STAC Browser</li> </ul> <p>The S2 cloud probability is created with the&nbsp;<a href="https://github.com/sentinel-hub/sentinel2-cloud-detector">sentinel2-cloud-detector</a>&nbsp;library (using&nbsp;<a href="https://github.com/microsoft/LightGBM">LightGBM</a>). All bands are upsampled using bilinear interpolation to 10m resolution before the gradient boost base algorithm is applied. The resulting&nbsp;<code>0..1</code>&nbsp;floating point probability is scaled to&nbsp;<code>0..100</code>&nbsp;and stored as a UINT8. Areas missing any or all of the bands are masked out. Higher values are more likely to be clouds or highly reflective surfaces (e.g. roof tops or snow).</p> <p>Sentinel-2 is a wide-swath, high-resolution, multi-spectral imaging mission supporting Copernicus Land Monitoring studies, including the monitoring of vegetation, soil and water cover, as well as observation of inland waterways and coastal areas.</p> <p>The Level-2 data can be found in the collection&nbsp;<a href="https://radiantearth.github.io/stac-browser/COPERNICUS_S2_SR">COPERNICUS/S2_SR</a>. The Level-1B data can be found in the collection&nbsp;<a href="https://radiantearth.github.io/stac-browser/COPERNICUS_S2">COPERNICUS/S2</a>. Additional metadata is available on assets in those collections.</p> <p>See&nbsp;<a href="https://developers.google.com/earth-engine/tutorials/community/sentinel-2-s2cloudless">this tutorial</a>&nbsp;explaining how to apply the cloud mask.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Probability of Detection applied to X-ray inspection using numerical simulations

<p>In this work, we apply and adapt established Probability of Detection (POD) methods on inline inspection of aluminium cylinder heads using X-ray computed tomography. The CT simulation tool SimCT [4] is used to acquire virtual images of the specimens including artificial defects, which avoids the manufacturing of calibrated defects of known type (e.g., pore, inclusion, crack etc.), size and location. One of the exemplary defects is discussed as representative result together with the generated POD curves as well as its characteristics (i.e., the minimum detected defect, the maximum missed defect, POD(a90) =0.90 and a90/95).</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record