Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

815

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

815 results for “Forecasting”

Learn how ShareScore rates datasets ↗
zenodo36/100

Supporting Data for pyCSEP: A Software Toolkit for Earthquake Forecast Developers

<p>Contains data needed to reproduce the figures from the publication of pyCSEP: A Software Toolkit for Earthquake Forecast Developers.</p> <p><br> &nbsp;&nbsp;&nbsp; evaluation_catalog.json<br> &nbsp;&nbsp;&nbsp; evaluation_catalog_zechar2013_merge.txt<br> &nbsp;&nbsp;&nbsp; SRL_2018031_esupp_Table_S1.txt<br> <br> &nbsp;&nbsp;&nbsp; bird_liu.neokinema-fromXML.dat<br> &nbsp;&nbsp;&nbsp; ebel.aftershock.corrected-fromXML.dat<br> &nbsp;&nbsp;&nbsp; helmstetter_et_al.hkj.aftershock-fromXML.dat<br> &nbsp;&nbsp;&nbsp; lombardi.DBM.italy.5yr.2010-01-01.dat<br> &nbsp;&nbsp;&nbsp; meletti.MPS04.italy.5yr.2010-01-01.dat<br> &nbsp;&nbsp;&nbsp; werner.HiResSmoSeis-m1.italy.5yr.2010-01-01.dat<br> <br> &nbsp;&nbsp;&nbsp; config.json<br> &nbsp;&nbsp;&nbsp; m71_event.json<br> &nbsp;&nbsp;&nbsp; results_complete.bin</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

History of daily forecast of cumulative COVID-19 mortality in multiple geographic entities across the world

<p>Forecasts are available from 2020-04-01 to 2021-10-20 for dozens to more than 200 geographic entities (GE) across the world (from 46 GE on 2020-04-01 to 246 GE on 2021-10-20).&nbsp;</p> <p>Each forecast of cumulative mortality is grounded on a probabilistic mixture of mortality trajectories of ahead-of-time geographic entities playing the role of real-life predictors eventually complemented by a parametric model based on a SIR representation. The methodology is presented in Soubeyrand, Ribaud et al. (2020, https://doi.org/10.1371/journal.pone.0238410) and Soubeyrand, Demongeot et al. (2020, https://doi.org/10.1016/j.onehlt.2020.100187).&nbsp;</p> <p>The forecast are daily implemented by a web application entitled &quot;COVID-19 Visualization&quot; available at https://shiny.biosp.inrae.fr/app_direct/mapCovid19/</p> <p>The original code is available here: https://gitlab.paca.inrae.fr/biosp/shinyMapCovid19</p> <p>Raw data for drawing the forecast are provided by the Johns Hopkins University Center for Systems Science and Engineering (JHU CSSE; https://systems.jhu.edu/) available at https://github.com/CSSEGISandData/COVID-19/ (Dong et al., 2020, https://doi.org/10.1016/S1473-3099(20)30120-1).&nbsp;</p> <p>Information about the data set and the code are provided in the&nbsp;readme.txt file.</p> <p>Data are provided in the forecast_data.rds file produced originally with the saveRDS() function of the R Statistical Software (https://cran.r-project.org/).&nbsp;</p> <p>A code for loading the data set and extracting some data corresponding to specific dates and geographic entities with the R Statistical Software is provided in the read_data.R file.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Unified Model Atmospheric Forecast Model Data for Machine Learning Cloud-Base Height

<p>Unified Model data, in pp format, for machine learning of cloud-base height based on profiles of temperature, humidity, pressure and cloud fraction. The model configuration is Global Atmosphere 6, running with a resolution of N320 (which is coarser than what was running operationally at the time). Each simulation is run for 24 hours, re-initialising every 24 hours. A separate data file is provided every 6 hours. Data points are on a latitude-longitude grid in the horizontal and on a stretched grid in the vertical. See https://gmd.copernicus.org/articles/10/1487/2017/ for details of the model configuration.</p> <p>Data from January 2016 is for training.</p> <p>Data from July 2017 is for development/validation</p> <p>Data from October 2017 is for final testing.</p> <p>&nbsp;</p>

openogl-uk-3.0Jul 2021View details →
zenodo36/100

Daily NOAA Global Ensemble Forecasting System forecasts for six National Ecological Observatory Network lakes (2021--05-18 to 2021-10-24)

<p>NOAA Global Ensemble Forecasting System output generated at 00 UTC that has been subsetted and temporally downscaled from 6-hr to 1-hr for&nbsp;six lakes in the National Ecological Observatory Network. &nbsp;The files include all ensembles and the set of variables required to run the General Lake Model. The NEON siteID for the&nbsp;lakes are BARC, SUGG, CRAM, LIRO, PRLA, PRPO. &nbsp;See&nbsp;https://www.neonscience.org for more information about each lake.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Data of manuscript "Forecasting day-ahead 1-minute irradiance variability from Numerical Weather Predictions" submitted to Solar Energy

<p>This is the data corresponding to manuscript &quot;Forecasting day-ahead 1-minute irradiance variability from Numerical Weather Predictions&quot; by Kreuwel et al., 2022, submitted to Solar Energy.</p> <p>&nbsp;</p> <p>The file `basic_stats.tar.gz` contains a broad set of standard statistics of surface meteorology and vertical profiles. The file `sw_flux_dn_xy.tar.gz` contains spatial cross sections of downwelling shortwave radiation.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Modeled data related to the article "Diffusive Wave Models for Operational Forecasting of Channel Routing at Continental Scale"

<p>The data holder contains modeled data on water level and discharge from some test cases.</p>

opencc-by-4.0Apr 2022View details →
dryad36/100

Forecasting suppression of invasive Sea Lamprey in Lake Superior: data and code for Bayesian forecast model

<p>Resource managers frequently are tasked with mitigating or reversing adverse effects of invasive species through management policies and actions.  In Lake Superior, of the Laurentian Great Lakes, invasive sea lamprey populations are suppressed to protect valuable fish stocks.  However, the relationship between choice of long-term control strategy and the future chance of achieving the suppression target is unclear.</p> <p>Using a 60+ year time-series of suppression effort and monitoring data from 50 assessment sites located on Lake Superior tributaries, we developed a Bayesian state-space model to forecast the probability of suppressing lamprey below the suppression target.</p> <p>With annual application of lampricide (i.e., lamprey-specific pesticide) at historical mean levels, we forecasted a 15% chance of achieving the Lake Superior sea lamprey suppression target in 2040.</p> <p>Increasing lampricide effort and/or supplementing lampricide control with age-1 recruitment reduction increased suppression chance.  Annual application of the maximum historical lampricide effort resulted in a 50% predicted chance of achieving the target, annual application of the mean historic lampricide effort plus a 40% reduction in recruitment resulted in a 54% chance, and the maximum amount of effort considered (maximum historic lampricide and 60% reduction in recruitment) resulted in a 94% chance.</p> <p><em><a>Policy </a>implications</em>. <a>We</a> developed a simulation model from a robust, long-term monitoring dataset that improves understanding of why long-term sea lamprey suppression objectives have been difficult to achieve in Lake Superior.  Furthermore, the model provides a means to gauge efficacy of sea lamprey control policy and action scenarios based on forecasted chance of achieving the suppression target. Creating processes for iteratively refining our forecasting model with stakeholder and technical-expert input and integration with a decision analysis framework could strengthen the link between ecological knowledge obtained from long-term monitoring and invasive sea lamprey management.</p>

opencc-zeroMay 2022View details →
dryad36/100

Practical guide to using Kendall's τ in the context of forecasting critical transitions

<p>Recent studies demonstrate that trends in indicators extracted from measured time series can indicate approaching to an impending transition. Kendall's τ coefficient is often used to study the trend of statistics related to the critical slowing down phenomenon and other methods to forecast critical transitions. Because statistics are estimated from time series, the values of Kendall'sτare affected by parameters such as window size, sample rate and length of the time series, resulting in challenges and uncertainties in interpreting results. In this study, we examine the effects of different parameters on the distribution of the trend obtained from Kendall's τ, and provide insights into how to choose these parameters. We also suggest the use of the non-parametric Mann-Kendall test to evaluate the significance of a Kendall'sτvalue. The non-parametric test is computationally much faster compared to the traditional parametric ARMA test.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Interpretable Deep Learning for Probabilistic MJO Prediction: CNN Forecasts

<p>This repository contains data produced for the paper &quot;Interpretable Deep Learning for Probabilistic MJO Prediction&quot; by A. Delaunay and H. M. Christensen (2021).</p> <p>&gt;&gt; mu_ens_XX.pt<br> contains the mean forecasts from each ensemble member at a lead time of XX&nbsp;days</p> <p>&gt;&gt; cov_alea_XX.pt<br> contains the aleatoric predictions of each ensemble member at a lead time of XX&nbsp;days</p>

opengpl-2.0-or-laterJul 2022View details →
dryad36/100

Using deep convolutional neural networks to forecast spatial patterns of Amazonian deforestation: supporting data and outputs

<p class="MsoNormal"><strong>1.    </strong>Tropical forests are subject to diverse deforestation pressures while their conservation is essential to achieve global climate goals. Predicting the location of deforestation is challenging due to the complexity of the natural and human systems involved but accurate and timely forecasts could enable effective planning and on-the-ground enforcement practices to curb deforestation rates. New computer vision technologies based on deep learning can be applied to the increasing volume of Earth observation data to generate novel insights and make predictions with unprecedented accuracy.</p> <p class="MsoNormal"><strong>2.    </strong>Here, we demonstrate the ability of deep convolutional neural networks (CNNs) to learn spatiotemporal patterns of deforestation from a limited set of freely available global data layers, including multispectral satellite imagery, the Hansen maps of annual forest change (2001-2020) and the ALOS PALSAR digital surface model, to forecast deforestation (2021). We designed four model architectures, based on 2D CNNs, 3D CNNs, and Convolutional Long Short-Term Memory (ConvLSTM) Recurrent Neural Networks (RNNs), to produce spatial maps that indicate the risk to each forested pixel (~30 m) in the landscape of becoming deforested within the next year. They were trained and tested on data from two ~80,000 km<sup>2</sup> tropical forest regions in the Southern Peruvian Amazon.</p> <p class="MsoNormal"><strong>3.</strong><strong>    </strong><span>The networks could predict the location of future forest loss to a high degree of accuracy (F</span><sub>1 </sub><span>= 0.58-0.71). Our best performing model (3D CNN) had the highest pixel-wise accuracy (F</span><sub>1 </sub><span>= 0.71) when validated on 2020 forest loss (2014-2019 training). Visual interpretation of the mapped forecasts indicated that the network could automatically discern the drivers of forest loss from the input data. For example, pixels around new access routes (e.g. roads) were assigned high risk whereas this was not the case for recent, concentrated natural loss events (e.g. remote landslides).</span></p> <p class="MsoNormal"><strong>4.</strong><strong>    </strong>CNNs can harness limited time-series data to predict near-future deforestation patterns, an important step in harnessing the growing volume of satellite remote sensing data to curb global deforestation. The modelling framework can be readily applied to any tropical forest location and used by governments and conservation organisations to prevent deforestation and plan protected areas.</p>

opencc-zeroJul 2022View details →
zenodo36/100

A successful short-term volcanic eruption forecasting using seismic features: datasets and Sotware

<p>Successful Short-Term Volcanic Eruption Forecasting Using Seismic Features, Suplementary Material</p> <p>by Rey-Devesa (1,2), Ben&iacute;tez (3), Prudencio, Ligdamis Guti&eacute;rrez (1,2), Cort&eacute;s (1,2), Titos (3), Koulakov (4,5), Zuccarello (6) and Ib&aacute;&ntilde;ez (1,2).</p> <p><br> Institutions associated:</p> <p>(1) Department of Theoretical Physics and Cosmos. Science Faculty. Avd. Fuentenueva s/n. University of Granada. 18071. Granada. Spain.</p> <p>(2) Andalusian Institute of Geophysiscs. Campus de Cartuja. University of Granada. C/Profesor Clavera 12. 18071. Granada. Spain.</p> <p>(3) Department of Signal Theory, Telematics and Communication. University of Granada. Informatics and Telecommunication School. 18071. Granada. Spain.</p> <p>(4) Trofimuk Institute of Petroleum Geology and Geophysics SB RAS, Prospekt Koptyuga, 3, 630090 Novosibirsk, Russia</p> <p>(5) Institute of the Earth&rsquo;s Crust SB RAS, Lermontova 128, Irkutsk, Russia</p> <p>(6) Istituto Nazionale di Geofisica e Vulcanologia, Sezione di Pisa (INGV-Pisa), via Cesare Battisti, 53, 56125, Pisa, Italy.</p> <p><br> Acknowledgment:</p> <p>This study was partially supported by the Spanish FEMALE project (PID2019-106260GB-I00).<br> P. Rey-Devesa was funded by the Ministerio de Ciencia e Innovaci&oacute;n del Gobierno de Espa&ntilde;a (MCIN),<br> Agencia Estatal de Investigaci&oacute;n (AEI), Fondo Social Europeo (FSE),<br> and Programa Estatal de Promoci&oacute;n del Talento y su Empleabilidad en I+D+I Ayudas para contratos predoctorales para la formaci&oacute;n de doctores 2020 (PRE2020-092719).<br> Ivan Koulakov was supported by the Russian Science Foundation (Grant No. 20-17-00075).<br> Luciano Zuccarello was supported by the INGV Pianeta Dinamico 2021 Tema 8 SOME project (grant no. CUP D53J1900017001)<br> funded by the Italian Ministry of University and Research<br> &ldquo;Fondo finalizzato al rilancio degli investimenti delle amministrazioni centrali dello Stato e allo sviluppo del Paese, legge 145/2018&rdquo;.<br> English language editing was performed by Tornillo Scientific, UK.</p> <p><br> Data availability statement:</p> <p>1.- Seismic data from Kilauea, Augustine, Bezymianny (2007), and Mount St. Helens are available from the IRIS data repository (http://ds.iris.edu/seismon/index.phtml).<br> &nbsp;&nbsp;&nbsp; (An example of the Python code to access the data is described below.)<br> 2.- Seismic data from Bezymianny (2017-2018) are available from Ivan Koulakov (ivan.science@gmail.com) upon request.<br> 3.- Seismic data from Mt. Etna are available from INGV-Italy upon request (http://terremoti.ingv.it/en/help),<br> &nbsp;&nbsp;&nbsp;&nbsp; also available from the Zenodo data repository (https://doi.org/10.5281/zenodo.6849621).</p> <p>&nbsp;</p> <p>Access code in Python to download the records of Kilauea, Augustine and Mount St. Helens volcanoes, from the IRIS data repository.</p> <p>&#39;&#39;&#39;To access the raw signals please first install ObsPy and then execute following commands in a python console: &#39;&#39;&#39;</p> <p>Example:</p> <p>from obspy.core import UTCDateTime<br> from obspy.clients.fdsn import Client<br> import obspy.io.mseed<br> client = Client(&#39;IRIS&#39;)<br> t1 = UTCDateTime(&#39;2006-01-10T00:00:00&#39;)<br> t2 = UTCDateTime(&#39;2006-01-12T00:00:00&#39;)<br> raw_data = client.get_waveforms(<br> &nbsp;&nbsp;&nbsp; network=&#39;AV&#39;,<br> &nbsp;&nbsp;&nbsp; station=&#39;AUH&#39;,<br> &nbsp;&nbsp;&nbsp; location=&#39;&#39;,<br> &nbsp;&nbsp;&nbsp; channel=&#39;HHZ&#39;,<br> &nbsp;&nbsp;&nbsp; starttime=t1,<br> &nbsp;&nbsp;&nbsp; endtime=t2)</p> <p>&#39;&#39;&#39;To further download station information execute: &#39;&#39;&#39;</p> <p>xml&nbsp; = client.get_stations(network=&#39;AV&#39;,station=&#39;AUH&#39;,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> channel=&#39;HHZ&#39;,starttime=t1,endtime=t2,level=&#39;response&#39;)</p> <p>&#39;&#39;&#39; &#39;To scale the data using the station&rsquo;s meta-data: &#39;&#39;&#39;</p> <p>data = raw_data.remove_response(inventory=xml)</p> <p>&#39;&#39;&#39; To filter, trim and plot the data execute: &#39;&#39;&#39;</p> <p>data.write(&quot;Augustine.mseed&quot;, format=&quot;MSEED&quot;)</p> <p>data.filter(&#39;bandpass&#39;,freqmin=1.0,freqmax=20)<br> data.trim(t1+60,t2-60)<br> data.plot()</p> <p>Contents:</p> <p>6 different Matlab codes. The principal code is called FeatureExtraction.<br> The codes rsac.m and ReadMSEEDFast.m are for reading different format of data. (Not developed by the group)<br> Seismic Data from Mt. Etna for using as an example.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

Data for: Forecasting shifts in habitat suitability of three marine predators suggests a rapid decline in inter-specific overlap under future climate change

<p><strong><span>Aim:</span></strong><span> To estimate spatiotemporal changes in habitat suitability and inter-specific overlap among three marine predators: Baltic grey seals (<em>Halichoerus grypus grypus</em>), harbour seals (<em>Phoca vitulina</em>), and harbour porpoises (<em>Phocoena phocoena</em>) under contemporary and future conditions.</span></p> <p><strong><span>Location: </span></strong><span>The southwestern region of the Baltic Sea, including the Danish Straits and the Kattegat, one of the fastest-warming semi-enclosed seas in the world.</span></p> <p><strong><span>Methods: </span></strong><span>Location data (&gt;200 tagged individuals) were analysed within the </span><span>maximum entropy (MaxEnt) </span><span>algorithm to estimate changes in total area size and overlap of species-specific habitat suitability between 1997-2020 and 2091-2100. A total of eleven candidate predictor variables were considered </span><span>representing anthropogenic activity, environmental, and climate sensitive oceanographic conditions in the area. Sea surface temperature and salinity</span><span> data were taken from </span><span>representative concentration pathways [RCPs] scenarios 6.0 and 8.5</span><span> to forecast potential </span><span>climate change effects</span><span>.</span></p> <p><strong><span>Results:</span></strong><span> Model output suggests that habitat suitability of Baltic grey seals will decline drastically over space and time, largely driven by changes in sea surface salinity and a loss of currently available haulout sites following sea level rise in the future. A similar though weaker response was observed for harbour seals, while suitability of habitat for harbour porpoises was predicted to remain fairly stable over space and time. Inter-specific overlap in highly suitable habitat was predicted to increase slightly under RCP scenario 6.0 when compared to contemporary conditions but to largely disappear under RCP scenario 8.5.</span></p> <p><strong><span>Main conclusions:</span></strong><strong> </strong><span>Marine predators in the southwestern Baltic Sea and adjacent waters may respond differently to future climatic conditions, leading to divergent shifts in habitat suitability that are likely to decrease inter-specific overlap.<strong> </strong>We, therefore, conclude that climate change can lead to a marked redistribution of area use by marine predators in the region, which may influence local food-web dynamics and ecosystem functioning.</span></p>

opencc-zeroJul 2022View details →
zenodo36/100

Dataset for Embedded Temporal Convolutional Networks for Essential Climate Variables Forecasting

<p>The dataset contains time series of surface soil moisture at three locations, namely Idaho, Indiana, and Oklahoma.</p> <p>For each location, the folder contains multiple gif images, one for each year, each containing&nbsp;12 frames corresponding to monthly averages.</p> <p>More information at&nbsp;</p> <p>https://github.com/gtsagkatakis/ETCN</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Forecasting 24-hour-averaged PM2.5concentration in the Aburrá Valley using tree-based ML models, global forecasts, and satellite information: Dataset

<p>Data necessary for the training and evaluating the 24-hourly-averaged PM2.5 forecast over 19 stations within the Aburr&aacute; Valley, Colombia,&nbsp;is included here.</p>

opencc-by-4.0Sep 2022View details →
dryad36/100

Forecasting the publication and citation outcomes of Covid-19 preprints

<p>The scientific community reacted quickly to the <em>Covid-19</em> pandemic in 2020, generating an unprecedented increase in publications. Many of these publications were released on preprint servers such as <em>medRxiv</em> and <em>bioRxiv</em>. It is unknown however how reliable these preprints are, and if they will eventually be published in scientific journals. In this study, we use crowdsourced human forecasts to predict publication outcomes and future citation counts for a sample of 400 preprints with high <em>Altmetric</em> scores. Most of these preprints were published within one year of upload on a preprint server (70%), and 46% of the published preprints appeared in a high-impact journal with a Journal Impact Factor of at least 10. On average, the preprints received 162 citations within the first year. We found that forecasters can predict if preprints will be published after one year and if the publishing journal has high impact. Forecasts are also informative with respect to preprints' rankings in terms of <em>Google</em> <em>Scholar</em> citations within one year of upload on a preprint server. For both types of assessment, we found statistically significant positive correlations between forecasts and observed outcomes. While the forecasts can help to provide a preliminary assessment of preprints at a faster pace than the traditional peer-review process, it remains to be investigated if such an assessment is suited to identify methodological problems in pre-prints. </p>

opencc-zeroSep 2022View details →
zenodo36/100

Forecast of the competitiveness of EU regions in the conditions of climate change

<p>The database contains forecasts of the values&nbsp;of climate change competitiveness (Regional Climate Change Competitiveness Index) of EU regions (NUTS2) and countries. The calculations were made for the period 2022-2032&nbsp;and for period 2022-2100 using the ARIMA method.</p> <p>The Regional Climate Change Competitiveness Index is a mean&nbsp;to evaluate the ability of a region to use factors of competitiveness for the formation of a competitive position of the region under climate change conditions.The structure of the index stems from the premise that it constitutes a function of pillars that can be grouped into six broad sub-indexes: Basic, Natural, Efficiency, Innovation, Sectoral, and Social.Twenty-eight pillars were used in order to determine the main index.&nbsp; The higher the value of Regional Climate Change Competitiveness Index, the higher the level of regional competitiveness. For detailed methodology see: Karman, A.; Miszczuk, A.; Bronisz, U. Regional Climate Change Competitiveness&mdash;Modelling Approach. Energies 2021, 14, 3704. https://doi.org/10.3390/en14123704.</p> <p>The forecast can be used by regional authorities to model regional policy (policy-mix, tools, measures) against the climate change.</p> <p>Funding: National Science Centre Poland, &bdquo;Modelling of climate change impacts on regional competitiveness&rdquo; 2019/35/B/HS5/01548</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Material for manuscript submitted to Earth and Space Science "Evaluation of a mesoscale coupled ocean-atmosphere configuration for tropical cyclone forecasting in the South West Indian Ocean basin"

<p>Configuration files for AROME Indian Ocean, NEMO and OASIS which are necessary to reproduce the results in the publication :</p> <p>Corale, L;&nbsp; Malardel S. , Bielli S. and M-N Bouin (2022) Evaluation of a mesoscale coupled ocean-atmosphere configuration for tropical cyclone forecasting in the South West Indian Ocean basin. <em>Earth and Space Science.</em></p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Supporting Data for pyCSEP: An Enhanced Python Toolkit for Earthquake Forecast Developers

<p>Contains the reproducibility package for pyCSEP: An Enhanced Python Toolkit for Earthquake Forecast Developers</p> <p>Please view the README.md in this archive for instructions on how to run the reproducibility package.</p> <p>The source code for pyCSEP used in this reproducibility package can be viewed on GitHub&nbsp;<a href="https://github.com/SCECcode/pycsep/releases/tag/v0.5.2">here</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Рис. 10. Блок-схема фиЗико-статистического прогноЗа уроЖайности спата приморского гребешка. in Review of methods for the forecast of mollusk's spat productivity in sea-farms of Primorye and probable ways of their enhancement

Рис. 10. Блок-схема фиЗико-статистического прогноЗа уроЖайности спата приморского гребешка.

opencc-by-4.0Dec 2018View details →
dryad36/100

Can ingredients based forecasting be learned? Disentangling a random forest's severe weather predictions

<p>Machine learning (ML)-based models have been rapidly integrated into forecast practices across the weather forecasting community in recent years. While ML tools introduce additional data to forecasting operations, there is a need for explainability to be available alongside the model output, such that the guidance can be transparent and trustworthy for the forecaster. This work makes use of the algorithm tree interpreter (TI) to disaggregate the contributions of meteorological features used in the Colorado State University Machine Learning Probabilities (CSU-MLP) system, a random forest-based ML tool that produces real-time probabilistic forecasts for severe weather using inputs from the Global Ensemble Forecast System v12. TI feature contributions are analyzed in time and space for CSU-MLP day-2 and 3 individual hazard (tornado, wind, and hail) forecasts and day-4 aggregate severe forecasts over a 2-yr period. For individual forecast periods, this work demonstrates that feature contributions derived from TI can be interpreted in an ingredients-based sense, effectively making the CSU-MLP probabilities physically interpretable. When investigated in an aggregate sense, TI illustrates that the CSU-MLP system's predictions use meteorological inputs in ways that are consistent with the spatiotemporal patterns seen in meteorological fields that pertain to severe storms climatology. This work concludes with a discussion on how these insights could be beneficial for model development, real-time forecast operations, and retrospective event analysis.</p>

opencc-zeroMay 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record