Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
815
datasets available to search
ShareScore release 0.7.1
Dataset results
815 results for “Forecasting”
Hellisheiði geothermal field: Hydraulic data for pseudo-prospective forecasting models (Dec. 2018 - Jan. 2021)
<p>This dataset comprises the compound volumes processed from injection and production rates in the Hellisheiði field between December 2018 and January 2021. This dataset corresponds to the input data for the ETAS-f and Seismogenic Index models of Ritz et al., 2023 (doi:<a href="https://doi.org/10.22541/essoar.168500354.49240043/v1">10.22541/essoar.168500354.49240043/v1</a>)</p><p>The hydraulic data was acquired and processed by Reykjavik Energy/ON power, the operator of the Hellisheiði geothermal field.</p>
Additional resources for "Day ahead electricity price forecasting with neural networks - one or multiple outputs?"
Open the record for dataset details and reuse information.
Monsoon Mission Coupled Forecast System Version 2.0: Model Description and Indian Monsoon Simulations Figures
<p>Monsoon Mission Coupled Forecast System Version 2.0: Model Description and Indian Monsoon Simulations Figures</p>
Application of artificial neural network to forecast indoor air temperature in a building with artificial ventilation: impact of early stopping.
<p>Indoor air temperature prediction can facilitate energy-saving actions without compromising the indoor thermal comfort of occupants. The aim of this study was to analyse the performance of various artificial neural networks with a view to proposing an optimal approach for predicting the indoor temperature of a tertiary building with artificial ventilation. The MLP, CNN, LSTM models and the CNN-LSTM combination (long short-term memory network) were used and coupled with the optimisation algorithms (Adam, SGD) and the independent hyper-parameters early stopping and dropout. The parameters used are outdoor ambient temperature, outdoor relative humidity, indoor relative humidity, wet bulb temperature, black globe temperature and mean radiant temperature. The data is collected in an artificially ventilated building in Yaoundé, Cameroon. A numerical code was developed in Python to run the simulations. In order to study the impact of the parameters on the prediction, two scenarios were distinguished in this work: (1) all the parameters are input to the network, (2) only the parameters whose absolute value of the correlation coefficient was greater than or equal to 0.5 were used. The impact of early stopping is assessed by distinguishing two case studies: the first without early stopping, the second with early stopping. The results showed that without early stopping, the MLP, CNN, LSTM and CNN-LSTM networks are adequate for predicting the temperature with the second scenario, mainly with both the SGD and Adam algorithms, and CNN-LSTM is the most appropriate model because the MSE and MAE values obtained in this case were closer to 0. With early stopping, the learning time is reduced and the learning curves are improved; the models optimised better with the SGD algorithm in general, but the best neural network model was obtained with the Adam algorithm and the LSTM network for the performances MSE=0.0005, MAE=0.0130 with the second scenario.</p><p><strong>Keywords: </strong>prediction, indoor temperature, artificial neural network, early stopping, artificially ventilated building.</p>
Data underlying the publication: "CAR36, a regional high-resolution ocean forecasting system for improving drift and beaching of Sargassum in the Caribbean Archipelago."
<p><strong>CAR36 dataset</strong></p><p>These data correspond to the <strong>1-year (2019)</strong> simulation from the regional ocean system CAR36. These <strong>daily hindcasts</strong> have been used in the study presented in the paper submitted in GMD editor and entitled: "CAR36, a regional high-resolution ocean forecasting system for improving drift and beaching of Sargassum in the Caribbean Archipelago", where the CAR36 system is fully described.</p><p><br>The uploaded files are in <strong>netcdf</strong> format:</p><ul><li><i>CAR36_daily_SSH_20190102-20191224.nc</i> = 1-year daily hindcasts of <strong>Sea Surface Height </strong></li><li><i>CAR36_daily_SST_20190102-20191224.nc </i>= 1-year daily hindcasts of <strong>Sea Surface Temperature</strong></li><li><i>CAR36_daily_SSU_20190102-20191224.nc</i> = 1-year daily hindcasts of <strong>Sea Surface Current Speed (zonal component)</strong></li><li><i>CAR36_daily_SSV_20190102-20191224.nc</i> = 1-year daily hindcasts of <strong>Sea Surface Current Speed (meridian component)</strong></li></ul><p>All data are projected on the native model tripolar<strong> ORCA grid</strong> <strong>in 1/36° </strong>horizontal resolution.</p><p>NB: In order to filter (in a 1st order) the semi-diurnal tidal signal (with a period of 12h30), the daily mean corresponds to a 25h-average. </p><p><strong>CAR36 software</strong></p><p>The NEMO_CAR36.tar file gathers the <strong>NEMO code configuration</strong> of the CAR36 model. This code follows the same license than NEMO one : <strong>CeCILL</strong>. A file named "License_CeCILL.txt" reminds the details of this license in the NEMO_CAR36.tar file.<br><br>NB.: This model have been renamed CAR36 (English acronym) for the paper instead of ARCAN36 (French initial acronym). In the provided NEMO code, the name ARCAN36 is still used. </p>
Weather forecasts from multiple models and observations at Norwegian synop stations
<p>The data consists of temperature (2m) and wind speed (10m) observations at 183 Norwegian synop stations and corresponding weather forecasts generated by the following models</p><ul><li>Pangu-Weather</li><li>ECMWF HRES</li><li>ECMWF ENS reforecast</li><li>MEPS</li><li>ECMWF ENS reforecast control member</li><li>MEPS control member</li></ul><p>The data was created for the article <a href="https://arxiv.org/abs/2309.01247"><i>Evaluation of forecasts by a global data-driven weather model with and without probabilistic post-processing at Norwegian stations</i> </a>with <a href="https://github.com/jbbremnes/pangu-asr">source codes</a> available in GitHub. The data is stored as a single JLD2 (HDF5) file containing a vector of six data frames.</p>
Forecasts, score summary files, target observational data and meteorological driver files to accompany the manuscript "Skill of process-based forecasts relative to multiple null models varies across time and depth for water temperature and dissolved oxygen"
<p>This data publication includes raw ensemble forecast output (forecasts.zip), as well as summary score files (scores.zip) for process-based forecasts produced with the Forecasting Lake and Reservoir Ecosystems (FLARE) framework. In addition, it includes scores for climatology (climatology_scores.csv) and random walk (RW_scores.csv) null forecasts, formatted observational data of target variables (sunp-targets-insitu.csv), and meteorological driver files required for analysis to accompany the manuscript "Skill of process-based forecasts relative to multiple null models varies across time and depth for water temperature and dissolved oxygen". Forecasts were made of water temperature and dissolved oxygen at Lake Sunapee, NH in 2021 and 2022.</p>
Supplement of "Algorithm for continual monitoring of fog life cycles based on geostationary satellite imagery as a basis for solar energy forecasting"
<p>The file uploaded here is an animation that visually illustrates the outputs of the a newly developed machine learning based FLS (<strong>F</strong>og and <strong>L</strong>ow <strong>S</strong>tratus) detection algorithm for the SEVIRI (<strong>S</strong>pinning <strong>E</strong>nhanced <strong>V</strong>isible and <strong>I</strong>nfra<strong>R</strong>ed <strong>I</strong>mager) instrument onboard the MSG (<strong>M</strong>eteosat <strong>S</strong>econd <strong>G</strong>eneration) geo-stationary satellites over the 24hr cycle of the day for the day of <strong>02/March/2021</strong> and compares them with the corresponding raw channel values observed by SEVIRI. The proposed algorithm classifies each SEVIRI pixel as "clear-sky", "FLS", or "non-FLS-cloud" (identified with Khaki, Red, and Blue in the animation) based on the SEVIRI pixel values of BT12.0, BT8.7 - BT12.0, BT10.8 - BT12.0, and BT12.0 - BT13.4 plus the standard deviation of each of these variables in a spatial window sized 3x3 pixels with the central pixel being the target pixel. </p><p><br>In this animation, the left-hand panel shows a false-color RGB image constructed based on the SEVIRI raw channel data with the red, green, and blue channels being BT12.0- BT13.4, BT8.7 - BT12.0, and BT10.8 - BT12.0, respectively. In this panel, the green color represents the high clouds, and the light and dark red colors represent the clear-sky and FLS, respectively. The right-hand panel of this animation also shows the outputs of the ML FLS detection algorithm developed in the present study.</p>
Quantitative precipitation estimate (QPE) and forecast (QPF) exceedance comparison with flash flood reports
<p class="ParagraphText">Flash flooding remains a challenging prediction problem, which is exacerbated by the lack of a universally accepted definition of the phenomenon. In this article, we extend prior analysis to examine the correspondence of various combinations of quantitative precipitation estimates (QPE) and precipitation thresholds to observed occurrences of flash floods, additionally considering short-term quantitative precipitation forecasts from a convection-allowing model. Consistent with previous studies, there is large variability between QPE datasets in the frequency of "heavy" precipitation events. There is also large regional variability in the best thresholds for correspondence with reported flash floods. In general, Flash Flood Guidance (FFG) exceedances provide the best correspondence with observed flash floods, except in the interior western US where recurrence interval thresholds (for the southwestern US) and static thresholds (for the northern and central Rockies) provide better correspondence. Six-hour QPE provides better correspondence with observed flash floods than 1-h QPE in all regions except the west coast and southwestern US. Exceedances of precipitation thresholds in forecasts from the operational High-Resolution Rapid Refresh (HRRR) generally do not correspond with observed flash flood events as well as QPE datasets, but they outperform QPE datasets in some regions of complex terrain and sparse observational coverage such as the southwestern US. These results can provide context for forecasters seeking to identify potential flash flood events based on QPE or forecast-based exceedances of precipitation thresholds. </p>
HyPhAICCast-1 2-hour forecast on 01/01/2021 at 12:00 p.m.
Open the record for dataset details and reuse information.
Model checkpoints for "SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models"
<p>Checkpoints for all SEEDS models in the paper <a href="https://arxiv.org/abs/2306.14066" rel="nofollow">https://arxiv.org/abs/2306.14066</a>, including all SEEDS-GEE and SEEDS-GPP models in the main manuscript, and the additional models trained in the Supplemental Material.</p> <p>Checkpoint naming convention:</p> <ul> <li><code>gee_c2_s7</code>: SEEDS-GEE trained conditioning on 2 seeds for 7-day lead time.</li> <li><code>gpp_c2_s7_g3_r4</code>: SEEDS-GPP trained conditioning on 2 seeds for 7-day leadtime, where the label mixture is 3 GEFS members and 4 ERA reanalyses.</li> </ul>
Data from: Forecasting nocturnal bird migration for dynamic aeroconservation: the value of short-term dataset
<p>Placing wind turbines within large migration flyways, such as the North Sea basin, can contribute to the decline of vulnerable migratory bird populations by increasing mortality through collisions. Curtailment of wind turbines limited to short periods with intense migration can minimize these negative impacts, and near-term bird migration forecasts can inform such decisions. Although near-term forecasts are usually created with long-term datasets, the pace of environmental alteration due to wind energy calls for urgent development of conservation measures that rely on existing data, even when it does not have long temporal coverage. Here, we use five years of tracking bird radar data collected off the western Dutch coast, weather, and phenological variables to develop seasonal near-term forecasts of low-altitude nocturnal bird migration over the southern North Sea. Overall, the models explained 71% of the variance and correctly predicted migration intensity above or below a threshold for intense hourly migration in more than 80% of hours in both seasons. However, the percentage of correctly predicted intense migration hours (top 5% of hours with the most intense migration) was low, likely due to the short-term dataset and their rare occurrence. We, therefore, advise careful consideration of a curtailment threshold to achieve optimal results. Synthesis and applications: Near-term forecasts of migration fluxes evaluated against measurements can be used to define curtailment thresholds for offshore wind energy. We show that to minimize collision risk for 50% of migrants, if predicted correctly, curtailments should be applied during 18 hours in spring and 26 in autumn in the focal year of model assessments, resulting in an estimated annual wind energy loss of 0.12%. Drawing from the Dutch curtailment framework, which pioneered the 'international first' offshore curtailment, we argue that using forecasts developed from limited temporal datasets alongside expert insight and data-driven policies can expedite conservation efforts in a rapidly changing world. This approach is particularly valuable in light of increasing interannual variability in weather conditions.</p>
HRRR Convective Objects for "Comparing Distributions of Overshooting Convection in HRRR Forecasts to Observations"
<p>Data files for all convective objects identified from High-Resolution Rapid Refresh (HRRR) model forecasts for the Contiguous United States initialized at 06 UTC every day of May and July 2021. Convection was defined in the following two ways, one corresponding to each file uploaded here:</p> <ul> <li>Meeting the criteria for convection defined by the Storm Labelling in Three-Dimensions (SL3D; Starzec et al. 2017) algorithm</li> <li>Satisfying two vertical velocity thresholds (2 m/s at 4 km altitude and 5 m/s at 8 km altitude)</li> </ul>
Observed and forecasted pollutant concentrations in the center of Sao Paulo
<p>This dataset corresponds to the observed and forecasted pollutant concentrations analyzed in an article submitted to <em><strong>Earth's Future</strong></em>, which is entitled: <em><strong>Air quality forecasts with observation-based scaling of anthropogenic emissions for urban agglomerations. </strong></em></p> <p>The pollutants analyzed are carbon monoxide (CO), nitrogen dioxide (NO2), ozone (O3), sulfur dioxide (SO2) and particulate matter concentrations (PM2.5 and PM10).</p> <p>Forecasted concentrations are presented for three regional simulations (F-REF, F-DAY and F-HOUR) and two flobal simulations (WACCM and CAMS) for the next day (day_p1 is d+1) and for the day after the next day (day_p2 is d+2). </p> <p>The observed data are interpolated in the center of Sao Paulo, following the methodology presented in <a href="https://doi.org/10.1029/2022JD038179">https://doi.org/10.1029/2022JD038179</a></p>
Seamless short- to mid-term probabilistic wind power forecasting: Forecasting Results
<div> <div> <div> <div> <div>This dataset contains the predicted time series and the forecast assessment conducted in the paper <strong>Seamless short- to mid-term probabilistic wind power forecasting</strong>.</div> </div> </div> </div> </div>
Dataset for evaluation of AROME and SURFEX-SA forecasts
<p>SQL Dataset for the verification of AROME (first reference) and SURFEX-SA (produced for the paper) predictions of Austrian cities (Vienna, Linz, Klagenfurt and Innsbruck) with TAWES weather station data. In addition, data from the ensemble prediction system C-LAEF are also included. AROME and C-LAEF are originally available on 2.5 km grid, but they were bilinear interpolated to the respective TAWES coodinates. Datasets can be easily read by the verification tool "harp" in R.</p> <p>TAWES observations are free available from the GeoSphere Austria Datahub (second reference).</p>
Global Flood Forecasting (GFF)
<h3>Dataset from paper: <strong><em>Off to new Shores: A Dataset & Benchmark for (near-)coastal Flood Inundation Forecasting</em></strong></h3> <p>Consists of target flood maps and various inputs used to predict said flood maps. Each input modality is contained in a separate zip file. Each zip file contains a collection of raster data, one file per ROI.</p> <p><em><strong>Files:</strong></em></p> <ul> <li><strong>base.zip</strong> - Contains flood maps and meta-data (e.g. geolocations). Includes <a href="https://github.com/Orion-AI-Lab/KuroSiwo" target="_blank" rel="noopener">Kuro Siwo labels</a>.</li> <li><strong>s1.zip</strong> - Contains Copernicus Sentinel-1 data from 2014-2020 obtained through Alaska Foundation, using an EarthData account.</li> <li><strong>era5.zip</strong> - Contains ERA5 and ERA5-Land data products obtained from Google Earth Engine (GEE), which was originally generated using Copernicus Climate Change Service information 2014-2020.<em><br></em></li> <li><strong>glofas.zip</strong> - Contains <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/cems-glofas-historical" target="_blank" rel="noopener">GLOFAS v4.0 data</a> from Copernicus Climate Change Service's Climate Data Store (CDS).</li> <li><strong>hydroatlas.zip</strong> - Contains <a href="https://www.hydrosheds.org/hydroatlas" target="_blank" rel="noopener">HydroATLAS v10 data</a>. Rasterized and resampled to match ERA5-Land.</li> <li><strong>hand.zip</strong> - Contains <a href="https://gis.asf.alaska.edu/arcgis/rest/services/GlobalHAND/GLO30_HAND/ImageServer" target="_blank" rel="noopener">HAND data</a> (CC0 license). Resampled to match Sentinel-1.</li> <li><strong>dem.zip</strong> - Contains <a href="https://spacedata.copernicus.eu/collections/copernicus-digital-elevation-model" target="_blank" rel="noopener">CopDEM GLO-30 data</a>. Resampled to match Sentinel-1 and ERA5-Land.</li> <li><strong>extras.zip</strong> - Contains 4 folders of extra data: <ul> <li><em>vit-snunet-extra</em> - Flood maps, meta-data and S1 of ROIs that fell outside era5 availability.</li> <li><em>vit-snunet-raw</em> - Raw floodmaps generated by Kuro Siwo model (after removing tiling effects, before other post-processing)</li> <li><em>vit</em> - Alternate flood maps; generated using FloodViT model only</li> <li><em>snunet</em> - Alternate flood maps; generated using SNUNet model only</li> </ul> </li> <li><strong>worldcover.zip</strong> - Contains <a href="https://worldcover2021.esa.int/download" target="_blank" rel="noopener">worldcover data</a> (CC-BY license) resampled (nearest) to Sentinel-1. Only used at evaluation time. </li> </ul> <p><em><strong>Instructions</strong></em>: </p> <ol> <li>Download whichever inputs you want to use.</li> <li>Extract all in the same folder.</li> <li>Point <a href="https://github.com/Multihuntr/gff" target="_blank" rel="noopener">the code</a> to that folder.</li> </ol> <p><em><strong>Licenses</strong></em>: The redistributed data products (s1, era5, hydroatlas, hand, dem and worldcover) retain their original licenses. They are, however, very permissive. Only HydroATLAS is not necessarily available for commercial use.</p>
CNN-Based Forecasting of Pitch Angle-Resolved Energetic Electron Flux at MEO Using Solar Wind and Geomagnetic Data
<div> <p> CNN-Based Forecasting of Pitch Angle-Resolved Energetic Electron Flux at MEO Using Solar Wind and Geomagnetic Data. The model and test set data are provided here. Data used for training, validating, and testing the 1.8MeV channel model is also offered as an example.</p> </div> <p><strong>initial_data: </strong>The test set data has been normalized and can be used as model input.</p> <p><strong>norm para: </strong>The normalization parameters used for data processing</p> <p><strong>model_test_dataset_performance.py: </strong>The script to obtain the outputs of the models at different energy levels on the test set. Before running it, unzip “initial_data.rar” and "norm_para.rar"</p> <p><strong>full_dataset_for_rept_ch0: </strong>Data used for training, validating, and testing the 1.8MeV channel model. It's not essential for model_test_dataset_performance.py</p>
A deep learning-based parametric inversion for forecasting water-filled bodies position using electromagnetic method
<p>We design a tunnel electromagnetic joint scan observation system and present a deep learning-based parametric inversion for improved tunnel electromagnetic imaging, designed specifically for tunnel prediction of water filled structures. It utilizes a configuration wherein transmitters scan along the surface while receivers are positioned within the tunnel, employing time-domain and frequency-domain transmitters and a multi-component receiver. The DL model for the first time provides parametric imaging of two different view, forming a self-checking mechanism, which can help constrain the predictions and reduce the non-uniqueness of the inversion. Trained by synthetic data, our system shows impressive adaptability to predict the 3D spatial position of water-filled anomalies and strong robustness in the tunnel environment with metal interference.</p> <p> </p> <p>Before prediction, you need to download the pre-trained weight model which contain UNet_model/FTEM.ckpt.data-00000-of-00001, UNet_model/FTEM.ckpt.index, UNet_model/FTEM.h5. Then place the directory containing the weight model in the same directory as the prediction code. Then run: python predi.py</p>
Forecasting of the Geomagnetic Activity for the Next 3 Days Utilizing Neural Networks Based on Parameters Related to Large-scale Structures of the Solar Corona
<p>These are supplementary data for the paper "Forecasting of the Geomagnetic Activity for the Next 3 Days Utilizing Neural Networks Based on Parameters Related to Large-scale Structures of the Solar Corona". They are:</p> <ul> <li>Python code to forecast Kp index</li> <li><span>Code to construct a nerual network model</span></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.