Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
182
datasets available to search
ShareScore release 0.9.0
Dataset results
182 results for “Weather forecast”
Influence of weather forecast resolution on the circulation of Lake George, NY.
This dataset contains outputs of numerical modeling for Lake George, New York, hydrodynamics. These numerical simulations were generated to assess the impact of increasing the resolution of weather forecasts on the lake’s thermal state. This research focused on June 2017, when an increase of biological activity was associated to the deepening of the thermocline in the south of the lake. Increasing the resolution of the weather forecast led to a more accurate representation of the water temperature in the lake, including the deepening of the thermocline. The dataset was used in support of “The influence of weather forecast resolution on the circulation of Lake George, NY”.
Impact of urban and shipping emissions on NASA-Unified Weather Research and Forecasting model results
<p>This dataset supports Huang et al. (2019, JGR-Atmospheres): "Impact of aerosols from urban and shipping emission sources on terrestrial carbon uptake and evapotranspiration: a case study in East Asia". The file named "NUWRFout.tar.gz" contains NUWRF base and sensitivity simulation results on 31 May 2016. The file named "LIS_soil_LAI.zip" contains model grid information, soil conditions and leaf area index (LAI) at NUWRF initialization times in late May 2016.</p>
Caravan MultiMet (Part 2, Forecasts): Extending Caravan with Multiple Weather Nowcasts and Forecasts
<p>Caravan MultiMet is a novel extension to Caravan, focusing on enriching the meteorological forcing data. Our extension adds three precipitation nowcast products (CPC, IMERG v07 Early, and CHIRPS) and three weather forecast products (ECMWF IFS HRES, GraphCast, and CHIRPS-GEFS). Since all data is kept in it's original time zone (UTC+0) and the ERA5-Land data in the original Caravan data set is shifted to local time of each gauge, we also include ERA5-Land reanalysis data in this extension, matching the UTC-0 timezone for all gauges of the other forcings. This part of the extension includes the <strong>forecast</strong> products.</p> <p>The inclusion of diverse data sources, particularly weather forecasts, enables more robust evaluation and benchmarking of hydrological models, especially for real-time forecasting scenarios. To the best of our knowledge, this extension makes Caravan the first open large-sample hydrology dataset to incorporate weather forecast data.</p> <p>The data is also publicly available on Google Cloud Platform (GCP), and we provide below a colab with an example of how to access it which does not require downloading the entire dataset.</p> <p>Additional resources:</p> <ul> <li>The <a href="https://github.com/kratzert/Caravan">Caravan GitHub repository</a> includes further information and links to other extensions.</li> <li>The original <a href="https://www.nature.com/articles/s41597-023-01975-w">Caravan paper</a></li> </ul> <p>------</p> <p>Channel log:</p> <ul> <li>21 November 2024: Version 1.1 - Fixed bug in the FAO Penman-Monteith potential evaporation - values are now clipped to 0 (previously, some values were negative). Only the ERA5-Land dataset was changed.</li> </ul>
Caravan MultiMet (Part 1, Nowcasts): Extending Caravan with Multiple Weather Nowcasts and Forecasts
<p>Caravan MultiMet is a novel extension to Caravan, focusing on enriching the meteorological forcing data. Our extension adds three precipitation nowcast products (CPC, IMERG v07 Early, and CHIRPS) and three weather forecast products (ECMWF IFS HRES, GraphCast, and CHIRPS-GEFS). Since all data is kept in it's original time zone (UTC+0) and the ERA5-Land data in the original Caravan data set is shifted to local time of each gauge, we also include ERA5-Land reanalysis data in this extension, matching the UTC-0 timezone for all gauges of the other forcings. This part of the extension includes the <strong>nowcast</strong> products.</p> <p>The inclusion of diverse data sources, particularly weather forecasts, enables more robust evaluation and benchmarking of hydrological models, especially for real-time forecasting scenarios. To the best of our knowledge, this extension makes Caravan the first open large-sample hydrology dataset to incorporate weather forecast data.</p> <p>The data is also publicly available on Google Cloud Platform (GCP), and we provide below a colab with an example of how to access it which does not require downloading the entire dataset.</p> <p>Additional resources:</p> <ul> <li>The <a href="https://github.com/kratzert/Caravan">Caravan GitHub repository</a> includes further information and links to other extensions.</li> <li>The original <a href="https://www.nature.com/articles/s41597-023-01975-w">Caravan paper</a></li> </ul> <p>------</p> <p>Channel log:</p> <ul> <li>21 November 2024: Version 1.1 - Fixed bug in the FAO Penman-Monteith potential evaporation - values are now clipped to 0 (previously, some values were negative). Only the ERA5-Land dataset was changed.</li> </ul>
Dataset: Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements
<p>The dataset presented is the companion data to the Journal of Hydrometeorology publication entitled “Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements.” The data that follows contains everything needed to reproduce the spatial inputs for the meteorological station model run using the Spatial Modeling for Resources Framework (SMRF, Havens et al., 2017).</p> <p> </p> <p>Software versions used:</p> <ul> <li>Image Processing Workbench v2.2.0 (Marks et al., 2017)</li> <li>Spatial Modeling for Resources Framework v0.5.3 (Havens et al., 2019)</li> </ul> <p> </p> <p><strong>NOTE:</strong> Reproducing the spatial inputs will generate 10 netCDF files at ~80GB per file.</p> <p> </p> <p><strong>topo.nc</strong> – Contains multiple static layers that are required to run SMRF and iSnobal. The netCDF layers are:</p> <ul> <li>dem – digital elevation model at 100 meter resolution, aggregated from the 10 meter National Elevation Dataset (Archuleta et al., 2017)</li> <li>mask – basin mask for the Boise River Basin</li> <li>veg_height – vegetation height in meters from the National Land Cover Database (Homer et al., 2015)</li> <li>veg_type – vegetation type from the National Land Cover Database</li> <li>veg_tau – vegetation fractional transmissivity derived from the vegetation type</li> <li>veg_k – vegetation emissivity derived from the vegetation type</li> </ul> <p> </p> <p><strong>maxus.nc</strong> – maximum upwind slope netCDF that contains 72 images for all wind directions in 5 degree increments using the algorithm described in Winstral and Marks (2002)</p> <p> </p> <p><strong>Station data:</strong></p> <ul> <li>Contains hourly meteorological station data downloaded from Mesowest (Horel et al., 2002). Data was cleaned and filtered prior to running SMRF.</li> <li>metadata.csv – metadata for 40 stations</li> <li>air_temp.csv – 38 stations</li> <li>cloud_factor.csv – 7 stations</li> <li>precip.csv – 21 stations</li> <li>vapor_pressure.csv – 19 stations</li> <li>wind_direction.csv – 14 stations</li> <li>wind_speed.csv – 14 stations</li> </ul> <p> </p> <p><strong>smrf_config.ini</strong> – Configuration file needed to reproduce the spatial inputs using SMRF. The paths will need to be changed to reflect the data location.</p>
Temperature and rainfall datasets for the paper "STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting"
<p>This page includes spatiotemporal datasets used in the paper <a href="https://doi.org/10.1016/j.neucom.2020.09.060">STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting.</a></p> <p>ARIMA and deep learning models use datasets, as follow:</p> <ul> <li>ARIMA models</li> </ul> <p>baseline-chirps-1981-2019.nc (rainfall data)<br> baseline-ucar-1979-2015.nc (temperature data)</p> <ul> <li>Deep learning models:</li> </ul> <p><em>5-step ahead:</em></p> <p>dataset-chirps-1981-2019-seq5-ystep5.nc (rainfall data)<br> dataset-ucar-1979-2015-seq5-ystep5.nc (temperature data)</p> <p><em>15-step ahead:</em></p> <p>dataset-chirps-1981-2019-seq5-ystep15.nc (rainfall data)<br> dataset-ucar-1979-2015-seq5-ystep15.nc (temperature data)</p>
Dataset of Machine Learning forecasted VTEC from paper: Uncertainty Quantification for Machine Learning-based Ionosphere and Space Weather Forecasting
<p>The *csv files contain forecasted one-day-ahead Vertical Total Electron Content (VTEC), consisting of the mean/median VTEC values and the upper and lower VTEC bounds of the 95% confidence intervals of 4 models based on machine learning for test data.</p> <p>The first part of the *csv file name corresponds to the type of model: SE stands for the super-ensemble VTEC model, QGB stands for the quantile gradient boosting VTEC model, BNN1 stands for the Bayesian neural network VTEC model, and BNN2 stands for the Bayesian neural network with negative log-likelihood (NLL) loss VTEC model. The second part of the file name refers to the geographic location of the VTEC points for which the forecast is performed, i.e., 10E70N for 10 degree of longitude and 70 degree of latitude, 10E40N for 10 degree of longitude and 40 degree of latitude, and 10E10N for 10 degree of longitude and 10 degree of latitude. The last part of the file name corresponds to the test year, i.e., year 2017.</p> <p>The SE_*_2017.csv file consists of 14 columns. The index column ("Date-time") is expressed in Coordinated Universal Time (UTC) as YYYY-MM-DD. Columns 1-3 contain the VTEC forecast results of Random Forest (RF) trained on three data subsets; columns 4-6 contain the VTEC forecast results of Adaptive Boosting (AB) trained on three data subsets; columns 7-9 contain the VTEC forecast results of Gradient Boosting (XGBoost) trained on three data subsets. Column 10 ("Mean") represents the mean of columns 1-9, i.e., the ensemble mean; column 11 ("Std") represents the standard deviation of columns 1-9, i.e., the ensemble spread; columns 12 ("UB") and 13 ("LB") contain the upper and lower bounds of the 95% confidence interval of VTEC, respectively; and column 14 contains the Global Ionosphere Maps (GIM) values of CODE, i.e., the ground-truth in this study.</p> <p>The QGB_*_2017.csv file consists of 4 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Median") contains the median VTEC forecast, column 2 ("LB") contains the lower VTEC bound of the 95% confidence interval, column 3 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 4 contains the GIM values of CODE, i.e., the ground-truth in this study.</p> <p>The BNN*_2017.csv file consists of 5 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Mean") contains the mean VTEC forecast, column 2 ("Std") contains the standard deviation, column 3 contains GIM values of CODE, i.e., ground-truth in this study; column 4 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 5 ("LB") contains the lower VTEC bound of the 95% confidence interval.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>Contact</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>If you have any questions regarding these data, please contact:</p> <p>Randa Natras</p> <p>Deutsches Geodätisches Forschungsinstitut (DGFI-TUM)</p> <p>Technical University of Munich</p> <p>Arcisstraße 21</p> <p>80333 München</p> <p>randa.natras@tum.de</p>
GNSS tomography data for assimilation into the Weather Research and Forecasting model
<p>The data set contains GNSS troposphere tomography estimations of 3D wet refractivity fields for a part of Central Europe (mostly Germany and Czech Republic), for the period of 29 May–14 June 2013 when heavy-precipitation events were observed. The refractivity fields were estimated using two different GNSS tomography models: ATom software package (https://github.com/GregorMoeller/ATom) developed at TU Wien, and the TOMO2 model (Rohm and Bosy, 2011; Rohm et al., 2014; Trzcina and Rohm, 2019) developed at the Wrocław University of Environmental and Life Sciences. Further description of the GNSS tomography processing can be found in the paper by Hanna et al. (2019).</p>
Participant Notes from Chapman Conference on Scientific Challenges Pertaining to Space Weather Forecasting Including Extremes
<p>Compilation of electronic meeting notes made by attendees at the Chapman Conference on Scientific Challenges Pertaining to Space Weather Forecasting Including Extremes.</p> <p>Files are provided for Days 1-3 of the meeting. Day 4 inputs are included in Discussion notes under a separate doi.</p> <p>The Chapman Conference was supported by NSF Award AGS 1848885 and NASA grants 936723.02.01.09.14 and 936723.02.01.11.21</p>
Assessment of uncertainty in weather forecasts
<p>Weather data from the Ebro River Basin Hydrographic Demarcation to train machine learning models to evaluate uncertainty in weather forecasts in real time.</p> <p>The dataset is divided into two parts. To see it and download it completely without splitting, here it is published:</p> <ul> <li><strong><a href="https://open.scayle.es/dataset/assessment-of-uncertainty-in-weather-forecasts">https://open.scayle.es/dataset/assessment-of-uncertainty-in-weather-forecasts</a></strong></li> </ul>
Weather data (forecast and observation) at three locations in France over 2021 for Machine Learning Training
<p>The data provided data are historical weather measurement and forecast at three location in France.</p> <p>Measurements are inside files named OBS_xxx</p> <p>Forecasts are inside files names YYY_xxx, with YYY is the name of the forecast simultion (GFS0.25, WRF12km or WRF3KM).</p> <p>In the two cases, xxx is the name of the site (Site 1, Site2 or Site3).</p> <p><br> <strong>Description of OBS_xxx files:</strong><br> - One line per measurement with hourly resolution<br> - columns are: Date(TU),Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> Date = date of measurement in TU and format DD/MM/YYYY HH:MM<br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The observation data go:<br> from 16/04/2021 00H <br> to 31/01/2022 23H</p> <p><br> <strong>Description of YYY_xxx files:</strong><br> - One line per forecast with hourly resolution<br> - columns are: First date run (TU),forecast hour,Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> First date run (TU) = date of start of the forecast in TU and format DD/MM/YYYY HH:MM. HH could be 00 and 12 according to the cycle of forecast start.<br> forecast hour = forecast hour from the start of the forecast date. 00 = forecast for "first date run". 01 = forecast for "First date run" + 1 hour. .... 95 = forecast for "First date run" + 95 hours.<br> For GFS0.25, forecast hour go from 00 to 95<br> For WRF12km, forecast hour go from 00 to 95<br> For WRF3m, forecast hour go from 00 to 95<br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The forecast data go:<br> from 13/04/2021 00H + 72H = first forecast for the 16/04/2021 00H<br> to 31/01/2022 12H + 11H = last forecast for the 31/01/2022 23H</p>
Lightning Assimilation in the Weather Research and Forecasting (WRF) Model: Technique Updates and Assessment of the Applications from Regional to Hemispheric Scales
<p>Figure 1. The data is proprietary, but it can be purchased from Vaisala Inc. (https:// <a href="http://www.vaisala.com/en/products/systems/lightning-detection">www.vaisala.com/en/products/systems/lightning-detection</a>), and the WWLLN raw data are also available for purchase at <a href="http://wwlln.net">http://wwlln.net</a>.</p> <p>Figure 2. Maps, data is not applicable.</p> <p>Figure 3. Data file: NLDN_WWLLN_Prism_Rainfall_Analysis.xlsx</p> <p>Figure 4. Data file: NLDN_WWLLN_METVARS_T2_Jul_2016.xlsx</p> <p>Figure 5. Data file: CONUSall_METOBS_q_Jul_2016.xlsx</p> <p>Figure 6. Data file: CONUSall_METOBS_ws_Jul_2016.xlsx</p> <p>Figure 7. Created using the R script: Hemi_Rain_ModelOnlyWGPM.R based on the R object files: AnnualRainFall_CFC_WRF_Hemi_BASE_*.rds, AnnualRainFall_CFC_WRF_Hemi_LTA_*.rds, and GPM_WRF_Paired_rain2Hemispheric_July2016.rds.</p> <p>Figure 8. Created using the R script: Hemi_Rain_Aanlysis.R based on the R object files: AnnualRainFall_CFC_WRF_Hemi_BASE_*.rds and AnnualRainFall_CFC_WRF_Hemi_LTA_*.rds.</p> <p>Figure 9. Data file: CPC_Model_Monthly_Prep_Hemi_Stats.xlsx</p> <p>Figure 10. Data file: CPC_Model_Monthly_Prep_Hemi_Stats.xlsx</p> <p>Figure 11. Data file: CPC_Model_Monthly_Prep_Hemi_Stats.xlsx</p> <p>Figure 12. Created using the R script: CreateCPCdataforUSdomain_vs_Prism.R based on the R oject files: Prism_CFC_WRF*.rds</p> <p>Figure 13. Data file: Hemi_lta_METOBS_T2_Jul_2016.xlsx</p> <p>Figure 14. Data file: Hemi_lta_METOBS_q_Jul_2016.xlsx</p> <p> </p>
An Objective Detection of Separation Scenario in Tropical Cyclone Trajectories Based on Ensemble Weather Forecast Data
<p>This repository contains the data used in "An Objective Detection of Separation Scenario in Tropical Cyclone Trajectories Based on Ensemble Weather Forecast Data" by Oettli and Kotsuki (submitted to Journal of Geophysical Research: Atmospheres).</p>
Datasets with weather forecasts (Temperature, Wind Direction, Humidity, Pressure, Wind Speed, GHI)
<p>Datasets with weather forecasts for HLU7 (Temperature, Wind Direction, Humidity, Pressure, Wind Speed, GHI) of the CROSSBOW project.</p>
Rye microgrid historical weather forecasts and stochastic scenarios
<p>This datasets connects historical weather forecasts from the Norwegian Meteorological Institute (met.no) and historical observations from Rye microgrid (https://doi.org/10.5281/zenodo.4448894).</p> <p>Each csv file represents a historical weather forecast for approximately 60 hours ahead. Each csv-file also contains the corresponding observations in the same time interval. Finally, the files also contain load, wind generation and solar PV generation predicitons.</p> <p>The predictions are generated using gradient boosting. The predictions ending with "_ls" are based on least square. The predictions ending with "_quantile_i" represent a quantile prediction. For example, "wind_quantile_2" means that there is a 20% probability the wind will be less than this value.</p> <p>The gradient boosting predictor has been trained to predict the wind power, solar power and load using the explanatory variables below:</p> <p>Solar PV: Cloud area fraction, initial production, clear sky production and forecast look-ahead time</p> <p>Wind power: wind speed, wind direction, wind power converted from wind speed forecast, initial production and forecast look-ahead time</p> <p>Load: hour of day, month of year</p>
Huge ensembles part I design of ensemble weather forecasts with spherical Fourier neural operators; Huge ensembles part II properties of a huge ensemble of hindcasts generated with spherical Fourier neural operators
Open the record for dataset details and reuse information.
Impact of Different Nesting Methods on the Simulation of a Severe Convective Event Over South Korea Using the Weather Research and Forecasting Model
<p>The data from various platforms (NCEP FNL, TRMM, ERA5, AWS) and WRF Model output utilised to generate the figures in the current study (https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2020JD033084) are available at this Zenodo data repository.</p>
Supplementary video content for "Three-dimensional visualization of ensemble weather forecasts", Parts 1 and 2 (Geoscientific Model Development, 2015)
<p>Supplementary video content (full resolution) for the papers "Three-dimensional visualization of ensemble weather forecasts - Part 1: The visualization tool Met.3D (version 1.0)" and "Three-dimensional visualization of ensemble weather forecasts - Part 2: Forecasting warm conveyor belt situations for aircraft-based field campaigns", Geoscientific Model Development, 2015. The corresponding papers can be found on http://geosci-model-dev.net/.</p>
Weather forecasts from multiple models and observations at Norwegian synop stations
<p>The data consists of temperature (2m) and wind speed (10m) observations at 183 Norwegian synop stations and corresponding weather forecasts generated by the following models</p><ul><li>Pangu-Weather</li><li>ECMWF HRES</li><li>ECMWF ENS reforecast</li><li>MEPS</li><li>ECMWF ENS reforecast control member</li><li>MEPS control member</li></ul><p>The data was created for the article <a href="https://arxiv.org/abs/2309.01247"><i>Evaluation of forecasts by a global data-driven weather model with and without probabilistic post-processing at Norwegian stations</i> </a>with <a href="https://github.com/jbbremnes/pangu-asr">source codes</a> available in GitHub. The data is stored as a single JLD2 (HDF5) file containing a vector of six data frames.</p>
Model checkpoints for "SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models"
<p>Checkpoints for all SEEDS models in the paper <a href="https://arxiv.org/abs/2306.14066" rel="nofollow">https://arxiv.org/abs/2306.14066</a>, including all SEEDS-GEE and SEEDS-GPP models in the main manuscript, and the additional models trained in the Supplemental Material.</p> <p>Checkpoint naming convention:</p> <ul> <li><code>gee_c2_s7</code>: SEEDS-GEE trained conditioning on 2 seeds for 7-day lead time.</li> <li><code>gpp_c2_s7_g3_r4</code>: SEEDS-GPP trained conditioning on 2 seeds for 7-day leadtime, where the label mixture is 3 GEFS members and 4 ERA reanalyses.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.