Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
815
datasets available to search
ShareScore release 0.7.1
Dataset results
815 results for “Forecasting”
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 2)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 2)</p>
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 1)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 1)</p>
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 3)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 3)</p>
AI4ER MRes Models & Forecasts
<p>A collection of model checkpoints and forecasts generated for the MRes thesis "Towards sharp sea ice concentration forecasts in the Arctic" by Andrew McDonald. See 4_forecast.ipynb and 5_evaluate.ipynb in the notebooks folder of the project GitHub repository at https://github.com/ampersandmcd/icenet-gan for a demo of how to make use of these files.</p> <p><strong>Checkpoints</strong></p> <ul> <li>radiant-sponge-59-great-unet-epoch=11-step=456 contains the PyTorch model checkpoint of our best performing UNet model</li> <li>stilted-armadillo-99-great-gan-epoch=7-step=3024.ckpt contains the PyTorch model checkpoint of our best performing GAN model</li> </ul> <p><strong>Forecasts</strong></p> <ul> <li>radiant-sponge-59-great-unet.nc contains a forecast generated by our best performing UNet model for the test set beginning February 2018 and ending June 2019</li> <li>stilted-armadillo-99-great-gan.nc contains a forecast generated by our best performing GAN model for the test set beginning February 2018 and ending June 2019</li> </ul> <p>Why is the GAN forecast file so much larger than the UNet forecast file? The GAN forecast file is an ensemble forecast comprising 25 ensemble members, with a mean forecast cached as a 26th member. Hence, it is 26x the size of the deterministic UNet forecast.</p>
Data and code for gmd-2023-113 "Parameter estimation for ocean background vertical diffusivity coefficients in the Community Earth System Model (v1.2.1) and its impact on ENSO forecast"
<p>Data and code for the paper "Parameter estimation for ocean background vertical diffusivity coefficients in the Community Earth System Model (v1.2.1) and its impact on ENSO forecast"</p> <p>includes: </p> <p>The model is Community Earth System Model (v1.2.1) (provided by www.cesm.ucar.edu)</p> <p>Data assimilation code is initially provided by Data Assimilation Research Testbed (DART) (https://dart.ucar.edu/), some modifications are made to enable parameter estimation function of ocean background vertical diffusivity coefficients. And the programs and scripts for deal with OISST and EN4 profiles are also developed.</p> <p>The parameter sensitivity experiment results are saved as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/sensitive2008-2012.nc">sensitive2008-2012.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/sensitive2008-2012salt.nc">sensitive2008-2012salt.nc</a> for temperature and salinity, respectively. And the python script to draw the results is </p> <p>The state estimation results are provided as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/Temp_05-17.nc">Temp_05-17.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/Temp_05-17.nc">Salt_05-17.nc</a> for temperature and salinity, respectively.</p> <p>The parameter estimation results are provided as <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/PE_Temp_05-17.nc">PE_Temp_05-17.nc</a> and <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/PE_Temp_05-17.nc">PE_Salt_05-17.nc</a> for temperature and salinity, respectively.</p> <p>the estimated paremeter ensemble is saved in <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/parameters.nc">parameters.nc</a></p> <p>the python script for comparing the SE and PE results is <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/plot_analysis.py">plot_analysis.py</a></p> <p>the nino3.4 indices computed by the forecast experiment is saved in <a href="https://zenodo.org/api/files/6e5b4bfa-61cf-44bc-8e15-e62fab467ea6/fcst_correlation.nc">fcst_correlation.nc</a></p> <p> </p>
Macrosystems EDDIE Module 5 version 2: Introduction to Ecological Forecasting (Instructor Materials)
Ecological forecasting is a tool that can be used for understanding and predicting changes in populations, communities, and ecosystems. Ecological forecasting is an emerging approach which provides an estimate of the future state of an ecological system with uncertainty, allowing society to prepare for changes in important ecosystem services. Ecological forecasters develop and update forecasts using the iterative forecasting cycle, in which they make a hypothesis of how an ecological system works; embed their hypothesis in a model; and use the model to make a forecast of future conditions. When observations become available, they can assess the accuracy of their forecast, which indicates if their hypothesis is supported or needs to be updated before the next forecast is generated. In this Macrosystems EDDIE (Environmental Data-Driven Inquiry & Exploration) module, students will apply the iterative forecasting cycle to develop an ecological forecast for a National Ecological Observation Network (NEON) site. Students will use NEON data to build an ecological model that predicts primary productivity. Using their calibrated model, they will learn about the different components of a forecast with uncertainty and compare productivity forecasts among NEON sites. The overarching goal of this module is for students to learn fundamental concepts about ecological forecasting and build a forecast for a NEON site. Students will work with an R Shiny interface to visualize data, build a model, generate a forecast with uncertainty, and then compare the forecast with observations. The A-B-C structure of this module makes it flexible and adaptable to a range of student levels and course structures. This EDI data package contains instructional materials necessary to teach the module. Intructional materials (instructor manual, introductory presentation for the module, and a presentation to introduce students and instructors to R Shiny) are provided in both pdf and editable formats with
Datasets used in "Verifying Operational Forecasts of Land-Sea Breeze and Boundary Layer Mixing Processes"
<p>Zip file containing datasets used in "Verifying Operational Forecasts of Land-Sea Breeze and Boundary Layer Mixing Processes". In particular, wind magnitude and direction data from automatic weather stations, the Australian Burea of Meteorology's official edited forecast, and unedited ACCESS and ECMWF model data. Data in NETCDF format. </p>
GNSS tomography data for assimilation into the Weather Research and Forecasting model
<p>The data set contains GNSS troposphere tomography estimations of 3D wet refractivity fields for a part of Central Europe (mostly Germany and Czech Republic), for the period of 29 May–14 June 2013 when heavy-precipitation events were observed. The refractivity fields were estimated using two different GNSS tomography models: ATom software package (https://github.com/GregorMoeller/ATom) developed at TU Wien, and the TOMO2 model (Rohm and Bosy, 2011; Rohm et al., 2014; Trzcina and Rohm, 2019) developed at the Wrocław University of Environmental and Life Sciences. Further description of the GNSS tomography processing can be found in the paper by Hanna et al. (2019).</p>
Participant Notes from Chapman Conference on Scientific Challenges Pertaining to Space Weather Forecasting Including Extremes
<p>Compilation of electronic meeting notes made by attendees at the Chapman Conference on Scientific Challenges Pertaining to Space Weather Forecasting Including Extremes.</p> <p>Files are provided for Days 1-3 of the meeting. Day 4 inputs are included in Discussion notes under a separate doi.</p> <p>The Chapman Conference was supported by NSF Award AGS 1848885 and NASA grants 936723.02.01.09.14 and 936723.02.01.11.21</p>
Verifying Spatial Structure in Ensembles of Forecast Fields
<p>Simulation results and GEFS data used in Jacobson et al. (2020). Please see the accompanying code <a href="https://github.com/joshhjacobson">repository</a> for further detail.</p>
Time Series used in the Forecasting Benchmark
<p>This data set contains the time series used in Libra (GitHub: <a href="https://github.com/DescartesResearch/ForecastBenchmark">https://github.com/DescartesResearch/ForecastBenchmark</a> ; CodeOcean: <a href="https://doi.org/10.24433/CO.3240518.v1">https://doi.org/10.24433/CO.3240518.v1</a>). Libra is a forecasting benchmark that automatically evaluates and ranks forecasting methods based on their performance in a diverse set of evaluation scenarios. The benchmark comprises four different use cases, each covering 100 heterogeneous time series taken from different domains.</p>
Non-Poissonian Forecast and Hazard source files - New Zealand National Seismic Hazard Model 2022
<h3>This repository contains:</h3><ul><li>The forecast's files for the Distributed Seismicity Model of the NZNSHM2022, as well as figures, and the Paraview files to explore them in the software interactively. (https://www.paraview.org/)</li><li>The Openquake source files (https://github.com/gem/oq-engine) to run the NZ-NSHM2022 model using the non-Poisson forecasts as single branches.</li></ul><h3>Installation instructions</h3><p>For reproducibility, this package should install OpenQuake (https://github.com/gem/oq-engine) in its version v3.16.4. However, Openquake should remain backward compatible for the Negative Binomial formulation in the future. To install the version 3.16.4, a virtual environment can be created used Anaconda/Miniconda/Micromamba (the latter is recommended, see installation instructions https://mamba.readthedocs.io/en/latest/installation.html) by using:</p><blockquote><p><i>conda env create -f environment.yml</i></p></blockquote><p>This environment should already contain the Openquake version. If the Openquake software should be installed manually into an environment created by the user:</p><blockquote><p><i>source activate {user_env}</i></p><p><i>git clone https://github.com/gem/oq-engine --depth=1 --branch=v3.16.4</i></p><p><i>cd oq-engine</i></p><p>pip install -e .</p></blockquote><p>For additional information, please see the README.md file, or visit <a href="https://github.com/pabloitu/nz_nshm2022_nonpoisson">https://github.com/pabloitu/nz_nshm2022_nonpoisson</a></p>
Ecological forecasts for marine resource management during climate extremes
<p><span>Forecasting weather has become commonplace, but as society faces novel and uncertain environmental conditions there is a critical need to forecast ecology. Forewarning of ecosystem conditions during climate extremes can support proactive decision-making, yet applications of ecological forecasts are still limited. We showcase the capacity for existing marine management tools to transition to a forecasting configuration and provide skilful ecological forecasts up to 12 months in advance. The management tools use ocean temperature anomalies to help mitigate whale entanglements and sea turtle bycatch, and we show that forecasts can forewarn of human-wildlife interactions caused by unprecedented climate extremes. <span>We further show that regionally downscaled forecasts are not a necessity for ecological forecasting and can be less skilful than global forecasts if they have fewer ensemble members.</span> Our results highlight capacity for ecological forecasts to be explored for regions without the infrastructure or capacity to regionally downscale, ultimately helping to improve marine resource management and climate adaptation globally.</span></p>
Generative deep learning for hydrological forecasting: CVAE-75 basins from CANOPEX_v1
<p>Data associated with https://doi.org/10.1016/j.jhydrol.2023.130498</p>
Dataset for monkeys A and B from AMAG: Additive, Multiplicative and Adaptive Graph Neural Network For Forecasting Neuron Activity
<p>ECoG data from two monkeys, affi (A) and beignet (B) used in AMAG: Additive, Multiplicative and Adaptive Graph Neural Network For Forecasting Neuron Activity. Jingyuan Li, Leo Scholl, Trung Le, Pavithra Rajeswaran, Amy L Orsborn, and Eli Shlizerman. NeurIPS. 2023. https://openreview.net/forum?id=7ntI4kcoqG</p><p>See also https://github.com/shlizee/AMAG</p>
Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation
<p><strong>Source codes and dataset of the research "Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation".</strong></p><p>Our work employed PyTorch, a framework for training Deep Learning models with GPU support and automatic back-propagation, to load the MViTv2 s models with Kinetics-400 weights. To simplify the code implementation, eliminating the need for an explicit loop to train and the automation of some hyperparameters, we use the PyTorch Lightning module. The inputs were batches of 10 samples with 16 sequenced images in 3-channel resized to 224 × 224 pixels and normalized from 0 to 1.</p><p>Most of the papers in our literature survey split the original dataset chronologically. Some authors also apply k-fold cross-validation to emphasize the evaluation of the model stability. However, we adopt a hybrid split taking the first 50,000 to apply the 5-fold cross-validation between the training and validation sets (known data), with 40,000 samples for training and 10,000 for validation. Thus, we can evaluate performance and stability by analyzing the mean and standard deviation of all trained models in the test set, composed of the last 9,834 samples, preserving the chronological order (simulating unknown data).</p><p>We develop three distinct models to evaluate the impact of oversampling magnetogram sequences through the dataset. The first model, Solar Flare MViT (SF MViT), has trained only with the original data from our base dataset without using oversampling. In the second model, Solar Flare MViT over Train (SF MViT oT), we only apply oversampling on training data, maintaining the original validation dataset. In the third model, Solar Flare MViT over Train and Validation (SF MViT oTV), we apply oversampling in both training and validation sets.</p><p>We also trained a model oversampling the entire dataset. We called it the "SF_MViT_oTV Test" to verify how resampling or adopting a test set with unreal data may bias the results positively.</p><p><strong>GitHub version</strong></p><p>The .zip hosted here contains all files from the project, including the checkpoint and the output files generated by the codes. We have a clean version hosted on GitHub (<a href="https://github.com/lfgrim/SFF_MagSeq_MViTs">https://github.com/lfgrim/SFF_MagSeq_MViTs</a>), without the magnetogram_jpg folder (which can be downloaded directly on <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip)</a> and the output and checkpoint files. Most code files hosted here also contain comments on the Portuguese language, which are being updated to English in the GitHub version.</p><p><strong>Folders Structure</strong></p><p>In the Root directory of the project, we have two folders: </p><ul><li>magnetogram_jpg: holds the source images provided by Space Environment Artificial Intelligence Early Warning Innovation Workshop through the link <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip. </a>It comprises 73,810 samples of high-quality magnetograms captured by HMI/SDO from 2010 May 4 to 2019 January 26. The HMI instrument provides these data (stored in hmi.sharp_720s dataset), making new samples available every 12 minutes. However, the images from this dataset were collected every 96 minutes. Each image has an associated magnetogram comprising a ready-made snippet of one or most solar ARs. It is essential to notice that the magnetograms cropped by SHARP can contain one or more solar ARs classified by the National Oceanic and Atmospheric Administration (NOAA).</li><li>Seq_Magnetogram: contains the references for source images with the corresponding labels in the next 24 h. and 48 h. in the respectively M24 and M48 sub-folders.<ul><li>M24/M48: both present the following sub-folders structure:<ul><li>Seqs16;</li><li>SF_MViT;</li><li>SF_MViT_oT;</li><li>SF_MViT_oTV;</li><li>SF_MViT_oTV_Test.</li></ul></li></ul></li></ul><p>There are also two files in root:</p><ul><li>inst_packages.sh: install the packages and dependencies to run the models.</li><li>download_MViTS.py: download the pre-trained MViTv2_S from PyTorch and store it in the cache.</li></ul><p>M24 and M48 folders hold reference text files (flare_Mclass...) linking the images in the magnetogram_jpg folders or the sequences (Seq16_flare_Mclass...) in the Seqs16 folders with their respective labels. They also hold "cria_seqs.py" which was responsible for creating the sequences and "test_pandas.py" to verify head info and check the number of samples categorized by the label of the text files. All the text files with the prefix "Seq16" and inside the Seqs16 folder were created by "criaseqs.py" code based on the correspondent "flare_Mclass" prefixed text files.</p><p>Seqs16 folder holds reference text files, in which each file contains a sequence of images that was pointed to the magnetogram_jpg folders.</p><p>All SF_MViT... folders hold the model training codes itself (SF_MViT...py) and the corresponding job submission (jobMViT...), temporary input (Seq16_flare...), output (saida_MVIT... and MViT_S...), error (err_MViT...) and checkpoint files (sample-FLARE...ckpt). Executed model training codes generate output, error, and checkpoint files. There is also a folder called "lightning_logs" that stores logs of trained models.</p><p><strong>Naming pattern for the files:</strong></p><ul><li>magnetogram_jpg: follows the format<i> </i>"hmi.sharp_720s.<SHARP-ID>.<date>.magnetogram.fits.jpg" and</li><li>Seqs16: follows the format "hmi.sharp_720s.<i><</i>SHARP-ID<i>></i>.<init-date>.to.<end-date>", where:<ul><li>hmi: is the instrument that captured the image</li><li>sharp_720s: is the database source of SDO/HMI.</li><li><SHARP-ID>: is the identification of SHARP region, and can contain one or more solar ARs classified by the (NOAA).</li><li><date>: is the date-time the instrument captured the image in the format yyyymmdd_hhnnss_TAI (y:year, m:month, d:day, h:hours, n:minutes, s:seconds).</li><li><init-date>: is the date-time when the sequence starts, and follow the same format of <date>.</li><li><end-date>: is the date-time when the sequence ends, and follow the same format of <date>.</li></ul></li><li>Reference text files in M24 and M48 or inside SF_MViT... folders follows the format "<prefix>flare_Mclass_<forecasting-horizon>_<dataset>.txt<over>", where:<ul><li><prefix>: is Seq16 if refers to a sequence, or void if refers direct to images.</li><li><forecasting-horizon>: "24h" or "48h".</li><li><dataset>: is "TrainVal<n>" or "Test". The <n> refers to the split of Train/Val.</li><li><over>: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>All SF_MViT...folders:<ul><li>Model training codes: "SF_MViT_<oversampling-type>_M+_<forecasting-horizon>_<split-type><gpu-type>", where:<ul><li><oversampling -type>: void or "oT" (over Train) or "oTV" (over Train and Val) or "oTV_Test" (over Train, Val and Test);</li><li><forecasting-horizon>: "24h" or "48h";</li><li><split-type>: "oneSplit" for a specific split or "allSplits" if run all splits.</li><li><gpu-type>: void is default to run 1 GPU or "2gpu" to run into 2 gpus systems;</li></ul></li><li>Job submission files: "jobMViT_<queue>", where:<ul><li><queue>: point the queue in Lovelace environment hosted on CENAPAD-SP (<a href="https://www.cenapad.unicamp.br/parque/jobsLovelace">https://www.cenapad.unicamp.br/parque/jobsLovelace</a>)</li></ul></li><li>Temporary inputs: "Seq16_flare_Mclass_<forecasting-horizon>_<dataset>.txt<over>:<ul><li><dataset>: train or val;</li><li><over>: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>Outputs: "saida_MViT_Adam_10-7<split>", where:<ul><li><split>: k0 to k4, means the correlated split of the output, or void if the output is from all splits.</li></ul></li><li>Error files: "err_MViT_Adam_10-7<split>", where:<ul><li><split>: k0 to k4, means the correlated split of the error log file, or void if the error file is from all splits.</li></ul></li><li>Checkpoint files: "sample-FLARE_MViT_S_10-7-epoch=<n-epoch>-valid_loss=<loss-value>-Wloss_k=<n-split>.ckpt", where:<ul><li><n-opoch>: epoch number of the checkpoint;</li><li><loss-value>: corresponding valid loss;</li><li><n-split>: 0 to 4.</li></ul></li></ul></li></ul>
Duferco forecast results using Kernel Ridge
<p>The data uploaded represent thirteen months of Duferco forecast (June 2021 - June 2022) using Kernel Ridge regression with hourly granularity for the Calabria wind farm.</p> <p>The three csv files represents the three different methods developed and tested during the project:</p> <ol> <li><strong>BL:</strong> Kernel Ridge regression with gaussian kernel (Baseline).</li> <li><strong>PC:</strong> The Baseline method with the pre-processing of the data using the power curve.</li> <li><strong>PC+K: </strong>The Baseline method with the pre-processing of the data using the power curve, and the post-processing of the forecast using Kalman smoothing filter.</li> </ol> <p> </p>
Forecasting Adversarial Actions Using Judgment Decomposition-Recomposition
<p>This repository contains the data and source code used in the research paper, "Forecasting Adversarial Actions Using Judgment Decomposition-Recomposition".</p>
Public IST:Forecasting input files
<p>Input files for Galaxy Clustering and Weak Lensing to reproduce the forecasts of the Euclid IST:Forecasting.</p>
CoronaCast 2024: 3-Day Forecast
<p>These are the model output data for the 3-day forecast of the total solar eclipse in April 8, 2024, using the Space Weather Modeling Framework (SWMF) at the University of Michigan. See README.txt for information on dataset contents, formats, and suggested software libraries.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.