Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
815
datasets available to search
ShareScore release 0.7.1
Dataset results
815 results for “Forecasting”
Traffic forecasting with Virtual Induction Loops - SUMO simulation dataset
<p>This repository is associated with my doctoral dissertation titled, "<strong>Smartphone based applications for Road Traffic Telematics</strong>". In particular this repository serves as the basis of Chapter 7 titled, "<strong>Traffic forecasting with Virtual Induction Loops (VIL)</strong>". The basic idea is to validate a traffic forecasting system which uses machine learning techniques on the simulation of real traffic flows on a real intersection in the City of Turin. This dataset contains simulation output from SUMO software for 56 real days between the months of October-2017 to April-2018. Details about these days are available in my thesis. For each day, 3 output files are available. Here is the description and naming convention:</p> <ol> <li>M1_100seed_100pr_dump.csv (This is the data dump file from SUMO. It contains flows of every single vehicle that was simulated. Naming convention is day_seed_vilPenetrationRate_dump.csv)</li> <li>M1_100seed_ilNorth_100pr.xml (This is the output from a simulated induction loop for Northbound traffic. Naming convention is day_seed_ilNorth_vilPenetrationRate.xml)</li> <li>M1_100seed_ilSouth_100pr.xml (This is the output from a simulated induction loop for Southbound traffic. Naming convention is day_seed_ilSouth_vilPenetrationRate.xml)</li> </ol> <p>For further details, please refer to my thesis.</p>
The Effect of QPF on Real-time Deterministic Hydrologic Forecast Uncertainty
<p>The use of Quantitative Precipitation Forecast (QPF) in hydrologic forecasting is commonplace, but QPF is subject to considerable error. When QPF is included as a model forcing in the hydrological forecast process, significant error is passed to subsequent hydrologic predictions. Two questions arise: (1) are the resulting observed hydrologic forecast errors sufficiently large to suggest the use of zero QPF in the forecast process; if the use of QPF is indicated, (2) how many periods (hours) of QPF (1-, 6-, 12-,..., 72-h...) should be used? Also, do forecast conditions exist under which the use of QPF should be different? This study presents results from two real-time hydrologic forecast experiments, focused on the NOAA/NWS Ohio River Forecast Center (OHRFC). The experiments rely on forecasts from subbasins at 38 forecast point locations, ranging in drainage area, geographic location within the Ohio River Valley, and watershed response time. Results from an experiment, spanning all flow ranges, for the August 10, 2007 - August 31, 2009 period, show that non-zero QPF produces smaller hydrologic forecast error than zero QPF. A second experiment, January 23, 2009 through September 15, 2010, suggests that QPF should be limited to 6- to 12-h duration for flood forecasts. Beyond 12-h, hydrologic forecast error increases substantially across all forecast ranges, but errors are much larger for flood forecasts. Increased durations of QPF produce smaller forecast error than shorter QPF durations only for non-flood forecasts. Experimental results are shown to be consistent with NWS, April 2001 to October 2016, forecast verification statistics for the OHRFC.</p>
Virtual Screening with Molecular Forecaster
<p>Commercially available compounds for USP5 zinc-finger ubiquitin binding domain (ZnF-UBD) were identified with Molecular Forecasters (MFI) FITTED docking platform. Preliminary assessment of docking for USP5 ZnF-UBD with FITTED can be found <a href="https://zenodo.org/record/2620208#.XQO2pYhKjIV">here</a>.</p>
Data for 'Improved predictability of the Indian Ocean Dipole using seasonally modulated ENSO forcing forecasts'
<p>Abstract of the associated paper: Despite recent progress in seasonal forecast development, the predictive skill for the Indian Ocean Dipole (IOD) remains typically limited to a lead time of one season or less in both dynamical and empirical models. Here we develop a simple stochastic-dynamical model (SDM) to predict the IOD using seasonally modulated El Niño-Southern Oscillation (ENSO) forcing together with a seasonal modulation of the Indian Ocean coupled ocean-atmosphere feedback. The SDM, with either observed or forecasted ENSO forcing, exhibits generally higher skill and longer lead times for predicting IOD events than the operational Climate Forecast System Version 2 and the SINTEX system. These results affirm our hypothesis that operational IOD predictability beyond persistence is largely controlled by ENSO predictability and the signal-to-noise ratio of the system. Therefore, potential future ENSO improvements in models should also translate to more skillful IOD predictions.</p>
Figure 4 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 4. Mean suitability for Arborophila crudigularis under baseline climate conditions and future climate scenarios. cccma and csiro represent two general circulation models; RCP2.6, and RCP8.5 represent two greenhouse gas emission scenarios; EN is entire suitable habitat; PR is presence records.
Figure 2 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 2. Performance of each model for predicting the suitable habitat for Arborophila crudigularis. GLM: Generalized linear model; GBM: generalized boosting model; GAM: generalized additive model; CTA: classification tree analysis; ANN: artificial neural network; FDA: flexible discriminant analysis; MARS: multiple adaptive regression splines; RF: random forest; MAXENT: maximum entropy model.
Figure 6 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 6. Changes in suitable habitat for Arborophila crudigularis under the RCP8.5 emission scenario. cccma and csiro represent two general circulation models.
Figure 5 in The potential effects of future climate change on suitable habitat for the Taiwan partridge (Arborophila crudigularis): an ensemble-based forecasting method
Figure 5. Changes in suitable habitat for Arborophila crudigularis under the RCP2.6 emission scenario. cccma and csiro represent two general circulation models.
Fig. 4 in Forecasting the impact of an invasive macrophyte species in the littoral zone through aquatic insect species composition
Fig. 4. Comparison among Bray-Curtis dissimilarity indices of aquatic insect assemblages associated with white ginger lily banks and native vegetation profile in the littoral zone of a tropical reservoir in the Brazilian Savanna (Group 1, white ginger lily; Group 2, invaded forest; Group 3, native macrophyte; Group 4, riparian vegetation).
Fig. 2 in Forecasting the impact of an invasive macrophyte species in the littoral zone through aquatic insect species composition
Fig. 2. Comparison between ecological variables of aquatic insect assemblages associated with invasive white ginger lily bank and other native vegetation banks in the littoral zone of a tropical reservoir in the Brazilian Savanna (A, abundance; B, richness; C, Simpson diversity; IM, invasive macrophyte; IF, invaded forest; NM, native macrophyte; RV, riparian vegetation).
Fig. 1 in Forecasting the impact of an invasive macrophyte species in the littoral zone through aquatic insect species composition
Fig. 1. Location and characterization of vegetation profile banks of the Fazzari reservoir in the Brazilian Savanna (Cerrado Biome, Brazil).
Fig. 3 in Forecasting the impact of an invasive macrophyte species in the littoral zone through aquatic insect species composition
Fig. 3. Analyses of non-metric MDS of aquatic insect assemblages associated with white ginger lilY banks and native vegetation profiles in the littoral zone of a tropical reservoir in the Brazilian Savanna (●, white ginger lilY; ○, invaded forest; ∆, native macrohYte; ▲, riparian vegetation).
Enhanced localized Temperature and Dewpoint forecast for Thessaloniki region
<p>2D Dataset of enhanced localized forcast for 2m Air Temperature and Dewpoint for the Thessaloniki region. The forecast has an hourly resolution and is is up to 72 hours. The compressed region is located at the city Thessaloniki with a little surrounding suburbs. The spacial resolution is 4x4km and comes from the SuperHD wheatehr forecast model at Meteologix.com. <br>The DMO model output was enhanced with statistical information of up to 10 weather stations located in Thessaloniki. Individual MOS-model forecasts for these 10 measurement stations where merge into the DMO model output. </p>
Forecasting the CBOE VIX and SKEW Indices Using Heterogeneous Autoregressive Models
<p>This dataset is used in the paper "Forecasting the CBOE VIX and SKEW Indices Using Heterogeneous Autoregressive Models" by Massimo Guidolin and Giulia F. Panzeri. The data contains daily observations of the VIX, SKEW, and SKEW− indices over the period January 4, 1996, - December 31, 2019. This period covers a total of 6,005 daily observations. The dataset is constructed from multiple sources as described in the paper and includes several files that correspond to different transformations and forecast errors of the indices.</p>
supplymentary for sediment load forecasting
<p>Forecasting results of sediment load from stations along the mainstream and different tributaries; Bayesian inference code to forecast sediment load</p>
MACHINE LEARNING APPROACHES FOR DEMAND FORECASTING: THE IMPACT OF CUSTOMER SATISFACTION ON PREDICTION ACCURACY
<p><span>This study investigates the effectiveness of various machine learning models in predicting product demand based on customer satisfaction data. Four models—Linear Regression, Random Forest, Gradient Boosting, and Support Vector Machine (SVM)—were evaluated using performance metrics, including Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² score. The results indicate that Gradient Boosting achieved the highest accuracy, with an MAE of 2.56, MSE of 12.75, RMSE of 3.57, and R² score of 0.82, effectively capturing the complex, non-linear relationships inherent in customer satisfaction factors. Random Forest also demonstrated strong performance, while Linear Regression and SVM showed limitations in handling intricate datasets. These findings underscore the importance of utilizing advanced machine learning techniques for accurate demand forecasting, highlighting the critical role of customer satisfaction data in enhancing predictive capabilities. The insights gained from this research can guide organizations in optimizing inventory management and improving customer satisfaction in a rapidly evolving market.</span></p>
Supplementary data for "Deep learning for industrial processes: Forecasting amine emissions from a carbon capture plant"
<p>A preliminary analysis of the data already has been discussed in <a href="https://dx.doi.org/10.2139/ssrn.3812299">10.2139/ssrn.3812299</a>.</p> <p><strong>Raw data</strong></p> <p>Raw measurement data is in the Excel files `day*_raw.xlsx`.</p> <p><strong>Model</strong></p> <p>Covariate and label scaler objects are serialized in joblib format in the following files:</p> <ul> <li>20210812_y_transformer_co2_ammonia_reduced_feature_set</li> <li>20210812_y_transformer__reduced_feature_set</li> <li>20210812_x_scaler_reduced_feature_set</li> </ul> <p>Checkpoints of the models are in the `*.pth.tar` files. An example for loading the models is:</p> <pre><code class="language-python">from pyprocessta.model.tcn import TCNModelDropout model_cov = TCNModelDropout( input_chunk_length=8, output_chunk_length=1, num_layers=5, num_filters=16, kernel_size=6, dropout=0.3, weight_norm=True, batch_size=32, n_epochs=100, log_tensorboard=True, optimizer_kwargs={"lr": 2e-4}, ) model_cov.load_from_checkpoint('20210814_2amp_pip_model_reduced_feature_set_darts')</code></pre> <p>which assumes that the checkpoints are placed as `model_best.pth.tar` in a folder called `20210812_2amp_pip_model_reduced_feature_set_darts`.</p> <p> </p>
Datasets with weather forecasts (Temperature, Wind Direction, Humidity, Pressure, Wind Speed, GHI)
<p>Datasets with weather forecasts for HLU7 (Temperature, Wind Direction, Humidity, Pressure, Wind Speed, GHI) of the CROSSBOW project.</p>
Rye microgrid historical weather forecasts and stochastic scenarios
<p>This datasets connects historical weather forecasts from the Norwegian Meteorological Institute (met.no) and historical observations from Rye microgrid (https://doi.org/10.5281/zenodo.4448894).</p> <p>Each csv file represents a historical weather forecast for approximately 60 hours ahead. Each csv-file also contains the corresponding observations in the same time interval. Finally, the files also contain load, wind generation and solar PV generation predicitons.</p> <p>The predictions are generated using gradient boosting. The predictions ending with "_ls" are based on least square. The predictions ending with "_quantile_i" represent a quantile prediction. For example, "wind_quantile_2" means that there is a 20% probability the wind will be less than this value.</p> <p>The gradient boosting predictor has been trained to predict the wind power, solar power and load using the explanatory variables below:</p> <p>Solar PV: Cloud area fraction, initial production, clear sky production and forecast look-ahead time</p> <p>Wind power: wind speed, wind direction, wind power converted from wind speed forecast, initial production and forecast look-ahead time</p> <p>Load: hour of day, month of year</p>
Spatial Damped Anomaly Persistence (SDAP) Forecasts of Sea Ice Presence in the Antarctic between 1999 and 2020
<p>Spatial Damped Anomaly Persistence Forecasts of Sea Ice Presence in the southern hemisphere between 1999 and 2020. Each netcdf file corresponds to a single initialisation, done at the start of the stated month, and the forecasts for the following 120 days, both probabilistic (SDAP) and deterministic (SAP) forecasts. The forecasts were derived using OSI SAF sea-ice concentration records (OSI SAF 450 and 430b) and follow the resolution of that dataset (25 km EASE-2 grid). Further details regarding the forecasting method and the results can be found in Niraula et Goessling, 2021 (in review).</p> <p> </p> <p>Please note that while the filenames say "DampedForecast", each file contains both Damped or Deterministic forecasts associated with the date.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.