Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
815
datasets available to search
ShareScore release 0.7.1
Dataset results
815 results for “Forecasting”
Caravan MultiMet (Part 2, Forecasts): Extending Caravan with Multiple Weather Nowcasts and Forecasts
<p>Caravan MultiMet is a novel extension to Caravan, focusing on enriching the meteorological forcing data. Our extension adds three precipitation nowcast products (CPC, IMERG v07 Early, and CHIRPS) and three weather forecast products (ECMWF IFS HRES, GraphCast, and CHIRPS-GEFS). Since all data is kept in it's original time zone (UTC+0) and the ERA5-Land data in the original Caravan data set is shifted to local time of each gauge, we also include ERA5-Land reanalysis data in this extension, matching the UTC-0 timezone for all gauges of the other forcings. This part of the extension includes the <strong>forecast</strong> products.</p> <p>The inclusion of diverse data sources, particularly weather forecasts, enables more robust evaluation and benchmarking of hydrological models, especially for real-time forecasting scenarios. To the best of our knowledge, this extension makes Caravan the first open large-sample hydrology dataset to incorporate weather forecast data.</p> <p>The data is also publicly available on Google Cloud Platform (GCP), and we provide below a colab with an example of how to access it which does not require downloading the entire dataset.</p> <p>Additional resources:</p> <ul> <li>The <a href="https://github.com/kratzert/Caravan">Caravan GitHub repository</a> includes further information and links to other extensions.</li> <li>The original <a href="https://www.nature.com/articles/s41597-023-01975-w">Caravan paper</a></li> </ul> <p>------</p> <p>Channel log:</p> <ul> <li>21 November 2024: Version 1.1 - Fixed bug in the FAO Penman-Monteith potential evaporation - values are now clipped to 0 (previously, some values were negative). Only the ERA5-Land dataset was changed.</li> </ul>
Caravan MultiMet (Part 1, Nowcasts): Extending Caravan with Multiple Weather Nowcasts and Forecasts
<p>Caravan MultiMet is a novel extension to Caravan, focusing on enriching the meteorological forcing data. Our extension adds three precipitation nowcast products (CPC, IMERG v07 Early, and CHIRPS) and three weather forecast products (ECMWF IFS HRES, GraphCast, and CHIRPS-GEFS). Since all data is kept in it's original time zone (UTC+0) and the ERA5-Land data in the original Caravan data set is shifted to local time of each gauge, we also include ERA5-Land reanalysis data in this extension, matching the UTC-0 timezone for all gauges of the other forcings. This part of the extension includes the <strong>nowcast</strong> products.</p> <p>The inclusion of diverse data sources, particularly weather forecasts, enables more robust evaluation and benchmarking of hydrological models, especially for real-time forecasting scenarios. To the best of our knowledge, this extension makes Caravan the first open large-sample hydrology dataset to incorporate weather forecast data.</p> <p>The data is also publicly available on Google Cloud Platform (GCP), and we provide below a colab with an example of how to access it which does not require downloading the entire dataset.</p> <p>Additional resources:</p> <ul> <li>The <a href="https://github.com/kratzert/Caravan">Caravan GitHub repository</a> includes further information and links to other extensions.</li> <li>The original <a href="https://www.nature.com/articles/s41597-023-01975-w">Caravan paper</a></li> </ul> <p>------</p> <p>Channel log:</p> <ul> <li>21 November 2024: Version 1.1 - Fixed bug in the FAO Penman-Monteith potential evaporation - values are now clipped to 0 (previously, some values were negative). Only the ERA5-Land dataset was changed.</li> </ul>
Power forecasting literature search
<p>Dataset of metadata of documents returned from a search in Scopus database and saved in .csv format. This dataset is used for a bibliographic analysis (about electricity consumption forecasting) in the PhD thesis of the author.</p> <p>The Scopus search was done on 5th November 2021. The query corresponds to:</p> <p>TITLE-ABS-KEY ( ( *power* OR "*load consumption*" OR "*load forecast*" OR "*load predict*" OR *consumption* ) AND ( *predict* OR *forecast* ) ) AND ( LIMIT-TO ( SUBJAREA , "ENGI" ) OR LIMIT-TO ( SUBJAREA , "ENER" ) OR LIMIT-TO ( SUBJAREA , "COMP" ) )</p> <p>In total, 245421 documents were returned from the search. The dataset contains the metadata that was possible to export and download from these documents.</p>
Outputs of the Jupyter Notebook - Sea ice forecasting using IceNet
<p>The dataset contains the outputs of the notebook "Sea ice forecasting using IceNet" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li>Alejandro Coca-Castro (author), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></li> <li>Tom R. Andersson (reviewer), British Antarctic Survey, <a href="https://github.com/tom-andersson">@tom-andersson</a></li> <li>Nick Barlow (reviewer), The Alan Turing Institute, <a href="https://github.com/nbarlowATI">@nbarlowATI</a></li> </ul> <p><em>Modelling codebase</em></p> <ul> <li>Tom R. Andersson (author), British Antarctic Survey, <a href="https://github.com/tom-andersson">@tom-andersson</a></li> <li>James Byrne (contributor), British Antarctic Survey, <a href="https://github.com/JimCircadian">@JimCircadian</a></li> <li>Tony Phillips (contributor), British Antarctic Survey</li> </ul>
Impedance-based forecasting of battery performance amid uneven usage
<p>Dataset of 88 commercial lithium-ion coin cells cycled under multistage constant current charging/discharging, with currents randomly changed between cycles to emulate realistic use patterns.</p> <p>raw-data.zip contains the following data:</p> <p>Variable Discharge: We subject 24 Powerstream LiR2032 coin cells (of nominal capacity 1C = 35mAh) to a sequence of randomly selected charge and discharge currents at room temperature for 110-120 full charge/discharge cycles. Each cycle consists of acquisition of the galvanostatic EIS spectrum, followed by a charging and discharging stage. We collect impedance measurements at 57 frequencies uniformly distributed in the log domain in the range 0.02Hz-20kHz. Charging consists of a two stage Constant Current (CC) protocol; currents are randomly selected in the ranges 70mA-140mA (2C-4C) and 35mA-105mA (1C-3C) in stages 1 and 2 respectively. If the safety threshold voltage of 4.3V is reached before the time limit then charging is stopped. During discharging, a single constant discharge current, randomly selected in the range 35mA-140mA (1C-4C), is applied, until the voltage drops to 3.0V.</p> <p>Fixed Discharge: We subject an additional 16 Powerstream LiR2032 coin cells (of nominal capacity 1C = 35mAh) to the same cycling conditions as above, except now fixing the discharge current for all cells and cycles at 52.5mA (1.5C) instead of randomly changing the discharge current at each cycle.</p> <p>chemistry2-25C.zip contains the following data:</p> <p>Variable Discharge @ 25C: We subject 32 RS-Pro LiR2032 coin cells (of nominal capacity 1C = 40mAh) to a sequence of randomly selected charge and discharge currents at room temperature for 110-120 full charge/discharge cycles. Each cycle consists of acquisition of the galvanostatic EIS spectrum, followed by a charging and discharging stage. We collect impedance measurements at 57 frequencies uniformly distributed in the log domain in the range 0.02Hz-20kHz. Charging consists of a two stage Constant Current (CC) protocol; currents are randomly selected in the ranges 70mA-140mA (2C-4C) and 35mA-105mA (1C-3C) in stages 1 and 2 respectively. The distributions of currents are varied across different cell batches. If the safety threshold voltage of 4.3V is reached before the time limit then charging is stopped. During discharging, a single constant discharge current, randomly selected in the range 35mA-140mA (1C-4C), is applied, until the voltage drops to 3.0V.</p> <p>Variable Discharge @ 35C: We repeat the experiment conducted above for 16 additional RSPro cells, except that now we cycle the cells at 35C instead of 25C.</p>
Supplement to "Probabilistic load forecasting for the low voltage network: forecast fusion and daily peaks"
<p>This deposit contains the scripts and data used in the research article "Probabilistic load forecasting for the low voltage network: forecast fusion and daily peaks", which proposed a novel method for electricity demand forecasting in low voltage networks.</p> <p>The scripts are written in the form of R markdown and include additional commentary on the methodology. Both input data and the resulting forecast data and evaluation results are provided, though the latter two may be regenerated by running the scripts.</p>
Dataset Global Warming Forecast using Acceleration Factors
<p>The dataset includes results of Global Warming forecast using four methods.</p> <p>The methods include a parabolic trendline of the last 61 years of global warming and cumulated CO2 emissions.</p> <p>Two other methods apply the velocity and the acceleration of global warming and cumulative CO2 emissions.</p> <p>The relation between the global surface temperature change and the change in the cumulative CO2 emissions was determined in previous publications as 0.000745°C/GtCO2.</p> <p>The average result from all four methods for the business as usual CO2 mitigation scenario is 4.4°C (4.1°C -5.0°C).</p> <p>According to this forecast, the global temperature change will reach 1.5°C in 2031 (9 years from now) and 2.0°C in 2047 (25 years from now).</p>
Spatially Aggregated Rain Radar Forecast for the Koeln Weiden, Germany
<p>radar_forecast.csv contains time series data generated by spatial aggregation of rain radar forecasts constructed using robust local optical flow extrapolation.</p>
Plant Phenology Forecasts
<p>These are forecasts of plant phenology for 66 species of plants in North America. The forecasts models are made using data from the National Phenology Network, and climate drivers from the NOAA CFSv2 forecast model and the PRISM Climate Group. Maps from the data are also available at the site http://phenology.naturecast.org.</p> <p> </p> <p> </p>
Dataset: Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements
<p>The dataset presented is the companion data to the Journal of Hydrometeorology publication entitled “Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements.” The data that follows contains everything needed to reproduce the spatial inputs for the meteorological station model run using the Spatial Modeling for Resources Framework (SMRF, Havens et al., 2017).</p> <p> </p> <p>Software versions used:</p> <ul> <li>Image Processing Workbench v2.2.0 (Marks et al., 2017)</li> <li>Spatial Modeling for Resources Framework v0.5.3 (Havens et al., 2019)</li> </ul> <p> </p> <p><strong>NOTE:</strong> Reproducing the spatial inputs will generate 10 netCDF files at ~80GB per file.</p> <p> </p> <p><strong>topo.nc</strong> – Contains multiple static layers that are required to run SMRF and iSnobal. The netCDF layers are:</p> <ul> <li>dem – digital elevation model at 100 meter resolution, aggregated from the 10 meter National Elevation Dataset (Archuleta et al., 2017)</li> <li>mask – basin mask for the Boise River Basin</li> <li>veg_height – vegetation height in meters from the National Land Cover Database (Homer et al., 2015)</li> <li>veg_type – vegetation type from the National Land Cover Database</li> <li>veg_tau – vegetation fractional transmissivity derived from the vegetation type</li> <li>veg_k – vegetation emissivity derived from the vegetation type</li> </ul> <p> </p> <p><strong>maxus.nc</strong> – maximum upwind slope netCDF that contains 72 images for all wind directions in 5 degree increments using the algorithm described in Winstral and Marks (2002)</p> <p> </p> <p><strong>Station data:</strong></p> <ul> <li>Contains hourly meteorological station data downloaded from Mesowest (Horel et al., 2002). Data was cleaned and filtered prior to running SMRF.</li> <li>metadata.csv – metadata for 40 stations</li> <li>air_temp.csv – 38 stations</li> <li>cloud_factor.csv – 7 stations</li> <li>precip.csv – 21 stations</li> <li>vapor_pressure.csv – 19 stations</li> <li>wind_direction.csv – 14 stations</li> <li>wind_speed.csv – 14 stations</li> </ul> <p> </p> <p><strong>smrf_config.ini</strong> – Configuration file needed to reproduce the spatial inputs using SMRF. The paths will need to be changed to reflect the data location.</p>
Temperature and rainfall datasets for the paper "STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting"
<p>This page includes spatiotemporal datasets used in the paper <a href="https://doi.org/10.1016/j.neucom.2020.09.060">STConvS2S: Spatiotemporal Convolutional Sequence to Sequence Network for weather forecasting.</a></p> <p>ARIMA and deep learning models use datasets, as follow:</p> <ul> <li>ARIMA models</li> </ul> <p>baseline-chirps-1981-2019.nc (rainfall data)<br> baseline-ucar-1979-2015.nc (temperature data)</p> <ul> <li>Deep learning models:</li> </ul> <p><em>5-step ahead:</em></p> <p>dataset-chirps-1981-2019-seq5-ystep5.nc (rainfall data)<br> dataset-ucar-1979-2015-seq5-ystep5.nc (temperature data)</p> <p><em>15-step ahead:</em></p> <p>dataset-chirps-1981-2019-seq5-ystep15.nc (rainfall data)<br> dataset-ucar-1979-2015-seq5-ystep15.nc (temperature data)</p>
CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting
<h2><strong>CESNET-TimeSeries24: The dataset for network traffic forecasting and anomaly detection</strong></h2> <p>The dataset called CESNET-TimeSeries24 was collected by long-term monitoring of selected statistical metrics for 40 weeks for each IP address on the ISP network CESNET3 (Czech Education and Science Network). The dataset encompasses network traffic from more than 275,000 active IP addresses, assigned to a wide variety of devices, including office computers, NATs, servers, WiFi routers, honeypots, and video-game consoles found in dormitories. Moreover, the dataset is also rich in network anomaly types since it contains all types of anomalies, ensuring a comprehensive evaluation of anomaly detection methods.<br><br>Last but not least, the CESNET-TimeSeries24 dataset provides traffic time series on institutional and IP subnet levels to cover all possible anomaly detection or forecasting scopes. Overall, the time series dataset was created from the 66 billion IP flows that contain 4 trillion packets that carry approximately 3.7 petabytes of data. The CESNET-TimeSeries24 dataset is a complex real-world dataset that will finally bring insights into the evaluation of forecasting models in real-world environments.<br><br></p> <p>Please cite the usage of our dataset as:</p> <blockquote> <p>Koumar, J., Hynek, K., Čejka, T. <em>et al.</em> CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting. <em>Sci Data</em> <strong>12</strong>, 338 (2025). https://doi.org/10.1038/s41597-025-04603-x<br><br>@Article{cesnettimeseries24,<br> author={Koumar, Josef and Hynek, Karel and {\v{C}}ejka, Tom{\'a}{\v{s}} and {\v{S}}i{\v{s}}ka, Pavel},<br> title={CESNET-TimeSeries24: Time Series Dataset for Network Traffic Anomaly Detection and Forecasting},<br> journal={Scientific Data},<br> year={2025},<br> month={Feb},<br> day={26},<br> volume={12},<br> number={1},<br> pages={338},<br> issn={2052-4463},<br> doi={10.1038/s41597-025-04603-x},<br> url={https://doi.org/10.1038/s41597-025-04603-x}<br>}<br><br></p> </blockquote> <p> </p> <h3>Time series</h3> <p>We create evenly spaced time series for each IP address by aggregating IP flow records into time series datapoints. The created datapoints represent the behavior of IP addresses within a defined time window of 10 minutes. The vector of time-series metrics v_{ip, i} describes the IP address ip in the i-th time window. Thus, IP flows for vector v_{ip, i} are captured in time windows starting at t_i and ending at t_{i+1}. The time series are built from these datapoints. </p> <p>Datapoints created by the aggregation of IP flows contain the following time-series metrics:</p> <ul> <li><strong><em>Simple volumetric metrics:</em></strong> the number of IP flows, the number of packets, and the transmitted data size (i.e. number of bytes)</li> <li><strong><em>Unique volumetric metrics:</em></strong> the number of unique destination IP addresses, the number of unique destination Autonomous System Numbers (ASNs), and the number of unique destination transport layer ports. The aggregation of \textit{Unique volumetric metrics} is memory intensive since all unique values must be stored in an array. We used a server with 41 GB of RAM, which was enough for 10-minute aggregation on the ISP network. </li> <li><strong><em>Ratios metrics:</em></strong> the ratio of UDP/TCP packets, the ratio of UDP/TCP transmitted data size, the direction ratio of packets, and the direction ratio of transmitted data size</li> <li><em><strong>Average metrics:</strong></em> the average flow duration, and the average Time To Live (TTL)</li> </ul> <p> </p> <p><strong>Multiple time aggregation: </strong> The original datapoints in the dataset are aggregated by 10 minutes of network traffic. The size of the aggregation interval influences anomaly detection procedures, mainly the training speed of the detection model. However, the 10-minute intervals can be too short for longitudinal anomaly detection methods. Therefore, we added two more aggregation intervals to the datasets--1 hour and 1 day.</p> <p><strong>Time series of institutions:</strong> We identify 283 institutions inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution's data. </p> <p><strong>Time series of institutional subnets:</strong> We identify 548 institution subnets inside the CESNET3 network. These time series aggregated per each institution ID provide a view of the institution subnet's data. </p> <p> </p> <h3>Data Records</h3> <p>The file hierarchy is described below:</p> <blockquote> <p>cesnet-timeseries24/</p> <p> |- institution_subnets/</p> <p> | |- agg_10_minutes/<id_institution>.csv</p> <p> | |- agg_1_hour/<id_institution>.csv</p> <p> | |- agg_1_day/<id_institution>.csv</p> <p> | |- identifiers.csv</p> <p> |- institutions/</p> <p> | |- agg_10_minutes/<id_institution_subnet>.csv</p> <p> | |- agg_1_hour/<id_institution_subnet>.csv</p> <p> | |- agg_1_day/<id_institution_subnet>.csv</p> <p> | |- identifiers.csv</p> <p> |- ip_addresses_full/</p> <p> | |- agg_10_minutes/<id_ip_folder>/<id_ip>.csv</p> <p> | |- agg_1_hour/<id_ip_folder>/<id_ip>.csv</p> <p> | |- agg_1_day/<id_ip_folder>/<id_ip>.csv</p> <p> | |- identifiers.csv</p> <p> |- ip_addresses_sample/</p> <p> | |- agg_10_minutes/<id_ip>.csv</p> <p> | |- agg_1_hour/<id_ip>.csv</p> <p> | |- agg_1_day/<id_ip>.csv</p> <p> | |- identifiers.csv</p> <p> |- times/</p> <p> | |- times_10_minutes.csv</p> <p> | |- times_1_hour.csv</p> <p> | |- times_1_day.csv</p> <p> |- ids_relationship.csv<br> |- weekends_and_holidays.csv</p> </blockquote> <p>The following list describes time series data fields in CSV files:</p> <ul> <li><strong>id_time: </strong>Unique identifier for each aggregation interval within the time series, used to segment the dataset into specific time periods for analysis.</li> <li><strong>n_flows: </strong>Total number of flows observed in the aggregation interval, indicating the volume of distinct sessions or connections for the IP address.</li> <li><strong>n_packets: </strong>Total number of packets transmitted during the aggregation interval, reflecting the packet-level traffic volume for the IP address.</li> <li><strong>n_bytes: </strong>Total number of bytes transmitted during the aggregation interval, representing the data volume for the IP address.</li> <li><strong>n_dest_ip: </strong>Number of unique destination IP addresses contacted by the IP address during the aggregation interval, showing the diversity of endpoints reached.</li> <li><strong>n_dest_asn: </strong>Number of unique destination Autonomous System Numbers (ASNs) contacted by the IP address during the aggregation interval, indicating the diversity of networks reached.</li> <li><strong>n_dest_port: </strong>Number of unique destination transport layer ports contacted by the IP address during the aggregation interval, representing the variety of services accessed.</li> <li><strong>tcp_udp_ratio_packets: </strong>Ratio of packets sent using TCP versus UDP by the IP address during the aggregation interval, providing insight into the transport protocol usage pattern. This metric belongs to the interval <0, 1> where 1 is when all packets are sent over TCP, and 0 is when all packets are sent over UDP.</li> <li><strong>tcp_udp_ratio_bytes:</strong> Ratio of bytes sent using TCP versus UDP by the IP address during the aggregation interval, highlighting the data volume distribution between protocols. This metric belongs to the interval <0, 1> with same rule as <em>tcp_udp_ratio_packets</em>.</li> <li><strong>dir_ratio_packets: </strong>Ratio of packet directions (inbound versus outbound) for the IP address during the aggregation interval, indicating the balance of traffic flow directions. This metric belongs to the interval <0, 1>, where 1 is when all packets are sent in the outgoing direction from the monitored IP address, and 0 is when all packets are sent in the incoming direction to the monitored IP address.</li> <li><strong>dir_ratio_bytes: </strong>Ratio of byte directions (inbound versus outbound) for the IP address during the aggregation interval, showing the data volume distribution in traffic flows. This metric belongs to the interval <0, 1> with the same rule as <em>dir_ratio_packets</em>.</li> <li><strong>avg_duration: </strong>Average duration of IP flows for the IP address during the aggregation interval, measuring the typical session length.</li> <li><strong>avg_ttl: </strong>Average Time To Live (TTL) of IP flows for the IP address during the aggregation interval, providing insight into the lifespan of packets.</li> </ul> <p>Moreover, the time series created by re-aggregation contains following time series metrics instead of <strong>n_dest_ip</strong>, <strong>n_dest_asn</strong>, and <strong>n_dest_port</strong>:</p> <ul> <li><strong>sum_n_dest_ip: </strong>Sum of numbers of unique destination IP addresses.</li> <li><strong>avg_n_dest_ip: </strong>The average number of unique destination IP addresses.</li> <li><strong>std_n_dest_ip: </strong>Standard deviation of numbers of unique destination IP addresses.</li> <li><strong>sum_n_dest_asn: </strong>Sum of numbers of unique destination ASNs.</li> <li><strong>avg_n_dest_asn: </strong>The average number of unique destination ASNs.</li> <li><strong>std_n_dest_asn: </strong>Standard deviation of numbers of unique destination ASNs)</li> <li><strong>sum_n_dest_port: </strong>Sum of numbers of unique destination transport layer ports.</li> <li><strong>avg_n_dest_port: </strong> The average number of unique destination transport layer ports.</li> <li><strong>std_n_dest_port: </strong>Standard deviation of numbers of unique destination transport layer ports.</li> </ul> <p> </p> <p>Moreover, files <em>identifiers.csv</em> in each dataset type contain IDs of time series that are present in the dataset. Furthermore, the <em>ids_relationship.csv</em> file contains a relationship between IP addresses, Institutions, and institution subnets. The <em>weekends_and_holidays.csv</em> contains information about the non-working days in the Czech Republic.</p>
Supplementary Material for "Probabilistic Forecasting of Regional Net-load with Conditional Extremes and Gridded NWP"
<p>Supplementary material to accompany pre-print of "Probabilistic Forecasting of Regional Net-load with Conditional Extremes and Gridded NWP" by Jethro Browell and Matteo Fasiolo available on on arXiv. This is version 3. The only changes from version 1 & 2 to forecast evaluation (significance testing and additional plots). Future releases are subject to change following revisions of this article.</p>
Mainshock+aftershock forecasts from Regional Earthquake Likelihood Models (RELM) experiment
<p>Contains mainshock+aftershock forecasts produced by various members of the working group for the development of Regional Earthquake Likelihood Models. These forecasts were obtained from the Collaboratory of the Study of Earthquake Predictability (CSEP) testing center hosted by the Southern California Earthquake Center at the University of Southern California.</p> <p>Forecasts are described by the following publications</p> <ol> <li>Helmstetter et al. (2007) with aftershocks</li> <li>Kagan et al. (2007)</li> <li>Shen et al. (2007)</li> <li>Bird & Liu (2007)</li> <li>Ebel et al. (2007) with aftershocks</li> </ol> <p>Forecasts are stored in tab separated value files with the following fields (the first row of data is shown as an example):</p> <pre>LON_0 LON_1 LAT_0 LAT_1 DEPTH_0 DEPTH_1 MAG_0 MAG_1 RATE FLAG -125.4 -125.3 40.1 40.2 0.0 30.0 4.95 5.05 5.8499099999999998e-04 1 </pre> <p>References</p> <p>Bird, P., and Z. Liu (2007). Seismic Hazard Inferred from Tectonics: California, Seismological Research Letters 78 37-48.</p> <p>Ebel, J. E., D. W. Chambers, A. L. Kafka, and J. A. Baglivo (2007). Non-Poissonian Earthquake Clustering and the Hidden Markov Model as Bases for Earthquake Forecasting in California, Seismological Research Letters 78 57-65.</p> <p>Helmstetter, A., Y. Y. Kagan, and D. D. Jackson (2007). High-resolution Time-independent Grid-based Forecast for M >= 5 Earthquakes in California, Seismological Research Letters 78 78-86.</p> <p>Kagan, Y. Y., D. D. Jackson, and Y. Rong (2007). A Testable Five-Year Forecast of Moderate and Large Earthquakes in Southern California Based on Smoothed Seismicity, Seismological Research Letters 78 94-98.</p> <p>Shen, Z.-K., D. D. Jackson, and Y. Y. Kagan (2007). Implications of Geodetic Strain Rate for Future Earthquakes, with a Five-Year Forecast of M5 Earthquakes in Southern California, Seismological Research Letters 78 116-120</p> <p> </p>
Consistency test scores for aftershock+mainshock RELM forecasts
<p><strong>Summary</strong></p> <p>Files are transcribed from Zechar et al. (2013) into comma separated values (csv) files. The consistency test scores are shown in the electronic supplement table S4 and the catalog is found in Table 1 of the main text.</p> <p><strong>Reference</strong></p> <p>Zechar, J. D., D. Schorlemmer, M. J. Werner, M. C. Gerstenberger, D. A. Rhoades, and T. H. Jordan (2013). Regional Earthquake Likelihood Models I: First-Order Results, Bulletin of the Seismological Society of America 103 787-798.</p> <p> </p>
Mainshock+aftershock M4.95+ seismicity forecasts derived from the Regional Earthquake Likelihood Models (RELM) and the multiplicative hybrid earthquake models developed by Rhoades et al. (2014)
<p>Contains six mainshock+aftershock seismicity forecasts developed by the Working Group of the Regional Earthquake Likelihood Models (RELM) experiment, sixteen multiplicative hybrid forecasts created by Rhoades et al. (2014), and the 2011-2020 M4.95+ ANSS earthquake catalog for California. Six additional forecast files are included to properly conduct the comparative tests implemented in the Collaboratory for the Study of Earthquake Predictability (CSEP) testing centre.</p> <p>Forecasts are stored in tab separated value files with the following fields (the first row of data is shown as an example):</p> <pre>LON_0 LON_1 LAT_0 LAT_1 DEPTH_0 DEPTH_1 MAG_0 MAG_1 RATE FLAG -125.4 -125.3 40.1 40.2 0.0 30.0 4.95 5.05 5.8499099999999998e-04 1 </pre> <p>Forecast are described in detail by the following publications:</p> <p>Bird, P., and Z. Liu (2007). Seismic Hazard Inferred from Tectonics: California. Seismological Research Letters, 78(1):37-48.</p> <p>Ebel, J. E., D. W. Chambers, A. L. Kafka, and J. A. Baglivo (2007). Non-Poissonian Earthquake Clustering and the Hidden Markov Model as Bases for Earthquake Forecasting in California. Seismological Research Letters, 78(1): 57-65.</p> <p>Helmstetter, A., Y. Y. Kagan, and D. D. Jackson (2007). High-resolution Time-independent Grid-based Forecast for M >= 5 Earthquakes in California. Seismological Research Letters, 78(1): 78-86.</p> <p>Holliday, J., Chen, C., Tiampo, K., Rundle, J., Turcotte, D., and Donnellan, A. (2007). A RELM earthquake forecast based on pattern informatics. Seismological Research Letters, 78(1):87–93.</p> <p>Kagan, Y. Y., D. D. Jackson, and Y. Rong (2007). A Testable Five-Year Forecast of Moderate and Large Earthquakes in Southern California Based on Smoothed Seismicity. Seismological Research Letters, 78(1): 94-98.</p> <p>Rhoades, D.A., Gerstenberger, M.C., Christophersen, A., Zechar, J.D., Schorlemmer, D., Werner, M.J. and Jordan, T.H., 2014. Regional earthquake likelihood models II: Information gains of multiplicative hybrids. Bulletin of the Seismological Society of America, 104(6):3072-3083.</p> <p>Shen, Z.-K., D. D. Jackson, and Y. Y. Kagan (2007). Implications of Geodetic Strain Rate for Future Earthquakes, with a Five-Year Forecast of M5 Earthquakes in Southern California. Seismological Research Letters, 78(1):116-120.</p> <p>Ward, S. (2007). Methods for evaluating earthquake potential and likelihood in and around California. Seismological Research Letters, 78(1):121–133.</p> <p>Wiemer, S. and Schorlemmer, D. (2007). ALM: An asperity-based likelihood model for California. Seismological Research Letters, 78(1):134–140.</p>
Dataset CO2 Emission per Capita Forecast 2020-2100
<p>The dataset includes Business As Usual (BAU) forecast of the world's global CO2 emissions per capita (CpC) for 2020-2100.</p> <p>The CO2 emission forecast is from the publication “<em>Dataset Global Warming Forecast using Acceleration Factors</em>” [3]. According to this publication, the CO2 emissions without international transport will change from 33,803 MtCO2/y in 2020 to 70,191 MtCO2/y in 2100, a 108% increase.</p> <p>The population forecast applies a parabolic trendline of the last 30 years. According to this calculation, the world population will change from 7,795 million in 2020 to 15,206 million in 2100, a 95% increase.</p> <p>CO2 emissions per capita (CpC) are calculated by dividing the CO2 emissions per year by the population in the same year.</p> <p>The world CpC was 4.3366 tCO2/y,cap in 2020. The CpC forecast for 2100 is 4.6160 tCO2/y,cap, 6.4% increase.</p>
ERA5 Dataset used for sea-ice forecasting with IceNet
<p>Monthly 2 metre temperature and other variables used for running IceNet. The data is from ECMWF ERA5 between January 2019 to December 2021.</p> <p> </p>
Data to reproduce the results: Statistical power of spatial earthquake forecast tests
<p>We provide data needed to reproduce the figures from the publication titled "Statistical power of spatial earthquake forecast tests".</p>
Dataset of Machine Learning forecasted VTEC from paper: Uncertainty Quantification for Machine Learning-based Ionosphere and Space Weather Forecasting
<p>The *csv files contain forecasted one-day-ahead Vertical Total Electron Content (VTEC), consisting of the mean/median VTEC values and the upper and lower VTEC bounds of the 95% confidence intervals of 4 models based on machine learning for test data.</p> <p>The first part of the *csv file name corresponds to the type of model: SE stands for the super-ensemble VTEC model, QGB stands for the quantile gradient boosting VTEC model, BNN1 stands for the Bayesian neural network VTEC model, and BNN2 stands for the Bayesian neural network with negative log-likelihood (NLL) loss VTEC model. The second part of the file name refers to the geographic location of the VTEC points for which the forecast is performed, i.e., 10E70N for 10 degree of longitude and 70 degree of latitude, 10E40N for 10 degree of longitude and 40 degree of latitude, and 10E10N for 10 degree of longitude and 10 degree of latitude. The last part of the file name corresponds to the test year, i.e., year 2017.</p> <p>The SE_*_2017.csv file consists of 14 columns. The index column ("Date-time") is expressed in Coordinated Universal Time (UTC) as YYYY-MM-DD. Columns 1-3 contain the VTEC forecast results of Random Forest (RF) trained on three data subsets; columns 4-6 contain the VTEC forecast results of Adaptive Boosting (AB) trained on three data subsets; columns 7-9 contain the VTEC forecast results of Gradient Boosting (XGBoost) trained on three data subsets. Column 10 ("Mean") represents the mean of columns 1-9, i.e., the ensemble mean; column 11 ("Std") represents the standard deviation of columns 1-9, i.e., the ensemble spread; columns 12 ("UB") and 13 ("LB") contain the upper and lower bounds of the 95% confidence interval of VTEC, respectively; and column 14 contains the Global Ionosphere Maps (GIM) values of CODE, i.e., the ground-truth in this study.</p> <p>The QGB_*_2017.csv file consists of 4 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Median") contains the median VTEC forecast, column 2 ("LB") contains the lower VTEC bound of the 95% confidence interval, column 3 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 4 contains the GIM values of CODE, i.e., the ground-truth in this study.</p> <p>The BNN*_2017.csv file consists of 5 columns. The index column ("Date-time") is expressed in UTC as YYYY-MM-DD. Column 1 ("Mean") contains the mean VTEC forecast, column 2 ("Std") contains the standard deviation, column 3 contains GIM values of CODE, i.e., ground-truth in this study; column 4 ("UB") contains the upper VTEC bound of the 95% confidence interval, and column 5 ("LB") contains the lower VTEC bound of the 95% confidence interval.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>Contact</p> <p>----------------------------------------------------------------------------------------------------------------------------------------</p> <p>If you have any questions regarding these data, please contact:</p> <p>Randa Natras</p> <p>Deutsches Geodätisches Forschungsinstitut (DGFI-TUM)</p> <p>Technical University of Munich</p> <p>Arcisstraße 21</p> <p>80333 München</p> <p>randa.natras@tum.de</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.