Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,118
datasets available to search
ShareScore release 0.9.0
Dataset results
1,118 results for “Time series”
Decadal time series of spatially enhanced relative humidity for Europe at 1000 m resolution (2000 - 2021) derived from ERA5-Land data
<p>Overview:<br> ERA5-Land is a reanalysis dataset providing a consistent view of the evolution of land variables over several decades at an enhanced resolution compared to ERA5. ERA5-Land has been produced by replaying the land component of the ECMWF ERA5 climate reanalysis. Reanalysis combines model data with observations from across the world into a globally complete and consistent dataset using the laws of physics. Reanalysis produces data that goes several decades back in time, providing an accurate description of the climate of the past.</p> <p>Processing steps:<br> The original hourly ERA5-Land air temperature 2 m above ground and dewpoint temperature 2 m data has been spatially enhanced from 0.1 degree to 30 arc seconds (approx. 1000 m) spatial resolution by image fusion with CHELSA data (V1.2) (<a href="https://chelsa-climate.org/">https://chelsa-climate.org/</a>). For each day we used the corresponding monthly long-term average of CHELSA. The aim was to use the fine spatial detail of CHELSA and at the same time preserve the general regional pattern and fine temporal detail of ERA5-Land. The steps included aggregation and enhancement, specifically:<br> 1. spatially aggregate CHELSA to the resolution of ERA5-Land<br> 2. calculate difference of ERA5-Land - aggregated CHELSA<br> 3. interpolate differences with a Gaussian filter to 30 arc seconds<br> 4. add the interpolated differences to CHELSA</p> <p>Subsequently, the temperature time series have been aggregated on a daily basis. From these, daily relative humidity has been calculated for the time period 01/2000 - 07/2021.</p> <p>Relative humidity (rh2m) has been calculated from air temperature 2 m above ground (Ta) and dewpoint temperature 2 m above ground (Td) using the formula for saturated water pressure from Wright (1997):</p> <p><code>maximum water pressure = 611.21 * exp(17.502 * Ta / (240.97 + Ta))</code></p> <p><code>actual water pressure = 611.21 * exp(17.502 * Td / (240.97 + Td))</code></p> <p><code>relative humidity = actual water pressure / maximum water pressure</code></p> <p>The resulting relative humidity has been aggregated to decadal averages. Each month is divided into three decades: the first decade of a month covers days 1-10, the second decade covers days 11-20, and the third decade covers days 21-last day of the month.</p> <p>Resultant values have been converted to represent percent * 10, thus covering a theoretical range of [0, 1000].</p> <p>The data have been reprojected to EU LAEA.</p> <p>File naming scheme (YYYY = year; MM = month; dD = number of decade):<br> <code>ERA5_land_rh2m_avg_decadal_YYYY_MM_dD.tif</code></p> <p>Projection + EPSG code:<br> EU LAEA (EPSG: 3035)</p> <p>Spatial extent:<br> north: 6874000<br> south: -485000<br> west: 869000<br> east: 8712000</p> <p>Spatial resolution:<br> 1000 m</p> <p>Temporal resolution:<br> Decadal</p> <p>Pixel values:<br> Percent * 10 (scaled to Integer; example: value 738 = 73.8 %)</p> <p>Software used:<br> GDAL 3.2.2 and GRASS GIS 8.0.0</p> <p>Original ERA5-Land dataset license:<br> <a href="https://apps.ecmwf.int/datasets/licences/copernicus/">https://apps.ecmwf.int/datasets/licences/copernicus/</a></p> <p>CHELSA climatologies (V1.2):<br> Data used: Karger D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E, Linder, H.P., Kessler, M. (2018): Data from: Climatologies at high resolution for the earth's land surface areas. Dryad digital repository. <a href="http://dx.doi.org/doi:10.5061/dryad.kd1d4">http://dx.doi.org/doi:10.5061/dryad.kd1d4</a><br> Original peer-reviewed publication: Karger, D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E., Linder, P., Kessler, M. (2017): Climatologies at high resolution for the Earth land surface areas. Scientific Data. 4 170122. <a href="https://doi.org/10.1038/sdata.2017.122">https://doi.org/10.1038/sdata.2017.122</a></p> <p>Processed by:<br> mundialis GmbH & Co. KG, Germany (<a href="https://www.mundialis.de/">https://www.mundialis.de/</a>)</p> <p>Reference: Wright, J.M. (1997): Federal meteorological handbook no. 3 (FCM-H3-1997). Office of Federal Coordinator for Meteorological Services and Supporting Research. Washington, DC</p> <p>Data is also available in Latitude-Longitude/WGS84 (EPSG: 4326) projection: <a href="http://https://doi.org/10.5281/zenodo.6147830">https://doi.org/10.5281/zenodo.6147830</a></p>
NH-SWE: Northern Hemisphere Snow Water Equivalent dataset based on in-situ snow depth time series and the regionalisation of the ΔSNOW model
<p>Time series of daily Snow Water Equivalent (SWE) and Snow Density over the Northern Hemisphere, based on in-situ station observations of snow depth converted to SWE using the ΔSNOW model (Winkler et al., 2021) and regionalised parameters. </p> <p>An extensive description of the dataset and the method to generate it can be found in the data descriptor manuscript published in the journal Earth System Science Data: <a href="https://essd.copernicus.org/preprints/essd-2023-31/">https://essd.copernicus.org/articles/15/2577/2023/essd-15-2577-2023</a> </p> <p><strong>Dataset:</strong> A total of 11,0071 time series of modelled SWE and estimated snow density at the point scale, spanning 1950-2022, at daily resolution.<em> "NH-SWE_dataset_MAP.png"</em> shows a Northern Hemisphere map with the location of all stations in the NH-SWE dataset and their elevation in meters. </p> <p><strong>Files: </strong>The dataset is provided in two different formats:</p> <ol> <li>Individual <em>.csv</em> files for each station in the NH-SWE dataset at <em>"NH_SWE_dataset_vector_files.zip"</em></li> <li>Full-dataset <em>.csv </em>matrices with dates as rows and NH-SWE stations as columns at <em>"NH_SWE_dataset_matrix_files.zip"</em></li> </ol> <p><strong>Metadata:<em> </em></strong><em>"NH_SWE_METADATA.csv"</em> Includes information on NH-SWE stations location (ID, country, station name, coordinates, elevation), data source, length of time series, model parameters and the climate variables used to estimate them, and average snow climatology such as average maximum snow depth, average peak SWE and average maximum snow cover duration. More details and units in the <em>"README_fileformats.txt"</em> file. </p> <p><strong>ΔSNOW model parameter regionalisation: </strong>The code to obtain the ΔSNOW model parameters based on climate variables for all the stations in the NH-SWE dataset is shared in<em><strong> </strong>"DeltaSNOW_parameter_regionalisation.zip"</em>. The method is extensively described in the data descriptor manuscript by Fontrodona-Bach et al., (2023) submitted to Earth System Science Data. More details in the <em>"README_regionalisation.txt"</em> file. </p> <p><strong>Data use: </strong>Free, provided adequate citation of both the data descriptor manuscript and the zenodo record. See <em>"README_datausage.txt"</em></p> <p><strong>Version history:</strong><br>v1: Initial upload. The ΔSNOW model regionalisation was missing.<br>v2: Manuscript submission version. Updated dataset and includes the ΔSNOW model regionalisation code.</p> <p><strong>Reported errors:</strong><br>The dataset accidentally contains one station from the Southern Hemisphere (NH-SWE ID 500001), located in Antarctica (Country code AY). <br>The longitude of a few stations exceeds +180 decimal degrees. To obtain the correct value within the [-180,180] decimal degree longitude bounds, the value exceeding +180 needs to be added to -180 degrees (e.g. +181.0 degrees is actually -179.0 degrees).<br>Swedish stations have two different country codes, SE for the ECA&D stations, and SW for the GHCNd stations. <br>Japan country code is "JA" in the metadata, although the official country code should be JP. </p>
Time-series of shoreline change for the Klamath River Littoral Cell (California)
<p>This repository contains 35 years of tidally-corrected shoreline change data at the Klamath River Littoral Cell in northern California. This dataset was used in <em>Warrick et al. 2023, "</em><strong>A Large Sediment Accretion Wave Along a Northern California Littoral Cell</strong>"<em>, </em>to investigate and track the movement of a large sediment wave.</p> <p><em>CoastSat </em>was used to map shoreline changes on Landsat 5, Landsat 7 and Landsat 8 imagery between 1984 and 2022. The <em>Coastsat </em>toolbox is publicly available at https://github.com/kvos/CoastSat and described in <em>Vos et al. 2019, </em><a href="https://doi.org/10.1016/j.envsoft.2019.104528">https://doi.org/10.1016/j.envsoft.2019.104528</a>. The time-series of shoreline change were tidally-corrected along cross-shore transects using tide levels from a global tide model (FES2014) and a satellite-derived estimate of the beach slope (as described in <em>Vos et al. 2020, "Beach slopes from satellite-derived shorelines", </em><a href="https://doi.org/10.1029/2020GL088365">https://doi.org/10.1029/2020GL088365</a><em>)</em>.</p> <p>The data is located in the <em>/shoreline_data</em> folder and structured as follows:</p> <ul> <li>The littoral cell is divided in 4 sections (kmt_01, kmt_02, kmt_03, kmt_04)</li> <li>For each section there is a folder with 4 CSV files: <ul> <li><em>time_series_tidally_corrected.csv</em>: this file contains the tidally-corrected time-series of shoreline change along each transect belonging to the site (e.g. kmt01-000, kmt01-001 etc). This is the final product used for coastal change analyses.</li> <li><em>time_series_raw.csv</em>: this file contains the raw time-series of shoreline change, which have not be tidally-corrected. Note that each image is taken at a different stage of the tide.</li> <li><em>tide_levels_fes2014</em>: this file contains the tide levels at the time of image acquisition extracted from FES2014 (global tide model publicly available on AVISO+).</li> <li><em>transect_coordinates_and_beach_slopes.csv</em>: this file contains the coordinates (in WGS84 lat/lon coordinates) as well as the estimated beach slope for each transect.</li> </ul> </li> </ul> <p>In addition, there are 3 geospatial layers (.GEOJSON) which contain important spatial information. All the geospatial layers are in EPSG:2163 - US National Atlas Equal Area:</p> <ul> <li> <em>Klamath_polygons.geojson</em>: this layer contains the polygons that were used to run CoastSat for each section of the littoral cell.</li> <li><em>Klamath_shorelines.geojson</em>: this layer contains the sandy shorelines that were used to generate the cross-shore transects (also used as reference shorelines in CoastSat).</li> <li><em>transects.geojson</em>: this layer contains the cross-shore transects, which are spaced 100 m alongshore.</li> </ul> <p>Finally, in the<em> /animations</em> folder, there is a clip showing the mapped shorelines on the satellite imagery.</p>
IRIS preprocessed data used in paper "Multi variables time series information bottleneck"
<p>Prprocessed data used in paper "Multi variables time series information bottleneck" with the <a href="https://github.com/DenisUllmann/IB-MTS">GitHub</a> code</p> <p>This dataset is created from a public available dataset of observations performed by IRIS, a NASA small explorer mission developed and operated by LMSAL with mission operations executed at NASA Ames Research Center and major contributions to downlink communications funded by ESA and the Norwegian Space Centre.</p> <p>Multiple Time Series of IRIS level 2 data are available <a href="https://iris.lmsal.com/search/">here</a></p> <p>The selected data was labeled using these definitions:</p> <p>QS: Quiet Sun<br> AR: Active Regions of the Sun<br> FL: Flare</p> <p>A time series is labeled QS when every single time step refer to a quiet sun activity.<br> When a given time series is partially composed of flaring events, the global time series is labeled as FL.</p> <p>The npz file is a numpy (np) compressed data and can be loaded using np.load with allow_pickle=True<br> Loaded data is then a python dict described bellow.</p> <p>Each sample 'data' is a np.ndarray with 2 dimensions: time (various length) and wavelength (length=240 representing a range between 2793.8401Å and 2806.02Å).</p> <p>Each sample is given a 'position' which is a list of length 4:<br> position[1] is a string that gives the name of the event<br> position[4] is a boolean vector that gives the time positionsof the corresponding sample in the original sequence of public IRIS level2 data</p> <p>Data file info :</p> <p>Type: .npz<br> Size: 11.89GB</p> <p>*** Key: 'data_TR_QS'<br> ndarray data of length 2467<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TR_AR'<br> ndarray data of length 1042<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TR_FL'<br> ndarray data of length 1055<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_VAL_QS'<br> ndarray data of length 325<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_VAL_AR'<br> ndarray data of length 1042<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_VAL_FL'<br> ndarray data of length 714<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TE_QS'<br> ndarray data of length 1428<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TE_AR'<br> ndarray data of length 792<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TE_FL'<br> ndarray data of length 356<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_TR'<br> ndarray data of length 4564<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'data_VAL'<br> ndarray data of length 2081<br> containing np.ndarray of shapes ['various', 240]</p> <p><br> *** Key: 'data_TE'<br> ndarray data of length 2576<br> containing np.ndarray of shapes ['various', 240]</p> <p> </p> <p>*** Key: 'position_TR_QS'<br> ndarray data of length 2467<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TR_AR'<br> ndarray data of length 1042<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TR_FL'<br> ndarray data of length 1055<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_VAL_QS'<br> ndarray data of length 325<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_VAL_AR'<br> ndarray data of length 1042<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_VAL_FL'<br> ndarray data of length 714<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TE_QS'<br> ndarray data of length 1428<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TE_AR'<br> ndarray data of length 792<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TE_FL'<br> ndarray data of length 356<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TR'<br> ndarray data of length 4564<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_VAL'<br> ndarray data of length 2081<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p> <p> </p> <p>*** Key: 'position_TE'<br> ndarray data of length 2576<br> containing ndarray data of length 4<br> containing mix of types {'ndarray', 'int', 'str'}</p>
Calcium time series of cortex in a rat model of cortical dysplasia
<p>In vitro Calcium time series of rat (P30) primary motor cortex, were recorder by a CCD camera coupled to stereo-fluoerscence<br> microscope, with a fs = 300ms, following the next sequence: <strong>Basal, <em>Stimulus</em>, Rest.</strong></p> <p>All data is included in a compressed file named <strong>calcium_timeseries.tar.gz.</strong></p> <p>There are two groups of rats: <strong>Control</strong> (control animals), and <strong>BCNU</strong> (experimental animals using the BCNU/carmustine model of cortical dysplasia [1]).</p> <p>Time series are stored in .<strong>csv</strong> files with file names as <strong>R?Pilo-KCl.csv</strong> (where <strong>?</strong> indicates the rat ID). Each of these files holds the two recordings, one for each <em>Stimulus</em>, the first being <em>pilocarpine</em>, followd by <em>KCl</em> used as a control of cellular activity. (pilocarpine, KCl). recording session: The first 150 seconds of these time series correspond to basal activity, followed by 30 s of pilocarpine stimulus, and the rest of spontaneous activity after stimulation, for a total of 15 minutes for each <em>Stimulus</em>. The number of cells recorded varied between animals, as indicated by the number of columns in these .csv files. All of these files have the same number of rows (6000), with each row indicating a frame in the time series. The file <strong>dataEx.png</strong> illustrates this organization.</p> <p>Files named <strong>R?-Coor.csv</strong> (<strong>?</strong> indicates rat ID) show the <em>x</em> and <em>y</em> coordinates of every recorded cell, one for each row, ordered as<br> they appear in the calcium activity recordings. </p> <p><br> Authors:</p> <ul> <li>Ana Aquiles anaaquiles@ciencias.unam.mx</li> <li>Tatiana Fiordelisio tfiorde@ciencias.unam.mx</li> <li>Hiram Luna-Munguía hiram_luna@inb.unam.mx</li> <li>Luis Concha lconcha@unam.mx</li> </ul> <p> </p> <p>1. Benardete, E. A., & Kriegstein, A. R. (2002). Increased excitability and decreased sensitivity to GABA in an animal model of dysplastic cortex. <em>Epilepsia</em>, <em>43</em>(9), 970-982.</p>
Time Series Data of Gaze, Head Pose, Hand Pose, and Object Positions for Object Approaches with a Given Intention
<p>This data set comprises time series data of gaze, head pose, hand pose, and object positions for object approaches with a given intention. The data was captured in the context of the following publication:</p> <ul> <li><em>Michael Fennel, Serge Garbay, Antonio Zea, Uwe D. Hanebeck</em>, <strong>Intention Estimation with Recurrent Neural Networks for Mixed Reality Environments</strong>, Proceedings of the 26th International Conference on Information Fusion (Fusion 2023) <em>(under review)</em></li> </ul> <p>A Microsoft Hololens 2 was used for recording the data at 60 fps under the modalities explained in detail in the above-mentioned paper.</p> <p>The file names are structured as follows:</p> <ul> <li><em>1st/2nd:</em> <ul> <li>The data with "1st" contains approaches to randomly placed objects on a grid, which are rendered in augmented reality. The user is informed about the object to approach using a visual cue. This corresponds to Section IV-A.</li> <li>The data with "2nd" contains approaches to real objects placed statically in a room. The user is informed about the object to approach using a voice command.</li> </ul> </li> <li><em>unfiltered:</em> Contains all approaches, including those where the user disrespects the given commands. Filtering is done as described in the paper.</li> <li><em>train/val/test:</em> The first dataset was split in a 70/20/10 ratio for training, validation, and test.</li> </ul> <p>Each data set contains the following columns. In each approach, 5 objects numbered from i=0 to i=4 are present.</p> <ul> <li>General: <ul> <li><em>time:</em> in seconds</li> <li><em>subject:</em> consecutive subject number</li> <li><em>handedness:</em> left (1), right (0)</li> <li><em>trial:</em> consecutive trial number per subject</li> <li><em>target_label:</em> index of the object to approach (0 to 4)</li> </ul> </li> <li>Data in world coordinates: <ul> <li><em>head_{x,y,z}:</em> head position</li> <li><em>head_quat_{w,x,y,z}:</em> head orientation quaternion</li> <li><em>W_gaze_{x,y,z}:</em> gaze direction</li> <li><em>W_r_hand_{x,y,z}:</em> right hand position</li> <li><em>W_r_hand_quat_{w,x,y,z}:</em> right hand orientation quaternion</li> <li><em>W_l_hand_{x,y,z}:</em> left hand position</li> <li><em>W_l_hand_quat_{w,x,y,z}:</em> left hand orientation quaternion</li> <li><em>W_object_i_{x,y,z}:</em> position of object i</li> <li><em>W_object_i_quat {w,x,y,z}</em>: orientation quaternion of object i</li> </ul> </li> <li>Data in egocentric coordinates (head coordinate system). This data is provided for convenience and can be derived from the other data: <ul> <li><em>gaze_{x,y,z}:</em> gaze direction</li> <li><em>r_hand_{x,y,z}:</em> right hand position</li> <li><em>r_hand_quat_{w,x,y,z}:</em> right hand orientation quaternion</li> <li><em>l_hand_{x,y,z}:</em> left hand position</li> <li><em>l_hand_quat_{w,x,y,z}:</em> left hand orientation quaternion</li> <li><em>object_i_{x,y,z}:</em> position of object i</li> <li><em>object_i_quat {w,x,y,z}</em>: orientation quaternion of object i</li> </ul> </li> </ul> <p><strong>Acknowledgment:</strong></p> <p>This work was supported by the <a href="https://robdekon.de/">ROBDEKON</a> project of the German Federal Ministry of Education and Research.</p>
InSAR Time Series Analysis (2018-2021) for Volcanic Monitoring in Northern Chile
<p>This dataset is for the paper "First onset of unrest captured at Socompa: A Recent Geodetic Survey at Central Andean volcanoes in Northern Chile" which is published in GRL: <a href="https://doi.org/10.1029/2022GL102480">https://doi.org/10.1029/2022GL102480</a>.</p> <p><strong>InSAR Data:</strong></p> <p>The folder of InSAR_149A.rar stores the InSAR time series analysis dataset on ascending track 149.</p> <ol> <li>The 'Imagedate' folder stores the empty *.rslc files to indicate the date of each SLCs.</li> <li>Data_Asc.mat stores the main InSAR time series data, which includes the UTC time of the acquisition (for accurate time calculation), the length of perpendicular baselines (unit is meter), the number of days counting from the first epoch, the unwrapped time series data (ifg), the unwrapped time series data with GACOS correction (ifg_aps), the look angles (la, unit is rad), and lat&lon.</li> <li>parms.mat stores the parameters used during the data processing by StaMPS.</li> <li>semi_fit.mat stores the results of the semi-variogram fitting of each interferogram on time series. It provides two versions for the original dataset (semi) and the GACOS-corrected dataset (semi_aps). This file is mainly used to weight the data during the time series fitting.</li> <li>runTSA.m, the main function to run the InSAR time series fitting. See more details in the Code part.</li> </ol> <p>The folder of InSAR_156D.rar stores the same content as the InSAR_149A.rar but for descending track 156.</p> <p><strong>Code:</strong></p> <p>This folder contains the codes of the InSAR time series fitting for this dataset, and the GBIS software.</p> <ol> <li>TSA_findref.m, this function is used to search the reference point of the InSAR data.</li> <li>TSA_EQ_fit.m, is the main function to perform InSAR time series fitting.</li> <li>rb_pixel_fit.m, is the robust way to fit the linear model.</li> <li>TSA_EQ_pixel.m, is the function used to plot the results.</li> </ol> <p>To perform the InSAR time series fitting, you need to put these four functions under your Matlab path, and then run the runTSA.m function in the data folder.</p> <p>The GBIS folder stores the updated version of the GBIS software, which allows you to perform the pCDM, CDM, and pECM. The core functions of these models are provided by Dr. Mehdi Nikkhoo, and you could find them here: https://www.volcanodeformation.com/software</p> <p><strong>GBIS_Modelling_Results:</strong></p> <p>This folder stores the data of InSAR and GPS joint inversion for Socompa Uplift.</p> <ol> <li>The folder Socompa stores the modelling results using the models of Okada(D), pECM(E), Mogi(M), pCDM(N), and Yang(Y), respectively. </li> <li>GPS_data.txt stores the cumulative displacements and the uncertainties of the SOCM station in three directions.</li> <li>Socompa.inp is the configuration file for GBIS running.</li> <li>Vol_asc.mat and Vol_asc_ds.mat stores the original and the downsampled ascending data, while Vol_dsc.mat and Vol_dsc_ds.mat store those of descending.</li> </ol> <p>Many thanks for using our dataset and please let me know if you have any further questions!</p>
EO4WildFires: An Earth Observation multi-sensor, time-series machine-learning-ready benchmark dataset for wildfire impact prediction
<p>This paper presents a benchmark dataset called EO4WildFires; a multi-sensor (multi spectral; Sentinel-2, Synthetic-Aperture Radar - SAR; Sentinel-1, meteorological parameters; NASA Power) time-series dataset that spans 45 countries, which can be used for developing machine learning and deep learning methods targeted for the estimation of the area that a forest wildfire might cover.</p> <p>This novel EO4WildFires dataset is annotated using EFFIS (European Forest Fire Information System) as forest fire detection and size estimation data source. A total of 31,742 wildfire events are gathered from 2018 to 2022. For each event, Sentinel-2 (multispectral), Sentinel-1 (SAR) and meteorological data are assembled into a single data cube. The meteorological parameters that are included in the data cube are: ratio of actual partial pressure of water vapor to the partial pressure at saturation, average temperature, bias corrected average total precipitation, average wind speed, fraction of land covered by snowfall, percent of root zone soil wetness, snow depth, snow precipitation, as well as percent of soil moisture.</p> <p>The main problem that this dataset is designed to address, is the severity forecasting before wildfires occur. The dataset is not used to predict wildfire events, but rather to predict the severity (size of area damaged by fire) of a wildfire event, if that happens in a specific place under the current and historical forest status, as recorded from multispectral and SAR images, and meteorological data.</p> <p>Using the data cube for the collected wildfire events, the EO4WildFires dataset is used to realize three (3) different preliminary experiments, in order to evaluate the contributing factors for wildfire severity prediction. The first experiment evaluates wildfire size using only the meteorological parameters, the second one utilizes both the multispectral and SAR parts of the dataset, while the third exploits all dataset parts. In each experiment, machine learning models are developed, and their accuracy is evaluated.</p>
The GNSS time series along the northern coastline of Java, Indonesia
<p>This repository contains the GNSS time series along the northern coastline of Java, both in .rneu and the GAGE's .pos format (https://www.unavco.org/data/gps-gnss/derived-products/docs/NOTICE-TO-DATA-PRODUCT-USERS-GPS-2013-03-15.pdf). The repository also contains the stations' coordinates. </p> <p>Notes:<br> In the .rneu format of the GNSS time series:<br> 1. The outliers have been removed.<br> 2. The offsets due to instrument changes at CGON in mid-2016 and at CSIT in late 2015 have been corrected.</p> <p>Please refer to:<br> Susilo, S., Salman, R., Hermawan, W. <em>et al.</em> GNSS land subsidence observations along the northern coastline of Java, Indonesia. <em>Sci Data</em> <strong>10</strong>, 421 (2023). https://doi.org/10.1038/s41597-023-02274-0</p>
Time series measurements of nitrogen fixation in the subtropical North Pacific (extended through 2019) (Reformatted)
<p>Rates of N2 fixation were measured using the 15N2 isotopic tracer technique. Sampling occurred during near-monthly Hawaii Ocean Time-series cruises. Whole seawater samples from six discrete depths (5, 25, 45, 75, 100, and 125 m) were subsampled into acid-washed 4.3 L polycarbonate bottles. The 15N2 gas was first dissolved into seawater and 100 mL of the resulting 15N2-enriched water was added to 4.3 L polycarbonate sampling bottles. The resulting atom % enrichment of stocks of 15N2-enriched seawater was measured using a membrane inlet mass spectrometer. Incubation bottles amended with the 15N2 tracer were attached to a free-drifting array and incubated at the discrete depths from which samples had been collected. The array was deployed before dawn and samples were incubated at in situ light and temperature for 24 h. After recovery of the array, the entire volume from each bottle was filtered onto a pre-combusted glass microfiber filter (Whatman 25 mm GF/F) and filters were placed onto pre-combusted pieces of foil in Petri dishes and stored frozen at -20°C. Filters were dried for 24 h at 60°C, pelleted, and the total mass of N and its isotopic signature on each filter were analyzed on an elemental analyzer-isotope ratio mass spectrometer (Carlo-Erba EA NC2500 coupled with ThermoFinnigan Delta S). Dataset has been reformatted to meet submission requirements for Simons CMAP.</p>
CAELUS: Classification of sky conditions from 1-min time series of global solar irradiance using variability indices and dynamic thresholds
<p>CAELUS, a novel classification algorithm that relies on various thresholds to separate all possible sky conditions into six classes, is presented in Ruiz-Arias and Gueymard (2023, doi: <a href="https://doi.org/10.1016/j.solener.2023.111895">10.1016/j.solener.2023.111895</a>).</p> <p>This dataset was used to develop, validate and benchmark CAELUS. It is made up by 1-min quality-assured observations of global horizontal irradiance (GHI) and diffuse horizontal irradiance at 54 stations of the Baseline Surface Radiation Network (BSRN) archive, which is publicly available (see download instructions in https://bsrn.awi.de/data). The dataset includes 5 years of data per station, except in two of them (Petrolina, Brazil, and Solar Village, Saudi Arabia), combined with other variables that are required to run CAELUS, namely: solar zenith angle (sza), extraterrestrial horizontal solar irradiance (eth), clear-sky GHI (ghics) and GHI in a clean and dry atmosphere (ghicda). In addition, the dataset also provides the sky classification obtained with CAELUS.</p> <p>Further details about CAELUS and the dataset compilation is available in Ruiz-Arias and Gueymard (2023, doi: <a href="https://doi.org/10.1016/j.solener.2023.111895">10.1016/j.solener.2023.111895</a>). A Python implementation of CAELUS is available in https://github.com/jararias/caelus.</p>
A Meta-Learner Approach to Multistep-Ahead Time Series Prediction
<p><strong>Abstract</strong></p> <p>The application of machine learning has become commonplace for problems in modern data science. The democratization of the decision process when choosing a machine learning algorithm has also received considerable attention through the use of meta features and automated machine learning for both classification and regression type problems. However, this is not the case for multistep-ahead time series problems. Time series models generally rely upon the series itself to make future predictions, as opposed to independent features used in regression and classification problems. The structure of a time series is generally described by features such as trend, seasonality, cyclicality, and irregularity. In this research, we demonstrate how time series metrics for these features, in conjunction with an ensemble based regression learner, were used to predict the standardized mean square error of candidate time series prediction models. These experiments used datasets that cover a wide feature space and enable researchers to select the single best performing model or the top N performing models. A robust evaluation was carried out to test the learner's performance on both synthetic and real time series. </p> <p><strong>Proposed Dataset</strong></p> <p>The dataset proposed here gives the results for 20 step ahead predictions for eight Machine Learning/Multi-step ahead prediction strategies for 5,842 time series datasets outlined <a href="https://www.sciencedirect.com/science/article/pii/S2215016121002521">here</a>. It was used as the training data for the Meta Learners in this research. The meta features used are columns C to AE. Columns AH outlines the method/strategy used and columns AI to BB (the error) is the outcome variable for each prediction step. The description of the method/strategies is as follows:</p> <p><strong>Machine Learning methods:</strong></p> <ul> <li>NN: Neural Network</li> <li>ARIMA: Autoregressive Integrated Moving Average</li> <li>SVR: Support Vector Regression</li> <li>LSTM: Long Short Term Memory</li> <li>RNN: Recurrent Neural Network</li> </ul> <p><strong>Multistep ahead prediction strategy:</strong></p> <ul> <li>OSAP: One Step ahead strategy</li> <li>MRFA: Multi Resolution Forecast Aggregation</li> </ul>
Datasets used in the study "Trends in medication use after the onset of the COVID-19 pandemic in the Republic of Ireland: an interrupted time series study"
<p>This record contains datasets analysed as part of the study "Trends in medication use after the onset of the COVID-19 pandemic in the Republic of Ireland: an interrupted time series study".</p> <p>Two datasets were used, one relating to therapeutic subgroups defined by ATC codes (atc_wide_freq_avg.csv) and one relating to individual medications (drugs_wide_freq_avg.csv). Datasets were collated by combining monthly data reported by HSE Primary Care Reimbursement Services in Ireland relating to dispensing on the General Medical Services scheme at https://www.sspcrs.ie/portal/annual-reporting/</p> <p>Code used to collate datasets and for data management is included in Stata format (compile_data_export_for_analysis_final.do).</p> <p>The study protocol is available at https://doi.org/10.17605/OSF.IO/B4RTM</p>
Network traffic datasets created by Single Flow Time Series Analysis
<p><strong>Network traffic datasets created by Single Flow Time Series Analysis</strong></p> <p>Datasets were created for the paper: Network Traffic Classification based on Single Flow Time Series Analysis -- Josef Koumar, Karel Hynek, Tomáš Čejka -- which was published at The 19th International Conference on Network and Service Management (CNSM) 2023. Please cite usage of our datasets as:<br> </p> <blockquote> <p>J. Koumar, K. Hynek and T. Čejka, "Network Traffic Classification Based on Single Flow Time Series Analysis," <em>2023 19th International Conference on Network and Service Management (CNSM)</em>, Niagara Falls, ON, Canada, 2023, pp. 1-7, doi: 10.23919/CNSM59352.2023.10327876.</p> </blockquote> <p>This Zenodo repository contains 23 datasets created from 15 well-known published datasets which are cited in the table below. Each dataset contains 69 features created by Time Series Analysis of Single Flow Time Series. The detailed description of features from datasets is in the file: <em>feature_description.pdf</em></p> <p> </p> <p>In the following table is a description of each dataset file:</p> <table> <tbody> <tr> <td><strong>File name</strong></td> <td><strong>Detection problem</strong></td> <td><strong>Citation of original raw dataset</strong></td> </tr> <tr> <td>botnet_binary.csv </td> <td>Binary detection of botnet </td> <td>S. García et al. An Empirical Comparison of Botnet Detection Methods. Computers & Security, 45:100–123, 2014. </td> </tr> <tr> <td>botnet_multiclass.csv </td> <td>Multi-class classification of botnet </td> <td>S. García et al. An Empirical Comparison of Botnet Detection Methods. Computers & Security, 45:100–123, 2014. </td> </tr> <tr> <td>cryptomining_design.csv</td> <td>Binary detection of cryptomining; the design part </td> <td>Richard Plný et al. Datasets of Cryptomining Communication. Zenodo, October 2022 </td> </tr> <tr> <td>cryptomining_evaluation.csv </td> <td>Binary detection of cryptomining; the evaluation part </td> <td>Richard Plný et al. Datasets of Cryptomining Communication. Zenodo, October 2022 </td> </tr> <tr> <td>dns_malware.csv </td> <td>Binary detection of malware DNS </td> <td>Samaneh Mahdavifar et al. Classifying Malicious Domains using DNS Traffic Analysis. In DASC/PiCom/CBDCom/CyberSciTech 2021, pages 60–67. IEEE, 2021. </td> </tr> <tr> <td>doh_cic.csv </td> <td>Binary detection of DoH </td> <td> <p>Mohammadreza MontazeriShatoori et al. Detection of doh tunnels using time-series classification of encrypted traffic. In DASC/PiCom/CBDCom/CyberSciTech 2020, pages 63–70. IEEE, 2020 </p> </td> </tr> <tr> <td>doh_real_world.csv </td> <td>Binary detection of DoH </td> <td>Kamil Jeřábek et al. Collection of datasets with DNS over HTTPS traffic. Data in Brief, 42:108310, 2022 </td> </tr> <tr> <td>dos.csv </td> <td>Binary detection of DoS </td> <td>Nickolaos Koroniotis et al. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst., 100:779–796, 2019.</td> </tr> <tr> <td>edge_iiot_binary.csv </td> <td>Binary detection of IoT malware </td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>edge_iiot_multiclass.csv</td> <td>Multi-class classification of IoT malware</td> <td>Mohamed Amine Ferrag et al. Edge-iiotset: A new comprehensive realistic cyber security dataset of iot and iiot applications: Centralized and federated learning, 2022.</td> </tr> <tr> <td>https_brute_force.csv</td> <td>Binary detection of HTTPS Brute Force</td> <td>Jan Luxemburk et al. HTTPS Brute-force dataset with extended network flows, November 2020</td> </tr> <tr> <td>ids_cic_binary.csv</td> <td>Binary detection of intrusion in IDS</td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108–116, 2018.</td> </tr> <tr> <td>ids_cic_multiclass.csv </td> <td>Multi-class classification of intrusion in IDS </td> <td>Iman Sharafaldin et al. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp, 1:108–116, 2018. </td> </tr> <tr> <td>ids_unsw_nb_15_binary.csv </td> <td>Binary detection of intrusion in IDS </td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1–6. IEEE, 2015.</td> </tr> <tr> <td>ids_unsw_nb_15_multiclass.csv </td> <td>Multi-class classification of intrusion in IDS </td> <td>Nour Moustafa and Jill Slay. Unsw-nb15: a comprehensive data set for network intrusion detection systems (unsw-nb15 network data set). In 2015 military communications and information systems conference (MilCIS), pages 1–6. IEEE, 2015.</td> </tr> <tr> <td>iot_23.csv </td> <td>Binary detection of IoT malware </td> <td>Sebastian Garcia et al. IoT-23: A labeled dataset with malicious and benign IoT network traffic, January 2020. More details here https://www.stratosphereips.org /datasets-iot23</td> </tr> <tr> <td>ton_iot_binary.csv </td> <td>Binary detection of IoT malware </td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>ton_iot_multiclass.csv </td> <td>Multi-class classification of IoT malware </td> <td>Nour Moustafa. A new distributed architecture for evaluating ai-based security systems at the edge: Network ton iot datasets. Sustainable Cities and Society, 72:102994, 2021</td> </tr> <tr> <td>tor_binary.csv </td> <td>Binary detection of TOR </td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253–262. SciTePress, 2017. </td> </tr> <tr> <td>tor_multiclass.csv </td> <td>Multi-class classification of TOR </td> <td>Arash Habibi Lashkari et al. Characterization of Tor Traffic using Time based Features. In ICISSP 2017, pages 253–262. SciTePress, 2017. </td> </tr> <tr> <td>vpn_iscx_binary.csv </td> <td>Binary detection of VPN </td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407–414, 2016. </td> </tr> <tr> <td>vpn_iscx_multiclass.csv </td> <td>Multi-class classification of VPN </td> <td>Gerard Draper-Gil et al. Characterization of Encrypted and VPN Traffic Using Time-related. In ICISSP, pages 407–414, 2016. </td> </tr> <tr> <td>vpn_vnat_binary.csv </td> <td>Binary detection of VPN </td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> <tr> <td>vpn_vnat_multiclass.csv</td> <td>Multi-class classification of VPN </td> <td>Steven Jorgensen et al. Extensible Machine Learning for Encrypted Network Traffic Application Labeling via Uncertainty Quantification. CoRR, abs/2205.05628, 2022</td> </tr> </tbody> </table> <p> </p>
Pre-processed and modeled GNSS time-series after the 2011 Tohoku Earthquake
<p>The raw, pre-processed, and modeled GNSS time-series of the 213 GEONET sites in the Tohoku region, Japan, from Mar. 12, 2011 to Nov. 20, 2021, relative to the Okhotsk plate (Argus et al., 2011, <em><em>Geochemistry, Geophysics, Geosystems</em></em>).</p> <p>The original GNSS time-series are F5 solutions, which are distributed by Geospatial Information Authority of Japan (GSI, https://www.gsi.go.jp/). The details and availability of F5 solutions are written in Takamatsu et al. (2023, Earth, Planets, and Space) https://doi.org/10.1186/s40623-023-01787-7.</p> <p>The GNSS time-series processing was performed by Tomita (submitted), and the following signals were excluded from the raw time-series: seasonal variation, coseismic step, antenna maintenance offset, and common mode errors. Then, the pre-processed time-series were modeled by a trajectory modeling method considering postseismic deformation of the 2011 Tohoku earthquake, the Boso SSEs, and postseismic deformations due to aftershocks and L-ASE (long-term aseismic slip event) since late 2019.<br> <br> "sitelist.txt" - Site information file<br> column 1: Full site ID<br> column 2: 4digits site ID<br> column 3: Longitude [deg]<br> column 4: Latitude [deg]<br> column 5: Height [m] <br> <br> "pre-process/xxxx.txt" - Time-series at xxxx (4digits site ID) site<br> column 1: days from Mar. 12, 2011 (1 corresponds to Mar. 12, 2011)<br> column 2: raw East-West displacement [m]<br> column 3: raw North-South displacement [m]<br> column 4: raw Up-down displacement [m]<br> column 5: pre-processed East-West displacement [m]<br> column 6: pre-processed North-South displacement [m]<br> column 7: pre-processed Up-down displacement [m]</p> <p>"model/xxxx/prediction_yy.txt" - Time-series for yy component (yy=EW, NS, UD) at xxxx (4digits site ID) site<br> column 1: days from Mar. 12, 2011 (1 corresponds to Mar. 12, 2011)<br> column 2: modeled displacement excluding the Boso SSEs [m]<br> column 3: modeled displacement excluding the Boso SSEs and postseismic deformation due to aftershocks caused one year after the 2011 Tohoku Eq. [m]<br> column 4: modeled displacement excluding the Boso SSEs, postseismic deformation due to aftershocks caused one year after the 2011 Tohoku Eq. and the 2019 L-ASE [m]</p> <p><br> The displacement on Mar. 12, 2011 was initially set to be zero before the pre-processing, but the removal of the above factors provided some deviation from zero.</p> <p>The raw time-series excluded outliers from the original F5 solutions, and the raw time-series were transformed into the Okhotsk plate reference.</p> <p>Following the above trajectory modeling, the fully-relaxed postseismic displacement fields due to 2015 Feb. 17 Sanriku-oki earthquake ("Table_displacement1.xlsx"), the 2015 May 13 Miyagi-oki earthquake ("Table_displacement2.xlsx"), and summation of the 2021 Feb. 13 Fukushima-oki, the 2021 Mar. 20 Miyagi-oki, and the 2021 May 1 earthquakes ("Table_displacement2.xlsx") were calculated. Moreover, the cumulative displacement field due to the 2019 L-ASE since Nov. 25, 2019 was also calculated. In those files, the estimation errors are also shown as 1σ standard deviation obtained from diagonal components of the model covariance matrices. </p> <p> </p> <p>The details of these data are introduced in the corresponding paper (Tomita, submitted).</p>
High resolution and high cadence time series of land surface categories, land use land cover, and land use land cover changes
<p>A prototype of monthly, 10 m resolution land surface categories, land use land cover (LULC) cover, and LULC change maps derived from Sentinel-2 data over three areas within Belgium, Portugal, and Sicily for the period 2018-2020. The LULC and LULC change maps were independently validated by IIASA. All products were generated within the framework of the RapidAI4EO project, funded from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 101004356.</p> <p>The data description can be found below. The validation report of the LULC and LULC change maps can be found in validation_LULC.pdf and validation_change.pdf, respectively, and the validation dataset can be found in Lesiv <em>et al.</em> (2023).</p> <p><strong>Data description</strong></p> <p>Increasing the cadence of the land cover updates from the typical (multi-)annual to monthly cadence poses several challenges. First, several land cover types are difficult to discriminate without any knowledge of temporal dynamics. For instance, croplands are characterized by a dynamic of vegetation growth and a harvest period (i.e. cycles of bare soil, sparsely vegetated and vegetated periods). This contrasts with grasslands that often lack the harvest period resulting in a bare soil cover. Without this temporal information, it is difficult to distinguish a vegetated cropland field from grassland. Second, phenological changes may introduce a large intra-class variability and thus also confusion between classes. For example, the shedding of leaves during autumn or wilting of herbaceous vegetation in dry summer periods introduces spectral variability within land cover classes.</p> <p>To overcome these challenges, we developed a workflow with two main phases. The first phase aims to map land surface categories (LSC) at a monthly resolution. The next phase uses the resulting monthly LSC probability time series to classify land cover.</p> <p><strong><em>Land surface category (LSC)</em></strong></p> <p>These LSC represent basic, observable bio-geophysical properties (categories) of the Earth surface that can be predicted directly from individual monthly composites. LSC classes contain a set of vegetated and non-vegetated surface categories.</p> <p>Discrete LSC classification legend:</p> <table> <tbody> <tr> <td> <p>Map code</p> </td> <td> <p>Land cover class</p> </td> </tr> <tr> <td> <p>11</p> </td> <td> <p>Tree (leaf-on)</p> </td> </tr> <tr> <td> <p>12</p> </td> <td> <p>Shrubland (leaf-on)</p> </td> </tr> <tr> <td> <p>13</p> </td> <td> <p>Grassland</p> </td> </tr> <tr> <td> <p>14</p> </td> <td> <p>Woody vegetation (leaf-off)</p> </td> </tr> <tr> <td> <p>15</p> </td> <td> <p>Wilted herbaceous vegetation</p> </td> </tr> <tr> <td> <p>21</p> </td> <td> <p>Bare/sparse vegetation</p> </td> </tr> <tr> <td> <p>22</p> </td> <td> <p>Water</p> </td> </tr> <tr> <td> <p>24</p> </td> <td> <p>Built-up</p> </td> </tr> </tbody> </table> <p>In order to predict the LSC, we trained a CatBoost model (Dorogush et al., 2018) using a DEM, spectral bands and vegetation indices, country, the timing (month) of the spectral data, and the pseudo-probability of a U-Net model trained to segment built-up surfaces as input. Labels were derived by post-processing the land cover labels of the ESA WorldCover product (Zanaga et al., 2021). Please note that the collection of these labels was suboptimal, likely having an impact on the LULC and change maps generated in the prototype.</p> <p><strong><em>Land use land cover</em></strong></p> <p>After predicting LSC over the three AOI’s, we trained a CatBoost model using the LSC probabilities over a window of one year, country, and the timing (month) as independent variable. The use of LSC probabilities over multiple months allows to incorporate information about dynamics, which is necessary to discriminate some classes (e.g. cropland and grassland or cropland and bare). Similar to the LSC labels, the LULC labels were derived from the ESA WorldCover product v100 (year 2020), resulting in a similar legend system.</p> <p>The use of a moving window approach to predict LULC allows to (i) incorporate temporal information that is necessary to discriminate land cover classes and (ii) is expected to lead to more consistent land cover maps. It however has the disadvantage that (i) no land cover predictions are available at the beginning and the end of the time series and (ii) the timing of the predicted land cover change is not always accurate. To resolve these issues, we applied a post-processing step that compares and integrates the LULC predictions and cleaned LSC predictions.</p> <p>Discrete LC classification legend:</p> <table> <tbody> <tr> <td> <p><strong>Map code</strong></p> </td> <td> <p><strong>Land cover class</strong></p> </td> </tr> <tr> <td> <p>10</p> </td> <td> <p>Tree cover</p> </td> </tr> <tr> <td> <p>20</p> </td> <td> <p>Shrubland</p> </td> </tr> <tr> <td> <p>30</p> </td> <td> <p>Grassland</p> </td> </tr> <tr> <td> <p>40</p> </td> <td> <p>Cropland</p> </td> </tr> <tr> <td> <p>50</p> </td> <td> <p>Built-up</p> </td> </tr> <tr> <td> <p>60</p> </td> <td> <p>Bare/sparse vegetation</p> </td> </tr> <tr> <td> <p>80</p> </td> <td> <p>Permanent water bodies</p> </td> </tr> <tr> <td> <p>90</p> </td> <td> <p>Herbaceous wetland</p> </td> </tr> </tbody> </table> <p><strong><em>Land use land cover change </em></strong></p> <p>Monthly change maps were finally derived from the land cover maps. The pixel values within the change maps represent the percentage of pixels that changed with respect to the previous month over an area of 90x90m. The maps contain values between 0-100, with larger values assigned to larger change patches. A value of 100 indicates that all pixels within an area of 90x90m around the pixel were flagged as change.</p> <p><strong><em>Files</em></strong></p> <p>The zip files contain the following data:</p> <ul> <li>lsc.zip: land surface category maps over the three AOI’s</li> <li>lc.zip: LULC maps over the three AOI’s</li> <li>change.zip: change maps over the three AOI’s</li> </ul> <p>These maps are generated for each month over the period 2018-2020 for each of the tiles (see tiles.gpkg for an overview of all tiles). The files names use the following naming convention: “<em>tile</em>-<em>year</em>-<em>month</em>.tif”.</p> <p><strong><em>References</em></strong></p> <p>Myroslava Lesiv, Halyna Bun, & Martina Duerauer. (2023). Validation data set on land cover changes for RapidAI4EO project [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7825963 </p> <p>Dorogush, A. V., Ershov, V., & Gulin, A. (2018). CatBoost: gradient boosting with categorical features support. arXiv preprint arXiv:1810.11363.</p> <p><em>Zanaga, D., </em><em>et al.</em><em>., 2021. ESA WorldCover 10 m 2020 v100. </em><a href="https://doi.org/10.5281/zenodo.5571936 "><em>https://doi.org/10.5281/zenodo.5571936 </em></a></p>
Duck94 time series
<p>Time series from pressure gages and current meters deployed along a cross-shore transect from near the shoreline to about 5 m water depth at the USACE Field Research Facility, Duck, NC in 1994.</p> <p>The README describes the 2Hz time series, the file with X-Y-Z locations, and provides links to bathymetry surveys and incident wave frequency-directional data.</p> <p>The Duck94 data set includes both offshore and onshore sandbar migration, incident waves from the north and the south of shore normal, and incident significant wave heights from 0.4 to 4.0 m.</p>
Controlled Anomalies Time Series (CATS) Dataset
<p>The Controlled Anomalies Time Series (CATS) Dataset consists of commands, external stimuli, and telemetry readings of a simulated complex dynamical system with 200 injected anomalies.</p> <p>The CATS Dataset exhibits a set of desirable properties that make it very suitable for benchmarking<strong> Anomaly Detection Algorithms in Multivariate Time Series </strong>[1]:</p> <ul> <li><strong>Multivariate (17 variables) </strong>including sensors reading and control signals. It simulates the operational behaviour of an arbitrary complex system including: <ul> <li><strong>4 Deliberate Actuations / Control Commands sent by a simulated operator / controller</strong>, for instance, commands of an operator to turn ON/OFF some equipment.</li> <li><strong>3 Environmental Stimuli / External Forces</strong> acting on the system and affecting its behaviour, for instance, the wind affecting the orientation of a large ground antenna.</li> <li><strong>10 Telemetry Readings</strong> representing the observable states of the complex system by means of sensors, for instance, a position, a temperature, a pressure, a voltage, current, humidity, velocity, acceleration, etc.</li> </ul> </li> <li><strong>5 million timestamps</strong>. Sensors readings are at 1Hz sampling frequency. <ul> <li><strong>1 million nominal </strong>observations (the first 1 million datapoints). This is suitable to start learning the "normal" behaviour.</li> <li><strong>4 million</strong> observations that include both <strong>nominal and anomalous segments</strong>. This is suitable to evaluate both semi-supervised approaches (novelty detection) as well as unsupervised approaches (outlier detection).</li> </ul> </li> <li><strong>200 anomalous segments. </strong>One anomalous segment may contain several successive anomalous observations / timestamps. Only the last 4 million observations contain anomalous segments.</li> <li><strong>Different types of anomalies </strong>to understand what anomaly types can be detected by different approaches. The categories are available in the dataset and in the metadata.</li> <li><strong>Fine control over ground truth.</strong> As this is a simulated system with deliberate anomaly injection, the start and end time of the anomalous behaviour is known very precisely. In contrast to real world datasets, there is no risk that the ground truth contains mislabelled segments which is often the case for real data.</li> <li><strong>Suitable for root cause analysis.</strong> In addition to the anomaly category, the time series channel in which the anomaly first developed itself is recorded and made available as part of the metadata. This can be useful to evaluate the performance of algorithm to trace back anomalies to the right root cause channel.</li> <li><strong>Affected channels.</strong> In addition to the knowledge of the root cause channel in which the anomaly first developed itself, we provide information of channels possibly affected by the anomaly. This can also be useful to evaluate the explainability of anomaly detection systems which may point out to the anomalous channels (root cause and affected).</li> <li><strong>Obvious anomalies.</strong> The simulated anomalies have been designed to be "easy" to be detected for human eyes (i.e., there are very large spikes or oscillations), hence also detectable for most algorithms. It makes this synthetic dataset useful for screening tasks (i.e., to eliminate algorithms that are not capable to detect those obvious anomalies). However, during our initial experiments, the dataset turned out to be challenging enough even for state-of-the-art anomaly detection approaches, making it suitable also for regular benchmark studies.</li> <li><strong>Context provided. </strong>Some variables can only be considered anomalous in relation to other behaviours. A typical example consists of a light and switch pair. The light being either on or off is nominal, the same goes for the switch, but having the switch on and the light off shall be considered anomalous. In the CATS dataset, users can choose (or not) to use the available context, and external stimuli, to test the usefulness of the context for detecting anomalies in this simulation.</li> <li><strong>Pure signal ideal for robustness-to-noise analysis.</strong> The simulated signals are provided without noise: while this may seem unrealistic at first, it is an advantage since users of the dataset can decide to add on top of the provided series any type of noise and choose an amplitude. This makes it well suited to test how sensitive and robust detection algorithms are against various levels of noise.</li> <li><strong>No missing data.</strong> You can drop whatever data you want to assess the impact of missing values on your detector with respect to a clean baseline.</li> </ul> <p><strong>Change Log</strong></p> <p>Version 2</p> <ul> <li><strong>Metadata:</strong> we include a metadata.csv with information about: <ul> <li>Anomaly categories</li> <li>Root cause channel (signal in which the anomaly is first visible)</li> <li>Affected channel (signal in which the anomaly might propagate) through coupled system dynamics</li> </ul> </li> <li><strong>Removal of anomaly overlaps:</strong> version 1 contained anomalies which overlapped with each other resulting in only 190 distinct anomalous segments. Now, there are no more anomaly overlaps.</li> <li><strong>Two data files: </strong>CSV and parquet for convenience.</li> </ul> <p>[1] Example Benchmark of Anomaly Detection in Time Series: “Sebastian Schmidl, Phillip Wenig, and Thorsten Papenbrock. Anomaly Detection in Time Series: A Comprehensive Evaluation. PVLDB, 15(9): 1779 - 1797, 2022. doi:10.14778/3538598.3538602”</p> <p><strong>About Solenix</strong></p> <p>Solenix is an international company providing software engineering, consulting services and software products for the space market. Solenix is a dynamic company that brings innovative technologies and concepts to the aerospace market, keeping up to date with technical advancements and actively promoting spin-in and spin-out technology activities. We combine modern solutions which complement conventional practices. We aspire to achieve maximum customer satisfaction by fostering collaboration, constructivism, and flexibility.</p>
Ensemble Machine Learning Prediction of Potential FAPAR: Monthly time-series 2021 and Long-Term Comparison with Actual FAPAR
<p><strong>General Description</strong></p> <p>The dataset contains composites at 250 m spatial resolution of (1) monthly potential FAPAR for the year 2021 from ensemble ML model predictions, (2) the model deviance for each prediction, (3) the yearly average of potential FAPAR, (4) the yearly average of actual FAPAR and (5) the yearly average of the difference between actual and potential (actual minus potential) FAPAR. The dataset is based on the <a href="https://zenodo.org/record/8392976">95th percentile of the monthly aggregated FAPAR</a> derived from <a href="http://glass.umd.edu/Overview.html">250 m 8 d GLASS V6 FAPAR</a>. Potential FAPAR was predicted by fitting an ensemble ML model using globally distributed training points (cca 3 Mio) and a set of 52 biophysical covariates including several layers related to human pressure. The code for modeling potential FAPAR is openly available at <a href="http://github.com/Open-Earth-Monitor/Global_FAPAR_250m">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m</a>. The dataset can be used in many applications like land degradation modeling, land productivity mapping, and land potential mapping. </p> <p><strong>Data Details</strong></p> <ul> <li><strong>Time period:</strong> January 2021 - December 2021</li> <li><strong>Type of data: </strong>Fraction of Absorbed Photosynthetically Active Radiation (FAPAR)</li> <li><strong>How the data was collected or derived:</strong> Derived from 250m 8 d GLASS V6 FAPAR</li> <li><strong>Statistical methods used: </strong>Ensemble machine learning</li> <li><strong>Limitations or exclusions in the data: </strong>The dataset does not include data for Antarctica.</li> <li><strong>Coordinate reference system:</strong> EPSG:4326</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.0008094, 179.9999424, 87.37000)</li> <li><strong>Spatial resolution:</strong> 1/480 d.d. = 0.00208333 (250m)</li> <li><strong>Image size: </strong>172,800 x 71,698</li> <li><strong>File format: </strong>Cloud Optimized Geotiff (COG) format.</li> </ul> <p><strong>Support</strong></p> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: <a href="https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues</a></p> <p><strong>Reference</strong></p> <p>Hackländer, J., Parente, L., Ho, Y.-F., Hengl, T., Simoes, R., Consoli, D., Şahin, M., Tian, X., Herold, M., Jung, M., Duveiller, G., Weynants, M., Wheeler, I., (2023?) "Land potential assessment and trend-analysis using 2000–2021 FAPAR monthly time-series at 250 m spatial resolution", submitted to PeerJ, preprint available at: <a href="https://doi.org/10.21203/rs.3.rs-3415685/v1">https://doi.org/10.21203/rs.3.rs-3415685/v1</a></p> <p> </p> <p><strong>Name convention</strong></p> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p> <ol> <li><strong>generic variable name:</strong> pot.fapar = Potential Fraction of Absorbed Photosynthetically Active Radiation</li> <li><strong>variable procedure combination: </strong>eml = ensemble machine learning</li> <li><strong>Position in the probability distribution / variable type:</strong> m = mean</li> <li><strong>Spatial support:</strong> 250m</li> <li><strong>Depth reference: </strong>s = surface</li> <li><strong>Time reference begin time:</strong> 20210101 = 2021-01-01</li> <li><strong>Time reference end time:</strong> 20211231 = 2021-12-31</li> <li><strong>Bounding box: </strong>go = global (without Antarctica)</li> <li><strong>EPSG code:</strong> epsg.4326 = EPSG:4326</li> <li><strong>Version code:</strong> v20230924 = 2023-09-24 (creation date)</li> </ol>
Host network traffic time series 2019/01
<p><em><strong>General info</strong></em></p> <p>Dataset was collected over one <strong>month period in January 2019</strong>. The observation points for the collection of IP flows were located at the borders of the university campus network. The campus university network has /16 CIDR IPv4 network range at disposal and contains various network segments from segments connecting dormitories, over server segments, to a segment containing working stations of university administrative workers. The size of the raw IP flows used to create the dataset was over 860GB. <strong>A host in our dataset is identified by its source IPv4 address. </strong><br> </p> <p><em><strong>Variables</strong></em></p> <p>The dataset contains the following variables:</p> <ul> <li><strong>Aggregations</strong> - created from five-minute total volumes aggregated over one-hour disjoint windows using mean/max/min aggregation functions <ul> <li><strong># of flows (FL) </strong>- number of flows for a given source IP </li> <li><strong># of packets (PKT)</strong> - number of packets for a given source IP</li> <li><strong># of bytes (BYT)</strong> - number of packets for a given source IP</li> <li><strong>flow duration (DUR)</strong> - average flow duration in seconds</li> </ul> </li> <li><strong>Distinct Counts </strong>- count of distinct values for each variable in five-minute window aggregated over one-hour disjoint windows using mean/max/min aggregation functions <ul> <li><strong># of peers (PEER)</strong> - number of distinct communication peers for a given source IP</li> <li><strong># of ports (PORTS)</strong> - number of distinct destination ports for a given source IP</li> <li><strong># of protocols (PROTO)</strong> - number of distinct communication protocols for a given source IP</li> <li><strong># of AS numbers (AS)</strong> - number of distinct destination AS numbers for a given source IP</li> <li><strong># of countries (CTRY)</strong> - number of distinct destination countries for a given source IP</li> </ul> </li> <li><strong>Labels</strong> <ul> <li><strong>Range (RNG)</strong> - a network range a host belongs to (anonymized)</li> <li><strong>Unit (UNT) </strong>- an administrative unit owning the network range</li> <li><strong>Sub-unit (SUB-UNT)</strong> - a sub-unit of the unit</li> </ul> </li> </ul> <p> </p> <p><em><strong>Dataset format</strong></em></p> <ul> <li>The dataset is in <strong>comma-separated values (CSV)</strong> format. </li> <li><strong>Header</strong> - multilevel, first 3 lines <ul> <li>1 level - aggregation type {mean|min|max}</li> <li>2 level - variable {see above}</li> <li>3 level - hour of a day {00,01,02,03,...,22,23}</li> </ul> </li> <li><strong>Lablels</strong> - last 4 columns</li> <li><strong>Dataset size </strong> <ul> <li>rows: 65536 host records + 3 headers</li> <li>columns: 648 variables + 4 labels</li> </ul> </li> </ul> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.