Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,118
datasets available to search
ShareScore release 0.9.0
Dataset results
1,118 results for “Time series”
Aeroelastic simulations of wind turbines affected by leading edge erosion: datasets for multivariate time-series classification
<p>This repository contains data generated and used for classification in the publication:<br> Duthé, G.; Abdallah, I.; Barber, S.; Chatzi, E. Modeling and Monitoring Erosion of the Leading Edge of Wind Turbine Blades. <em>Energies</em> <strong>2021</strong>, <em>14</em>, 7262. https://doi.org/10.3390/en14217262</p> <p>The data is generated via OpenFAST aeroelastic simulations coupled with a Non-Homogeneous Compound Poisson Process for degradation modelling and was used to train a Transformer deep learning model.</p> <p>One degradation run generates 1200 samples (1 sample every 6 days corresponding to a 20 year degradation period). In total 20 degradation runs are made available (20x1200 = 24'000 multivariate time-series samples). This repo can serve to benchmark long multivariate time-series classification algorithms. There are 10 possible classes of erosion severity.</p> <p>Each sample is a multivariate time-series of length 60'000, with the following 4 channels extracted from the simulations for a section at the tip of the blade:</p> <ul> <li>Inflow velocity</li> <li>Angle of attack</li> <li>Lift coefficient</li> <li>Drag coefficient</li> </ul> <p>Please see the publication above for more information as well as the included readme for information about the data and an example of how to load it into to PyTorch.</p> <p> </p>
GPS time series of solid Earth deformation for the southern Antarctic Peninsula
<p>This dataset contains vertical GPS time series observed from selected sites in the southern Antarctic Peninsula. The time series were processed using GAMIT-GLOBK software (Herring et al., 2016) in combination with a globally distributed network that includes all available data from the International GNSS Service (IGS), US POLENET-ANET and UKANET networks. The provided archive contains raw and corrected time series for the effect of elastic deformation-induced from the RACMO surface mass balance (SMB) model with a 5.5 km horizontal resolution (van Wessem et al. 2016). This data is used in the manuscript "GPS-observed elastic deformation due to surface mass balance variability in the Southern Antarctic Peninsula" to study how modelled elastic deformation due to SMB variation can explain vertical land motion observed derived from GPS signals.</p>
TimeSpec4LULC: A Smart-Global Dataset of Multi-Spectral Time Series of MODIS Terra-Aqua from 2000 to 2021 for Training Machine Learning models to perform LULC Mapping
<p>TimeSpec4LULC is a smart open-source global dataset of multi-spectral time series for 29 Land Use and Land Cover (LULC) classes ready to train machine learning models. It was built based on the seven spectral bands of the MODIS sensors at 500 m resolution from 2000 to 2021 (262 observations in each time series). Then, was annotated using spatial-temporal agreement across the 15 global LULC products available in Google Earth Engine (GEE).</p> <p>TimeSpec4LULC contains two datasets: the original dataset distributed over 6,076,531 pixels, and the balanced subset of the original dataset distributed over 29000 pixels.</p> <p>The original dataset contains 30 folders, namely "Metadata", and 29 folders corresponding to the 29 LULC classes. The folder "Metadata" holds 29 different CSV files describing the metadata of the 29 LULC classes. The remaining 29 folders contain the time series data for the 29 LULC classes. Each folder holds 262 CSV files corresponding to the 262 months. Inside each CSV file, we provide the seven values of the spectral bands as well as the coordinates for all the LULC class-related pixels.</p> <p>The balanced subset of the original dataset contains the metadata and the time series data for 1000 pixels per class representative of the globe. It holds 29 different JSON files following the names of the 29 LULC classes.</p> <p>The features of the dataset are:</p> <p>- ".geo": the geometry and coordinates (longitude and latitude) of the pixel center.</p> <p>- "ADM0_Code": the GAUL country code.</p> <p>- "ADM1_Code": the GAUL first-level administrative unit code.</p> <p>- GHM_Index": the average of the global human modification index.</p> <p>- "Products_Agreement_Percentage": the agreement percentage over the 15 global LULC products available in GEE.</p> <p>- "Temporal_Availability_Percentage": the percentage of non-missing values in each band.</p> <p>- "Pixel_TS": the time series values of the seven spectral bands.</p>
Virtual stations (TeroVIR ) and water level time series (TeroWAT) in West Africa and Arctic regions
<p>The dataset contains a sample of locations across Siberia and Africa, for which water-level time series were automatically derived from Sentinel-3 altimeters (methodology described in Machefer et al. 2022<sup>1</sup>) from year 2016 to year 2021, together with the in-situ station records and the area covered by the altimetry measurements. The purpose of this dataset is validation and exemplification of the methodology. </p> <p>The methodology described produces comprehensive water level records at a global scale based on altimetry satellite data. The validation against in-situ data was assessed in numerous environments in West Africa and complex locations such as Arctic rivers partially covered with ice.<br> <br> This dataset offers a sample of the records at 3 locations in West Africa (Kemacina [Mali], Koulikouro [Mali], Lokoja [Niger]) and in the sub-arctic region (Yakutsk [Russia]). The data are organised by Level 1 of <a href="http://www.hydrosheds.org/">HydroBASINS</a><sup>2 </sup>definition (ex: africa) in two folders, each containing: virtual stations (teroVIR) and insitu stations (insitu) as shapefiles with their associated metadata, the corresponding water level time series (teroWAT) in NetCDF, and the level 3 of HydroBASINS, corresponding to the largest river basins of each continent. Finally, a csv file (validation) presents the computed metrics assessing the accuracy of the processors.</p> <p>N.B.: time series with less than two common date points between insitu and teroWAT have not been assessed. </p> <p>[1] Machefer, M., Perpinyà-Vallès M., Escorihuela M.J., Gustafsson D., Romero L. (2022): Challenges and evolution of water level monitoring towards a comprehensive, world-scale coverage with remote sensing. Earth System Science Data (Under Reviewing)</p> <p>[2] Lehner, B., Grill G. (2013): Global river hydrography and network routing: baseline data and new approaches to study the world’s large river systems. Hydrological Processes, 27(15): 2171–2186. Data is available at www.hydrosheds.org.</p>
Dataset and R code: Effects of temperature and air pollution on emergency ambulance dispatches: a time series analysis in a medium-sized city in Germany
<p>Dataset and R script to replicate results in the manuscript "Effects of temperature and air pollution on emergency ambulance dispatches: a time series analysis in a medium-sized city in Germany", currently under review.</p>
European Stillbirth Rate Time Series Dataset
<p>This dataset contains mostly annual time series of stillbirth rates from the mid-eighteenth century to the present for six countries including previously unpublished data for Denmark and for the Dutch region of Zeeland. There are several sheets within the Excel file, one of which explains the source of each series, potential stillbirth registration problems for each series and articles or book chapters with further information. This data was compiled by Eric Schneider from disparate sources and would not be possible without contributions from Anne Løkke (University of Copenhagen) and Frans van Poppel (NIDI).</p>
A tempοral Deep Convolutional Neural Network model on Sentinel-1 Image Time Series for pixel-wise Flood Classification (dataset)
<p>This is a dataset which has been designed to be used for flood time series classification. Each time series is annotated as flood or no-flood and represents a pixel-wise time series derived from stack of Sentinel-1 IW GRD images that have been pre-processed according to <a href="http://doi.org/10.5281/zenodo.6510223">https://doi.org/10.5281/zenodo.6510223</a>.</p>
Assessing land surface phenology in Araucaria-Nothofagus forests in Chile with Landsat 8/Sentinel-2 time series - Data and Material
<p>This dataset contains the Enhanced Vegetation Index (EVI) data used in our research work about land surface phenology of Andean Araucaria-Nothofagus forests as well as the phenology information derived from it.</p> <p>Study area: Conguillío National Park, Chile<br> Study period: 2016-2020</p> <p>Description of datasets:</p> <p>conguillio.sen2.lnd8.evi.2016.2020.nc - A raster dataset (NetCDF) of EVI values (resolution 10m). EVI was calculated from Level-2 Sentinel-2 and Landsat 8 data. To ensure harmonization, the Landsat 8 data was resampled and reprojected to Sentinel-2 properties prior to the index calculation.</p> <p>evi_gb_beck_white.tif - A raster dataset (GeoTiff) of phenological metrics per year (resolution 10m). Metrics were derived by fitting a double logistic function (see Beck et al., 2006) to the smoothed and interpolated EVI pixel time series. Subsequently, the main phenological variables SOS (start of season) and EOS (end of season) were extracted using a 50% threshold value. The dataset itself is a result of the R package "greenbrown" and the layers are named accordingly (see https://greenbrown.r-forge.r-project.org/phenology.php). It is available as GeoTIFF and as R rasterfile.</p> <p>Details about the methodology and results describing this dataset can be found in the following publication:<br> Kosczor, E., Forkel, M., Hernández, J., Kinalczyk, D., Pirotti, F. & Kutchartt, E., 2022. Assessing land surface phenology in Araucaria-Nothofagus forests in Chile with Landsat 8/Sentinel-2 time series. Int. J. Appl. Earth Obs. Geoinf. 112, 102862. https://doi.org/10.1016/j.jag.2022.102862</p>
Data associated with the manuscript "Simple statistical models can be sufficient for testing hypotheses with population time series data"
<p>This is a revised version of the archive of R code and data used in the manuscript, <em>Simple statistical models can be sufficient for testing hypotheses with population time series data. </em>The data are in three files. <em>etodata1.csv</em> and <em>etodata2.csv</em> contain two versions of the same data for shoal-dwelling fishes in the Etowah River and associated environmental covariates. <em>knz_dat</em> contains data for small mammals collected in the Konza Prairie Biological Station and associated environmental covariates. The R code consists of four primary files that call nine auxiliary files. CaseStudy1-main_code and CaseStudy2-main_code are the primary files for running the two case studies. Simulations1 and Simulations2 are the files for running the two batteries of simulations. We thank the Konza Prairie Biological Station and Konza Prairie Long-Term Ecological Research Program supported by the National Science Foundation (DEB-1440484) for collecting and providing access to mammal community data. More details are in the manuscript and supporting information. </p>
Outputs of the Jupyter Notebook - Concatenating a gridded rainfall reanalysis dataset into a time series
<p>The dataset contains the outputs of the notebook "Concatenating a gridded rainfall reanalysis dataset into a time series" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Timothy Lam (author), University of Exeter, <a href="https://github.com/timo0thy">@timo0thy</a></p> </li> <li> <p>Marlene Kretschmer (author), University of Reading, <a href="https://github.com/MarleneKretschmer">@MarleneKretschmer</a></p> </li> <li> <p>Samantha Adams (author), Met Office Informatics Lab, <a href="https://github.com/svadams">@svadams</a></p> </li> <li> <p>Rachel Prudden (author), Met Office Informatics Lab, <a href="https://github.com/RPrudden">@RPrudden</a></p> </li> <li> <p>Elena Saggioro (author), University of Reading, <a href="https://github.com/ESaggioro">@ESaggioro</a></p> </li> <li> <p>Nick Homer (reviewer), University of Edinburgh, <a href="https://github.com/NHomer">@NHomer</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>NOAA National Center for Environmental Prediction (creator)</p> </li> </ul> <p><em>Dataset authors</em></p> <ul> <li> <p>Eugenia Kalnay, Director, NCEP Environmental Modeling Center</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>E. Kalnay, M. Kanamitsu, R. Kistler, W. Collins, D. Deaven, L. Gandin, M. Iredell, S. Saha, G. White, J. Woollen, Y. Zhu, M. Chelliah, W. Ebisuzaki, W. Higgins, J. Janowiak, K. C. Mo, C. Ropelewski, J. Wang, A. Leetmaa, R. Reynolds, Roy Jenne, and Dennis Joseph. The ncep/ncar 40-year reanalysis project. Bulletin of the American Meteorological Society, 77(3):437 – 472, 1996. URL: <a href="https://journals.ametsoc.org/view/journals/bams/77/3/1520-0477_1996_077_0437_tnyrp_2_0_co_2.xml">https://journals.ametsoc.org/view/journals/bams/77/3/1520-0477_1996_077_0437_tnyrp_2_0_co_2.xml</a>, <a href="https://doi.org/10.1175/1520-0477(1996)077%3C0437:TNYRP%3E2.0.CO;2">doi:10.1175/1520-0477(1996)077<0437:TNYRP>2.0.CO;2</a>.</p> </li> </ul> <p><em>Pipeline documentation</em></p> <ul> <li> <p>Marlene Kretschmer, Samantha V. Adams, Alberto Arribas, Rachel Prudden, Niall Robinson, Elena Saggioro, and Theodore G. Shepherd. Quantifying causal pathways of teleconnections. Bulletin of the American Meteorological Society, 102(12):E2247 – E2263, 2021. URL: <a href="https://journals.ametsoc.org/view/journals/bams/102/12/BAMS-D-20-0117.1.xml">https://journals.ametsoc.org/view/journals/bams/102/12/BAMS-D-20-0117.1.xml</a>, <a href="https://doi.org/10.1175/BAMS-D-20-0117.1">doi:10.1175/BAMS-D-20-0117.1</a>.</p> </li> </ul>
Code and Data associated with "Discovery of positive and purifying selection in metagenomic time series of hypermutator microbial populations"
<p>Code and data sufficient to reproduce analyses in "Discovery of positive and purifying selection in metagenomic time series of hypermutator microbial populations".</p>
Respiration regimes in rivers: Partitioning source‐specific respiration from metabolism time series
<p>This repository hosts preprocessed input data and model codes of the paper:</p> <p>"Respiration regimes in rivers: Partitioning source‐specific respiration from metabolism time series" <br> Limonology and Oceanography, 2022.<br> DOI: 10.1002/lno.12207</p>
Single pixel s*t Landsat time series training data for CNN
<p>Single pixel s*t Landsat time series classification using 1D CNN</p> <p>Sep 22, 2022 update (version 2): <br> The 1D CNN classification codes are available at https://github.com/hankui/cnn_Landsat_time_series_classification_v2-Python</p> <p>The NLCD training data is available at 10.5281/zenodo.7106054</p> <p>The NLCD training data is derived from Landsat 5/7 analysis ready data (ARD) in year 2011 (as x predictor variable) and National Land Cover Database (NLCD) 2011 (as y response variable)</p> <p>The NLCD training data is distributed across Continental United States (CONUS) with 3,314,439 30m pixel locations</p> <p>The NLCD training data include (i) NLCD label with 15 classes, i.e., all NLCD classes except ice (https://www.mrlc.gov/data/legends/national-land-cover-database-class-legend-and-description)<br> (ii) year 2011 growing season Landsat ARD percentiles for Landsat 5/7 bands 2, 3, 4, 5 and 7 and for 8 band ratios derived from the five bands <br> (iii) percentiles include 10th, 20th, 25th, 30th, 35th, 40th, 50th (median), 60th, 65th, 70th, 75th, 80th, 90th so that <br> one can use 5 percentiles (10th, 25th, 50th, 75th, and 90th)<br> 7 percentiles (10th, 20th, 35th, 50th, 65th, 80th, and 90th)<br> 9 percentiles (10th, 20th, 30th, 40th, 50th, 60th, 70th, 80th, and 90th)<br> (iv) the pixel location represented in Landsat ARD tile h and v no. and the pixel i and j locations in the tile<br> (v) the no. of the cloud free observations in 2011 growing season derived for the pixel location<br> </p> <p>#*************************************************************************************************************#</p> <p>A munuscript describing how the data were derived and how the 1D CNN was adapted to the data is in review </p> <p><br> #*************************************************************************************************************#</p> <p>The codes were written in python (v3.7) and tensorflow (v2.6). </p> <p>The parameters are:</p> <p>(1) learning rate: cnn training initial learning rate 0.01 used in the paper </p> <p>(2) epoch: cnn training epochs 70 used in the paper </p> <p>(3) method: cnn training optimizer method 1: Adam method 2: dynamic learning rate used in the paper</p> <p>(4) L2: L2 regularization value; 0.001 used in paper </p> <p>(5) layer: no. of CNN layers (can be 4, 5 and 8) and 5 and 8 used in the paper</p> <p>(6) perc: training data percentages (can be 0.1, 0.5 and 0.9) tested in the paper; the evaluation is used the left 10% </p> <p>(7) gpui: which gpu process it will use (only applicable with multi-gpus) </p> <p>(8) IMG_HEIGHT: the no. of percentiles (can be 3, 5, 7 and 9) and 5, 7 and 9 used in the paper </p> <p>An example would be: </p> <p>version=7_4 </p> <p>layer=5; perc=0.1; gpui=0;IMG_HEIGHT=5</p> <p>method=0; learning_rate=0.01; epoch=10; iter=1; L2=0.001; sleep ${SLEEP}; ## Hank layer=5; perc=0.1; </p> <p>echo "python Pro_2d1d_CNN_v${version}.py ${learning_rate} ${epoch} ${method} ${L2} ${layer} ${perc} ${gpui} ${IMG_HEIGHT} "</p> <p>python Pro_2d1d_CNN_v${version}.py ${learning_rate} ${epoch} ${method} ${L2} ${layer} ${perc} ${gpui} ${IMG_HEIGHT} > layer${layer}.p${perc}.d${IMG_HEIGHT}.rate${learning_rate}.e${epoch}.L${L2}.v${version} & </p> <p><br> #*************************************************************************************************************#</p> <p>Aug 29, 2021 (version 1): <br> Training data: There are 2 input text files (csv) storing the 3,314,439 NLCD and 484,476 CDL land cover training samples:<br> NLCD training: ./NLCD/metric.ard.nlcd.Mar01.18.40.txt<br> CDL training: ./CDL/metric.ard.nlcd.Mar01.18.40.txt</p> <p>The codes and their usages are at: <br> https://github.com/hankui/cnn_Landsat_time_series_classification_v1-R<br> </p>
Library of multivariate time series
<p>A database of many different types of multivariate time series, each with between 5-25 processes and between 100-2500 observations.</p> <p>The database contains a serialized Python dictionary of 1053 datasets, where the key for the dictionary is the dataset name, and each value is another dictionary of: "data", an MxT numpy array of processes-by-observations; and "labels", a list of descriptive labels for the dataset.</p>
Apulian Aqueduct demo site: daily time series of reservoirs inflows 2010-2019
<p>This dataset contains the time series of net inflows to the main reservoirs of Apulian aqueduct (demo site 1), computed with a mass-balance equation at a daily time step.</p> <ul> <li>Temporal coverage: 2010-2019</li> <li>Spatial coverage: Reservoirs: Conza, Locone, Monte Cotugno, Occhito, Pertusillo</li> <li>Unit of measure: m<sup>3</sup>/sec</li> </ul> <p>More information and details on the content of this dataset can be found in Project Ô Deliverable D4.1.</p>
FrenchPiezo: the French mainland groundwater level multivariate time series
<p><strong>FrenchPiezo </strong>is the French mainland multivariate time series dataset of groundwater levels (a.k.a piezometric levels) across the French mainland. The dataset contains 1026 multivariate time series, each made of 3 dimensions which are:</p> <ul> <li>the piezometric level (<strong>p</strong>),</li> <li>the precipitation (<strong>t</strong><strong>p</strong>),</li> <li>the evapotranspiration (<strong>e</strong>).</li> </ul> <p>Each time series is identified by a code (<strong>bss</strong>) which is the identifier of the piezometer used to measure the piezometric level. The measurements are sampled daily from January 2015 to January 2021, corresponding to 2221 days. The piezometric levels are collected from <a href="https://hubeau.eaufrance.fr/">Hub'Eau</a>, the French free API for accessing water data. The climate data <strong>tp </strong>and <strong>e </strong>are collected from the Copernicus ERA5 archive using their free API. These data are in the file <em>dataset_2015_2021.csv</em><strong> </strong>in which time series with more than 50 missing values have been dropped. The file <em>dataset_2015_2021_nomissing_linear.csv</em> is the same dataset in which missing values have been imputed using linear interpolation.</p> <p>In addition to the climate data, information about the nature of the soil where each time series is collected are downloaded from <a href="https://bdlisa.eaufrance.fr/decouvrir-la-bdlisa"><em>BDLISA</em></a> and are given in the file <em>dataset_stations.csv </em>. BDLISA classifies soils as hydrologic entities (EH), characterized by a set of attributes described <a href="https://www.sandre.eaufrance.fr/ftp/documents/fr/ddd/saq/2.1/sandre_dictionnaire_SAQ__2.1.pdf">here</a>: </p> <p>More information on how we created the dataset, the scripts used to collect the data, and some experiments on forecasting the future values of groundwater levels using global and local foresting methods can be found on our <a href="https://github.com/frankl1/piezoforecast">GitHub page</a>.</p>
Train and test datasets used for the paper "Neural network time-series classifiers for gravitational-wave searches in single-detector periods"
<p>This repository contains the datasets used for training and testing during the work discussed in the paper "<a href="https://iopscience.iop.org/article/10.1088/1361-6382/ad40f0" target="_blank" rel="noopener">Neural network time-series classifiers for gravitational-wave searches in single-detector periods</a>". Please refer to this paper for more details on how the dataset was produced and cite it if you use these data:</p> <p><em>A. Trovato et al "Neural network time-series classifiers for gravitational-wave searches in single-detector periods", Class. Quant. Grav. 2024 DOI 10.1088/1361-6382/ad40f0.</em></p> <p>In this repository you will find six files in format npz, three of which refer to the test dataset and three to the train dataset. Each file name is of the type {label}_{train or test}.npz where "label" can be "glitch", "noise" or "signal", while the second part of the name indicates whether the file was used for training or testing.</p> <p>Each file is a collection of numpy arrays so it should be read with python. It contains 3 numpy arrays: 'X', 'Y' and 'metadata'. 'X' is a matrix containing 1-second segments of data sampled at 2048 Hz of the LIGO-Livingston detector, so it has shape: (number of samples, 2048). 'Y' contains the label for each segment, which is 0 for noise, 1 for signal and 2 for glitch, so it has shape: (number of samples,). In this case, the information on 'Y' is redundant since it's given directly by the filename. The 'metadata' matrix contains 17 metadata for each sample only for the case of signals, for glitch or noise it contains just 17 zeros for each sample. The shape of 'metadata' is thus: (number of samples, 17). For the signal files, for each sample the metadata is an array with these components:</p> <ol> <li>GPS start of the file from which this segment comes</li> <li>starting GPS time of this segment</li> <li>duration of the segment [s]</li> <li>mass1 [solar masses]</li> <li>mass2 [solar masses]</li> <li>spin1z</li> <li>spin2z</li> <li>inclination [radians]</li> <li>coalescence phase [radians]</li> <li>distance [Mpc]</li> <li>right_ascension [radians]</li> <li>declination [radians]</li> <li>polarization [radians]</li> <li>SNR (signal to noise ratio)</li> <li>shift of the signal w.r.t. the timeseries [s]</li> <li>length of the signal [s]</li> <li>fraction of the signal contained in the time window</li> </ol> <p>Number of samples:</p> <ul> <li>80000 for the file glitch_test.npz</li> <li>69998 for the file glitch_train.npz</li> <li>500000 for the file noise_test.npz</li> <li>250000 for the file noise_train.npz</li> <li>500000 for the file signal_test.npz</li> <li>250000 for the file signal_train.npz</li> </ul> <p>An example of few lines of python code to read each file is:</p> <pre><code>import numpy as np f = np.load("filename.npz") X = f['X'] Y = f['Y'] m = f['metadata'] </code></pre> <p>For the preparation of these data, we acknowledge the use of the following software packages: GWpy [1], PyCBC [2] and LALSuite [3]. </p> <p>This research has made use of data or software obtained from the Gravitational Wave Open Science Center (<a href="https://gwosc.org/" target="_blank" rel="noopener">gwosc.org</a>), a service of the LIGO Scientific Collaboration, the Virgo Collaboration, and KAGRA. This material is based upon work supported by NSF's LIGO Laboratory which is a major facility fully funded by the National Science Foundation, as well as the Science and Technology Facilities Council (STFC) of the United Kingdom, the Max-Planck-Society (MPS), and the State of Niedersachsen/Germany for support of the construction of Advanced LIGO and construction and operation of the GEO600 detector. Additional support for Advanced LIGO was provided by the Australian Research Council. Virgo is funded, through the European Gravitational Observatory (EGO), by the French Centre National de Recherche Scientifique (CNRS), the Italian Istituto Nazionale di Fisica Nucleare (INFN) and the Dutch Nikhef, with contributions by institutions from Belgium, Germany, Greece, Hungary, Ireland, Japan, Monaco, Poland, Portugal, Spain. KAGRA is supported by Ministry of Education, Culture, Sports, Science and Technology (MEXT), Japan Society for the Promotion of Science (JSPS) in Japan; National Research Foundation (NRF) and Ministry of Science and ICT (MSIT) in Korea; Academia Sinica (AS) and National Science and Technology Council (NSTC) in Taiwan.</p> <p>[1] https://gwpy.github.io<br>[2] https://pycbc.org<br>[3] https://lscsoft.docs.ligo.org/lalsuite</p>
Historical time-series reconstruction benchmark dataset of Landsat bi-monthly aggregates from GLAD ARD-2 at 30-m resolution with stratified sampling based on ESA CCI
<h2>Description</h2> <p>Historical time-series reconstruction benchmark dataset presented here is designed for evaluating and comparing the performance of time series reconstruction methods in the context of land cover change detection. The dataset is based on the European Space Agency Climate Change Initiative (ESA CCI) land cover dataset, which has been aggregated into 18 classes to facilitate analysis. The dataset includes information on land cover dynamics from 2000 to 2020, focusing on identifying and characterizing changes in land cover over time.</p> <h3><strong>Data Collection and Processing:</strong></h3> <p>The dataset is derived from the ESA CCI land cover dataset, which provides information on land cover classes at a global scale. The original dataset, containing 37 land cover classes, was aggregated into 18 classes based on similarity. Pixels with stable land cover over the study period and pixels with one or multiple land cover changes were identified and grouped into strata for sampling purposes.</p> <p>Sampling points were selected using a stratified sampling design, ensuring representation across different land cover classes and change scenarios. Approximately 2600 points were selected from each stratum, resulting in a total of 51,978 sampling points. The selected points were uniformly distributed along the strata, with spatial variations accounted for.</p> <p>Bimonthly time series data were extracted for each sampling point from 1997 to 2022, capturing temporal dynamics in land cover. Artificial gaps were introduced into the time series data to simulate real-world data loss, allowing for the evaluation of time series reconstruction methods under varying gap densities.</p> <p>The time series values were extracted from Landsat GLAD imagery using the specified spectral bands, including blue, green, red, NIR, SWIR1, SWIR2, and thermal bands. Additionally, a clear quality band was also extracted.</p> <h3>Data Details</h3> <ul> <li><strong>Time Period:</strong> 1997-01-01 to 2022-12-31</li> <li><strong>Type of Data: </strong>R data frame / Geopackage points.</li> <li><strong>Collection/Derivation:</strong> Derived from Landsat ARD v2, processed with Scikit-map.</li> <li><strong>Coordinate Reference System:</strong> EPSG:4326</li> <li><strong>Bounding Box:</strong> All the globe</li> <li><strong>File Format:</strong> RDS</li> </ul> <p> </p> <h3><strong>Reclassified Classes of ESA CCI Land Cover Dataset</strong></h3> <table> <tbody> <tr> <td> <div> <div> <p><strong>Aggregated Class Code</strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Aggregated Class Label</strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Original ESA CCI Classes</strong></p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>10</p> </div> </div> </td> <td> <div> <div> <p>Cropland rainfed</p> </div> </div> </td> <td> <div> <div> <p>10, 11, 12</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>30</p> </div> </div> </td> <td> <div> <div> <p>Mosaic cropland | natural vegetation</p> </div> </div> </td> <td> <div> <div> <p>30, 40</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>50</p> </div> </div> </td> <td> <div> <div> <p>Tree cover broadleaved evergreen</p> </div> </div> </td> <td> <div> <div> <p>50</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>60</p> </div> </div> </td> <td> <div> <div> <p>Tree cover broadleaved deciduous</p> </div> </div> </td> <td> <div> <div> <p>60, 61, 62</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>70</p> </div> </div> </td> <td> <div> <div> <p>Tree cover needleleaved evergreen</p> </div> </div> </td> <td> <div> <div> <p>70, 71, 72</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>80</p> </div> </div> </td> <td> <div> <div> <p>Tree cover needleleaved deciduous</p> </div> </div> </td> <td> <div> <div> <p>80, 81, 82</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>90</p> </div> </div> </td> <td> <div> <div> <p>Tree cover mixed leaf type</p> </div> </div> </td> <td> <div> <div> <p>90</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>100</p> </div> </div> </td> <td> <div> <div> <p>Mosaic tree and shrub | herbaceous cover</p> </div> </div> </td> <td> <div> <div> <p>100, 110</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>120</p> </div> </div> </td> <td> <div> <div> <p>Shrubland</p> </div> </div> </td> <td> <div> <div> <p>120, 121, 122</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>150</p> </div> </div> </td> <td> <div> <div> <p>Sparse vegetation</p> </div> </div> </td> <td> <div> <div> <p>150, 151, 152, 153</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>160</p> </div> </div> </td> <td> <div> <div> <p>Tree cover flooded</p> </div> </div> </td> <td> <div> <div> <p>160, 170</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>180</p> </div> </div> </td> <td> <div> <div> <p>Shrub or herbaceous cover flooded</p> </div> </div> </td> <td> <div> <div> <p>180</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>200</p> </div> </div> </td> <td> <div> <div> <p>Bare areas</p> </div> </div> </td> <td> <div> <div> <p>200, 201, 202</p> </div> </div> </td> </tr> </tbody> </table> <p>In the table, each row represents a reclassified land cover class, identified by a unique code. The 'Original ESA CCI Classes' column lists the specific land cover classes from the European Space Agency Climate Change Initiative dataset that are grouped together to form each broader category. Note that land cover classes not listed in this table were retained in their original value and were not reclassified.</p> <h3><strong>File Format</strong></h3> <p>The dataset comprises observations spanning from January 1997 to November 2022, capturing data for 51,978 samples.</p> <ul> <li>blue.rds: Time series data for the blue spectral band.</li> <li>green.rds: Time series data for the green spectral band.</li> <li>red.rds: Time series data for the red spectral band.</li> <li>nir.rds: Time series data for the near-infrared (NIR) spectral band.</li> <li>swir1.rds: Time series data for the shortwave infrared 1 (SWIR1) spectral band.</li> <li>swir2.rds: Time series data for the shortwave infrared 2 (SWIR2) spectral band.</li> <li>thermal.rds: Time series data for the thermal infrared band.</li> <li>clear.rds: Time series data for the clear quality band, used for masking out cloudy observations.</li> </ul> <p>How open the files in R:</p> <p><code>blue <- readRDS("blue.rds")</code></p> <p>To open the files in Python, you need to the <code>pyreadr</code> library:</p> <p><code>import pyreadr</code><br><code>blue = pyreadr.read_r('blue.rds')</code></p> <p> </p>
Station M time series study (NE Pacific) CTD data (cruises 2006-2022, surface to 4000 m depth)
<p>These datasets are from sensors mounted on remotely operated vehicle deployments (ROVs Tiburon and Doc Ricketts) to Station M (approx 4000 m) in the NE Pacific. Collection dates were from 2006 to 2022 as the ROV operated from the surface to the abyssal seafloor.</p> <p>The CTD was a Seabird SBE 21, Oxygen came from a pair of Seabird SBE 43s, Beam transmission from a Wetlabs C-Star, 25cm path, 720nm(red) </p> <div>21 and 43s were calibrated annually at Seabird, and the 43s were corrected a couple of times a year with bottle titration. </div> <div> </div> <div>Units:</div> <div> <table> <tbody> <tr> <td>depth (meters)</td> </tr> <tr> <td>heading (degrees)</td> </tr> <tr> <td>temperature (degrees C)</td> </tr> <tr> <td>salinity (unitless)</td> </tr> <tr> <td>oxygen (ml/l)</td> </tr> </tbody> </table> </div>
Data from: Evaluating Window Size Effects on Univariate Time Series Forecasting with Machine Learning
<p>In the realm of time series prediction modeling, the window size (w) is a critical hyperparameter that determines the number of time units included in each example provided to a learning model. This hyperparameter is crucial because it allows the learning model to recognize both long-term and short-term trends, as well as seasonal patterns, while reducing sensitivity to random noise. This study aims to elucidate the impact of window size on the performance of machine learning algorithms in univariate time series forecasting tasks. To achieve this, we employed 40 time series from two different domains, conducting experiments with varying window sizes using four types of machine learning algorithms: Bagging, Boosting, Stacking, and a Recurrent Neural Network (RNN) architecture. The results reveal that increasing the window size generally enhances the evaluation metric values up to a stabilization point, beyond which further increases do not significantly improve predictive accuracy. This stabilization effect was observed in both domains when w values exceeded 100 time steps. Moreover, the study found that RNN architectures do not consistently outperform ensemble models in various univariate time series forecasting scenarios.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.