Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

327

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

327 results for “Time-Series”

Learn how ShareScore rates datasets ↗
zenodo44/100

LAI_TS_Val: LAI time-series validation datasets in the 1-km pixel grid at global scale from 2001 to 2011

<p>Leaf area index (LAI), which is defined as one half of the total green leaf area per unit ground surface area, is a critical structural variable for quantifying the exchange processes of energy and matter between the land surface and atmosphere, it is thus identified as a key parameter in most terrestrial ecosystem models. To acquire long-term LAI records at the global scale, several remote sensing LAI products have been generated from various satellite sensors. However, assessing the uncertainties associated with these LAI products through comparisons with independent ground-truth measurements is pivotal for an effective application of products. Many sites from global networks have collected and provided invaluable ground LAI measurements covering a wide range of biome types and spatial variabilities. These site-based LAI measurements have been obtained about 30 years (1990-now). However, the spatial scale mismatch between site and pixel observations restricts the utilization of LAI measurements for product time-series validation. This datasets were generated from site-based LAI measurements of FLUXET and Chinese Ecosystem Research Network (CERN), using the proposed GUGM (Grading and Upscaling of Ground Measurements) method to resolve the scale-mismatch issue between site and sensor observations and maximize the utility of time-series of site-based LAI measurements, which can achieve the goal of product time-series validation. This GUGM approach first ingests both high-resolution images and site-based LAI measurements to capture the spatiotemporal variability in the product pixel grid. Then, a strategy was employed to grade the spatial representativeness of LAI measurements in the product pixel grid. For those LAI measurements which cannot be directly used in the validation of products, a strategy was adopted to calculate the spatial upscaling coefficient based on site-based LAI measurements and aggregated high-resolution reference maps to derive reliable LAI time-series validation datasets. The GUGM method has been applied to the site-based LAI measurements to generate global time-series LAI validation datasets from 2001 to 2011 in the 1 km pixel grid. The datasets include 28 sites which are mainly located in North America and Asia, providing 924 validation data in total. Among these sites, 16 sites with 508 (55.0%) validation data were obtained for forest, while 11 sites with 341 (36.9%) validation data and one site with 75 (8.1%) were obtained for crops and grasses, respectively. This datasets were saved in two formats: *.xls and *.kmz and each format was zipped for 63&nbsp;KB and 31 KB, respectively.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

1988-2009 time-series of land-use/land-cover maps for the Mar Menor / Campo de Cartagena watershed by means of supervised classification of Landsat images.

<p>Serie de mapas de usos y coberturas de la cuenca del Mar Menor (SE España): 2009, 2000, 1997 y 1998. Así como el documento completo de tesis en las que se generaron y analizaron.</p> <p>Time-series of land-use / land-cover maps of Mar Menor watershed (SE Spain): 2009, 2000, 1997 y 1998. As well as the complete thesis document in which they were generated and analyzed.</p>

opencc-by-4.0Oct 2015View details →
zenodo44/100

Long time-series ecological niche modelling using archaeological settlement data.

<p><strong>CR_settlement_niche_[N]_[Yr]_[BC/AD].tif</strong></p> <p>Ecological niche models in GeoTIFF format generated with the MaxEnt software based using prehistoric settlement evidence as training data and environmental layers (elevation, mean annual precipitation, mean annual temperature, landscape water balance, soil types) as background data. Raster values represent the probability of presence of a settlement.<br> <strong>N</strong> - chronological ordering<br> <strong>Yr, BC/AD</strong> - calendar years BC or AD</p> <p>&nbsp;</p> <p><strong>CR_settlement_niche_combined.tif</strong></p> <p>All models combined by averaging.</p> <p>&nbsp;</p> <p><strong>CR_settlement_archeo.zip</strong></p> <p>Archaeological data used to train the MaxEnt models in ESRI SHP format with the following fields:</p> <p><strong>Site_Type:</strong> Cemetery or Settlement</p> <p><strong>Archeo_Dat:</strong> Archaeological dating (culture or period)</p> <p><strong>Source:</strong> Source dataset (AMCR or LONGWOOD)</p> <p>AMCR: Archeologick&aacute; mapa Česk&eacute; republiky &ndash; Archaeological Map of the Czech Republic. Retrieved from https://digiarchiv.aiscr.cz/.</p> <p>LONGWOOD: Kol&aacute;ř, J., Tk&aacute;č, P., Macek, M., &amp; Szab&oacute;, P. (2016).&nbsp; Archaeology and Historical Ecology: the Archaeological Database of the LONGWOOD ERC Project. Arch&auml;ologisches Korrespondenzblatt 46/4, 539-554.</p> <p><strong>Yrs_BP_Avg:</strong> Average dating in calendar years BP (based on the archaeological dating)</p> <p><strong>Yrs_BP_Unc:</strong> Temporal uncertainty of the dating (half of the culture or period&#39;s duration)</p> <p><strong>Loc_Accur:</strong> Spatial accuracy derived from the recorded degree of the accuracy of location (radius in meters around the center point)</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Pacific salmon population time-series dataset to support Appendix S1: Data and additional information on declines of Pacific Salmon

<p>Dataset used to support the main paper &#39;Protecting our coast for everyone&rsquo;s future: Indigenous and scientific knowledge support marine spatial protections proposed by Central Coast First Nations in Pacific Canada&#39; by Reid et al. 2022. Dataset cited in Appendix S1 regarding trends in adult salmon abundances in the Central Coast. The data were as compiled by Will Atlas from the <a href="https://wildsalmoncenter.org/">Wild Salmon Center</a>&nbsp;to describe trends in the abundance of adult salmon returning to the Central Coast, which is the sum of escapement and harvest, as derived from the following sources:</p> <ol> <li>Escapement data from DFO: <a href="https://open.canada.ca/data/en/dataset/c48669a3-045b-400d-b730-48aafe8c5ee6">NuSEDS-New Salmon Escapement Database System - Open Government Portal (canada.ca)</a></li> <li>Harvest rates estimated by Karl English and colleagues and available at: <a href="https://data.salmonwatersheds.ca/data-library/">Salmon Watersheds Program - Data Library</a>.</li> <li>Information on total harvest that is reported in the DFO post season review (DFO 2020).</li> </ol>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Annual maps of cropland abandonment, land cover, and other derived data for time-series analysis of cropland abandonment

<p>This archive contains raw annual land cover maps, cropland abandonment maps, and accompanying derived data products to support:</p> <blockquote> <p>Crawford C.L., Yin, H., Radeloff, V.C., and Wilcove, D.S. 2022. Rural land abandonment is too ephemeral to provide major benefits for biodiversity and climate. <em>Science Advances</em>&nbsp;<a href="https://doi.org/10.1126/sciadv.abm8999">doi.org/10.1126/sciadv.abm8999</a><em>.</em></p> </blockquote> <p>An archive of the analysis scripts developed for this project can be found at: <a href="https://github.com/chriscra/abandonment_trajectories">https://github.com/chriscra/abandonment_trajectories</a> (<a href="https://doi.org/10.5281/zenodo.6383127">https://doi.org/10.5281/zenodo.6383127</a>).</p> <p>Note that the label &quot;_2022_02_07&quot; in many file names refers to the date of the primary analysis. &quot;dts&rdquo; or &ldquo;dt&rdquo; refer to &ldquo;data.tables,&quot; large .csv files that were manipulated using the data.table package in R (Dowle and Srinivasan 2021, <a href="http://r-datatable.com/">http://r-datatable.com/</a>). &ldquo;Rasters&rdquo; refer to &ldquo;.tif&rdquo; files that were processed using the raster and terra packages in R (Hijmans, 2022; <a href="https://rspatial.org/terra/">https://rspatial.org/terra/</a>; <a href="https://rspatial.org/raster">https://rspatial.org/raster</a>).</p> <p>Data files fall into one of four categories of data derived during our analysis of abandonment: <strong>observed</strong>, <strong>potential</strong>, <strong>maximum</strong>, or <strong>recultivation</strong>. Derived datasets also follow the same naming convention, though are aggregated across sites. These four categories are as follows (using &ldquo;age_dts&rdquo; for our site in Shaanxi Province, China as an example):</p> <ol> <li><strong>observed</strong> abandonment identified through our primary analysis, with a threshold of five years. These files do not have a specific label beyond the description of the file and the date of analysis (e.g., shaanxi_age_2022_02_07.csv);</li> <li><strong>potential</strong> abandonment for a scenario without any recultivation, in which abandoned croplands are left abandoned from the year of initial abandonment through the end of the time series, with the label &ldquo;_potential&rdquo; (e.g., shaanxi_potential_age_2022_02_07.csv);</li> <li><strong>maximum</strong> age of abandonment over the course of the time series, with the label &ldquo;_max&rdquo; (e.g., shaanxi_max_age_2022_02_07.csv);</li> <li><strong>recultivation </strong>periods, corresponding to the lengths of recultivation periods following abandonment, given the label &ldquo;_recult&rdquo; (e.g., shaanxi_recult_age_2022_02_07.csv).</li> </ol> <p>&nbsp;</p> <p><strong>This archive includes multiple .zip files, the contents of which are described below:</strong></p> <ul> <li><strong>age_dts.zip</strong> - Maps of abandonment age (i.e., how long each pixel has been abandoned for, as of that year, also referred to as length, duration, etc.), for each year between 1987-2017 for all 11 sites. These maps are stored as .csv files, where each row is a pixel, the first two columns refer to the x and y coordinates (in terms of longitude and latitude), and subsequent columns contain the abandonment age values for an individual year (where years are labeled with &quot;y&quot; followed by the year, e.g., &quot;y1987&quot;). Maps are given with a latitude and longitude coordinate reference system. Folder contains observed age, potential age (&ldquo;_potential&rdquo;), maximum age (&ldquo;_max&rdquo;), and recultivation lengths (&ldquo;_recult&rdquo;) for all sites. Maximum age .csv files include only three columns: x, y, and the maximum length (i.e., &ldquo;max age&rdquo;, in years) for each pixel throughout the entire time series (1987-2017). Files were produced using the custom functions &quot;cc_filter_abn_dt(),&quot;&nbsp;&ldquo;cc_calc_max_age(),&quot;&nbsp;&ldquo;cc_calc_potential_age(),&rdquo;&nbsp;and &ldquo;cc_calc_recult_age();&rdquo;&nbsp;see &quot;_util/_util_functions.R.&quot;</li> <li><strong>age_rasters.zip</strong> - Maps of abandonment age (i.e., how long each pixel has been abandoned for), for each year between 1987-2017 for all 11 sites. Maps are stored as .tif files, where each band corresponds to one of the 31 years in our analysis (1987-2017), in ascending order (i.e., the first layer is 1987 and the 31st layer is 2017). Folder contains observed age, potential age (&ldquo;_potential&rdquo;), and maximum age (&ldquo;_max&rdquo;) rasters for all sites. Maximum age rasters include just one band (&ldquo;layer&rdquo;). These rasters match the corresponding .csv files contained in &quot;age_dts.zip.&rdquo;</li> <li><strong>derived_data.zip</strong> - summary datasets created throughout this analysis, listed below.</li> <li><strong>diff.zip</strong> - .csv files for each of our eleven sites containing the year-to-year lagged differences in abandonment age (i.e., length of time abandoned) for each pixel. The rows correspond to a single pixel of land, and the columns refer to the year the difference is in reference to. These rows do not have longitude or latitude values associated with them; however, rows correspond to the same rows in the .csv files in &quot;input_data.tables.zip&quot; and &quot;age_dts.zip.&quot;&nbsp;These files were produced using the custom function &quot;cc_diff_dt()&quot; (much like the base R function &quot;diff()&quot;), contained within the custom function &quot;cc_filter_abn_dt()&quot; (see &quot;_util/_util_functions.R&quot;). Folder contains diff files for observed abandonment, potential abandonment (&ldquo;_potential&rdquo;), and recultivation lengths (&ldquo;_recult&rdquo;) for all sites.</li> <li><strong>input_dts.zip</strong> - annual land cover maps for eleven sites with four land cover classes (see below), adapted from Yin et al. 2020 <em>Remote Sensing of Environment </em>(<a href="https://doi.org/10.1016/j.rse.2020.111873">https://doi.org/10.1016/j.rse.2020.111873</a>)<em>. </em>Like &ldquo;age_dts,&rdquo; these maps are stored as .csv files, where each row is a pixel and the first two columns refer to x and y coordinates (in terms of longitude and latitude). Subsequent columns contain the land cover class for an individual year (e.g., &quot;y1987&quot;). Note that these maps were recoded from Yin et al. 2020 so that land cover classification was consistent across sites (see below). This contains two files for each site: the raw land cover maps from Yin et al. 2020 (after recoding), and a &ldquo;clean&rdquo; version produced by applying 5- and 8-year temporal filters to the raw input (see custom function &ldquo;cc_temporal_filter_lc(),&rdquo;&nbsp;in &ldquo;_util/_util_functions.R&rdquo; and &ldquo;1_prep_r_to_dt.R&rdquo;). These files correspond to those in &quot;input_rasters.zip,&quot; and serve as the primary inputs for the analysis.</li> <li><strong>input_rasters.zip</strong> - annual land cover maps for eleven sites with four land cover classes (see below), adapted from Yin et al. 2020 <em>Remote Sensing of Environment. </em>Maps are stored as &quot;.tif&quot; files, where each band corresponds one of the 31 years in our analysis (1987-2017), in ascending order (i.e., the first layer is 1987 and the 31st layer is 2017). Maps are given with a latitude and longitude coordinate reference system. Note that these maps were recoded so that land cover classes matched across sites (see below). Contains two files for each site: the raw land cover maps (after recoding), and a &ldquo;clean&rdquo; version that has been processed with 5- and 8-year temporal filters (see above). These files match those in &quot;input_dts.zip.&quot;</li> <li><strong>length.zip</strong> - .csv files containing the length (i.e., age or duration, in years) of each distinct individual period of abandonment at each site. This folder contains length files for observed and potential abandonment, as well as recultivation lengths. Produced using the custom function &quot;cc_filter_abn_dt()&quot; and &ldquo;cc_extract_length();&rdquo;&nbsp;see &quot;_util/_util_functions.R.&quot;</li> </ul> <p><strong>derived_data.zip</strong> contains the following files:</p> <ul> <li>&quot;<strong>site_df.csv</strong>&quot; - a simple .csv containing descriptive information for each of our eleven sites, along with the original land cover codes used by Yin et al. 2020 (updated so that all eleven sites in how land cover classes were coded; see below).</li> <li><strong>Primary derived datasets </strong>for both observed abandonment (&ldquo;area_dat&rdquo;) and potential abandonment (&ldquo;potential_area_dat&rdquo;). <ul> <li><strong>area_dat</strong> - Shows the area (in ha) in each land cover class at each site in each year (1987-2017), along with the area of cropland abandoned in each year following a five-year abandonment threshold (abandoned for &gt;=5 years) or no threshold (abandoned for &gt;=1 years). Produced using custom functions &quot;cc_calc_area_per_lc_abn()&quot; via &quot;cc_summarize_abn_dts()&quot;. See scripts &quot;cluster/2_analyze_abn.R&quot; and &quot;_util/_util_functions.R.&quot;</li> <li><strong>persistence_dat</strong> - A .csv containing the area of cropland abandoned (ha) for a given &quot;cohort&quot; of abandoned cropland (i.e., a group of cropland abandoned in the same year, also called &quot;year_abn&quot;) in a specific year. This area is also given as a proportion of the initial area abandoned in each cohort, or the area of each cohort when it was first classified as abandoned at year 5 (&quot;initial_area_abn&quot;). The &quot;age&quot; is given as the number of years since a given cohort of abandoned cropland was last actively cultivated, and &quot;time&quot; is marked relative to the 5th year, when our five-year definition first classifies that land as abandoned (and where the proportion of abandoned land remaining abandoned is 1). Produced using custom functions &quot;cc_calc_persistence()&quot; via &quot;cc_summarize_abn_dts()&quot;. See scripts &quot;cluster/2_analyze_abn.R&quot; and &quot;_util/_util_functions.R.&quot;&nbsp;This serves as the main input for our linear models of recultivation (&ldquo;decay&rdquo;) trajectories.</li> <li><strong>turnover_dat</strong> - A .csv showing the annual gross gain, annual gross loss, and annual net change in the area (in ha) of abandoned cropland at each site in each year of the time series. Produced using custom functions &quot;cc_calc_abn_diff()&quot; via &quot;cc_summarize_abn_dts()&quot; (see &quot;_util/_util_functions.R&quot;), implemented in &quot;cluster/2_analyze_abn.R.&quot;&nbsp;This file is only produced for observed abandonment.</li> </ul> </li> <li><strong>Area summary files </strong>(for observed abandonment only) <ul> <li><strong>area_summary_df</strong> - Contains a range of summary values relating to the area of cropland abandonment for each of our eleven sites. All area values are given in hectares (ha) unless stated otherwise. It contains 16 variables as columns, including 1) &quot;site,&quot; 2) &quot;total_site_area_ha_2017&quot; - the total site area (ha) in 2017, 3) &quot;cropland_area_1987&quot; - the area in cropland in 1987 (ha), 4) &quot;area_abn_ha_2017&quot; - the area of cropland abandoned as of 2017 (ha), 5) &quot;area_ever_abn_ha&quot; - the total area of those pixels that were abandoned at least once during the time series (corresponding to the area of potential abandonment, as of 2017), 6) &quot;total_crop_extent_ha&quot; - the total area of those pixels that were classified as cropland at least once during the time series, 7) &quot;total_area_abn_remaining_2017&quot; - duplicate of &quot;area_abn_ha_2017,&quot; the area abandoned as of 2017 (ha), taken from &quot;area_recult_threshold,&quot; 8) &quot;total_initial_area_abn&quot; - the sum of the initial area of each cohort of abandonment when it is first classified as &quot;abandoned,&quot; i.e., at the 5 year mark (note that this is cumulative, and because it counts those pixels that were abandoned more than once, it is therefore larger than &quot;area_ever_abn_ha&quot;), taken from &quot;area_recult_threshold&quot; 9) &quot;total_area_abn_recultivated_2017&quot; - the area of abandoned land that was recultivated as of 2017 (cumulatively, i.e., &quot;total_initial_area_abn&quot; - &quot;area_abn_ha_2017&quot;), taken from &quot;area_recult_threshold,&quot; 10) &quot;proportion_recultivated&quot; - the proportion of all abandoned cropland (including multiple periods per pixel) that was recultivated by 2017, taken from &quot;area_recult_threshold,&quot; 11) &quot;area_2017_as_prop_site&quot; - area abandoned as of 2017 as a proportion of the total site area, 12) &quot;area_2017_as_prop_total_crop&quot; - area abandoned as of 2017 as a proportion of the total crop extent, 13) &quot;area_2017_as_prop_crop87&quot; - area abandoned as of 2017 as a proportion of cropland area in 1987, 14) &quot;area_ever_abn_as_prop_site&quot; - area ever abandoned as a proportion of the total site area, 15) &quot;area_ever_abn_as_prop_total_crop&quot; - area ever abandoned as a proportion of the total crop extent, 16) &quot;area_ever_abn_as_prop_crop87&quot; - area ever abandoned as a proportion of cropland area in 1987. See script &quot;1_summary_stats.Rmd.&quot;</li> <li><strong>area_recult_threshold</strong> - Contains data on the proportion of observed abandoned cropland area that is recultivated by the end of our time series. This includes the area of abandoned cropland as of 2017 (&quot;total_area_abn_remaining_2017&quot;) and the sum of the initial area of each cohort of abandonment when it is first classified as abandoned (at year 5; &quot;total_initial_area_abn&quot;). This &quot;total_initial_area_abn&quot; is cumulative, and allows for pixels that were abandoned multiple times during the time series to be counted multiple times. The difference between these two columns yields the &quot;total_area_abn_recultivated_2017,&quot;&nbsp;which in turn is used to calculate the &quot;proportion_recultivated,&quot;&nbsp;and the (ascending) &quot;order&quot; of sites based on this proportion. This file includes recultivation stats for each site for three abandonment definitions: 5, 7, and 10 years. See script &quot;1_summary_stats.Rmd.&quot;</li> <li><strong>abn_lc_area_2017</strong> - Contains the number of pixels and corresponding area (in ha) of abandoned cropland in the year 2017 at each site, according to the land cover class (either woody vegetation [2], or herbaceous vegetation [4]) and the age in 2017 (5 to 30 years). See script &quot;cluster/6_lc_of_abn.R.&quot;</li> <li><strong>abn_prop_lc_2017 </strong>- Contains the number of pixels and corresponding area (ha) of cropland abandoned in the year 2017 in each land cover type (woody vegetation [2], or herbaceous vegetation [4]). It also shows this area as a proportion of the total area abandoned at each site (i.e., in either land cover class: 2 or 4). See script &quot;cluster/6_lc_of_abn.R.&quot;</li> </ul> </li> <li><strong>Carbon</strong> <ul> <li><strong>carbon_df </strong>&ndash; contains the observed and potential carbon accumulation in abandoned croplands in each site in each year (in Mg C), for two abandonment thresholds: 5 years (our default abandonment definition) and 1 year (i.e., no threshold). Each data point corresponds to one of two scenarios (&ldquo;type&rdquo; column), either &ldquo;observed&rdquo; or &ldquo;potential.&rdquo; Carbon accumulation figures are for both the sum of forest and soil carbon at each site in a given year. Carbon accumulation is listed in three columns: 1) &ldquo;C_up_to_20&rdquo; contains the total carbon accumulated in those abandoned croplands with abandonment durations between 5 and 20 years. 2) &ldquo;C_21_30&rdquo; contains the total carbon accumulation in croplands with durations between 21 and 30 years, which are differentiated in order to account for non-linear carbon accumulation rates in soils over time, and 3) &ldquo;total_C_Mg&rdquo; contains the sum of the previous two columns, representing the total carbon accumulated across all abandoned croplands in each year.</li> <li><strong>soc_mean</strong> &ndash; contains mean soil organic carbon accumulation rates for years 1-20 and years 21-80, derived from Sanderman et al. 2020 (in Mg C; <a href="https://doi.org/10.7910/DVN/HA17D3">https://doi.org/10.7910/DVN/HA17D3</a>). These values correspond to accumulation rates in croplands upon abandonment and regeneration to natural vegetation (Sanderman et al. 2020&rsquo;s &ldquo;rewilding&rdquo; scenario). These mean values are calculated across those pixels identified as cropland by Sanderman et al. 2020 at each site. Mean values in year 20 and 80 are contained in columns &ldquo;mean_soc_20&rdquo; and &ldquo;mean_soc_80&rdquo; respectively, and the annualized rate over the first 20 years and the subsequent years 21 through 80 are contained in columns &ldquo;mean_annual_soc_1_20&rdquo; and &ldquo;mean_annual_soc_21_80&rdquo; respectively.</li> </ul> </li> <li><strong>Decay model data</strong> &ndash; two R data files containing data products for our linear models of abandonment recultivation trajectories. <ul> <li><strong>decay_endpoints_files</strong> &ndash; an R data file (.rds) containing seven data products produced as part of our common endpoint analysis, which calculated mean trajectories for each site across a range of common endpoints, ensuring that means were based on coefficient estimates derived from a consistent number of observations for each cohort. These files are: <ul> <li><strong>common_endpoint_dat &ndash; </strong>a .csv containing subsets of &ldquo;persistence_dat&rdquo; for each &ldquo;endpoint&rdquo; (7 through 29).</li> <li><strong>endpoint_n &ndash; </strong>a .csv describing, for each endpoint, the corresponding number of observations per cohort (&ldquo;n_obs&rdquo;), the number of cohorts (&ldquo;n_cohorts&rdquo;), the total number of observations across cohorts included (&ldquo;total_obs&rdquo;), and the cohorts that meet the endpoint threshold (&ldquo;cohorts&rdquo;).</li> <li><strong>coef_l3_endpoints &ndash; </strong>corresponding model coefficients for our primary model (&ldquo;l3&rdquo;) parameterized by the range of subsets across endpoints.</li> <li><strong>augment_endpoints &ndash; </strong>fitted values (i.e., model predictions) for linear models produced across the full range of endpoint subsets.</li> <li><strong>fitted_endpoints &ndash; </strong>a simplified .csv containing the mean linear and log coefficients for each site at each endpoint, and the corresponding predicted proportion remaining abandoned through time (based on the &ldquo;age,&rdquo; or duration, of abandonment).</li> <li><strong>time_to_endpoints &ndash; </strong>a .csv containing, for mean trajectories for each endpoint at each site, the estimated time required for a given amount of abandoned cropland in a cohort to be recultivated (deciles, 10% through 100%).</li> <li><strong>endpoint_half_lives &ndash; </strong>a .csv containing the half-lives calculated for the mean trajectories for each endpoint at each site.</li> </ul> </li> <li><strong>decay_mod_archive</strong> - an R data file (.rds) containing eleven data products derived from linear models of abandonment recultivation (&quot;decay&quot;): <ul> <li><strong>lm_mega_lin_log_lin_l</strong> &ndash; the primary linear model produced in our analysis. This model is referred to as &ldquo;lin_log_lin&rdquo; (or &ldquo;l3&rdquo;) because the model predicts linear persistence (&ldquo;lin&rdquo;) as a function of a log term of time (&ldquo;log&rdquo;) and a linear term of time (&ldquo;lin&rdquo;). &ldquo;mega&rdquo; refers to the fact that this model is run for the full dataset, pooled across all 11 sites.</li> <li><strong>coef_l3_mega</strong> &ndash; a .csv containing model coefficients for our primary linear model of recultivation (&ldquo;lin_log_lin&rdquo;, or &ldquo;l3&rdquo;), with a single row each for the linear term of time and the log term of time, for 26 cohorts at 11 sites.</li> <li><strong>mean_coef_l3_mega</strong> &ndash; a data frame containing the mean coefficient values for the log and linear terms of time across cohorts at each site. This also contains the mean of the low and high coefficient estimates, based on the 95% confidence interval.</li> <li><strong>half_lives_all_cohorts_l3</strong> &ndash; half-lives calculated for each cohort at each site, for our primary model.</li> <li><strong>half_life_mean_coefs_l3</strong> &ndash; half-lives calculated based on the mean trajectory for each site (based on the mean log coefficients and mean linear coefficients across all cohorts), for our primary model.</li> <li><strong>mod_AIC_mega</strong> &ndash; Akaike Information Criterion (AIC) values for all tested model specifications.</li> <li><strong>fitted_combo</strong> &ndash; fitted values (i.e., model predictions) for our primary model (&ldquo;l3&rdquo;) and a series of alternative model specifications (&ldquo;l3_trim&rdquo; &ndash; excluding cohorts with fewer than 5 observations; &ldquo;lin_log&rdquo; &ndash; a model including only one log time term; &ldquo;log2_lin&rdquo; &ndash; in which the log of persistence is predicted by log and linear time terms; and &ldquo;l3_no_cohort&rdquo; &ndash; our primary model, predicting linear persistence as a function of log time and linear time, but without cohort-level fixed effects).</li> <li><strong>time_to_combo</strong> &ndash; contains the estimated time required for a certain amount of abandoned cropland in a cohort to be recultivated (deciles, 10% through 100%). See script &quot;2_decay_models.Rmd.&quot;&nbsp;These values are calculated for a range of alternative model specifications (&quot;l3_trim&quot;, &ldquo;lin_log&rdquo;, &quot;log2_lin&quot;, and &quot;l3_no_cohort&quot;; see above).</li> </ul> </li> </ul> </li> <li><strong>Length data</strong> &ndash; includes &ldquo;_distill_df&rdquo; files and &ldquo;mean_length_df&rdquo; files for observed, potential, and recultivation. <ul> <li><strong>length_distill_df</strong> - .csvs containing the number (&quot;freq&quot;) of abandonment periods of a specific &quot;length&quot; of time (i.e., age) at each site over the course of the entire time series. Derived from the &quot;length&quot; files in &quot;length.zip.&quot;&nbsp;See script &quot;cluster/5_distill_lengths.R.&quot;</li> <li><strong>mean_length_df</strong> - .csvs with the mean, median, and standard deviation, for each site, for both &quot;all&quot; lengths or just the &quot;max&quot; length per pixel, and for a range of abandonment definitions (1, 3, 5, 7, and 10 years). Derived from &quot;length_distill_df.&quot;&nbsp;See script &quot;1_summary_stats.Rmd.&quot;</li> </ul> </li> <li><strong>Duration summary files</strong> &ndash; includes &ldquo;summary_stats_all_sites&rdquo; and &ldquo;summary_stats_all_sites_pooled,&rdquo; for observed and potential abandonment, and recultivation periods following abandonment. <ul> <li><strong>&ldquo;summary_stats_all_sites&rdquo;</strong> - A simple .csv derived from &quot;mean_length_df&quot; files containing summary stats across the 11 sites. This includes the mean of the mean abandonment duration (&quot;length&quot;, in years) for each of our 11 sites (&quot;mean_of_means&quot;), the standard deviation of these site mean abandonment lengths (&quot;sd_of_means&quot;), the mean of the standard deviation at each site (&quot;mean_of_sds&quot;), the mean median (&quot;mean_of_medians&quot;), and the mean number of abandonment periods (&quot;mean_n_abn_periods&quot;). Note that length &quot;all&quot; indicates that these stats account for all periods (including multiple per pixel), rather than just the max duration per pixel. See script &quot;1_summary_stats.Rmd.&quot;</li> <li><strong>&ldquo;summary_stats_all_sites_pooled&rdquo;</strong> - A summary .csv similar to &quot;summary_stats_all_sites,&quot;&nbsp;but calculated by pooling all distinct periods of abandonment across all eleven sites, and then calculating the mean, median, and standard deviation of abandonment duration. See script &quot;1_summary_stats.Rmd.&quot;</li> </ul> </li> <li><strong>Comparing annual approach to identifying abandonment to a two-timepoint (&ldquo;2yr&rdquo;) approach:</strong> <ul> <li><strong>abn_2yr_ages_df</strong> - Contains the age of former croplands identified as &quot;abandoned&quot; using a two-timepoint method (i.e., 2017 - 1987), where age values (as of 2017) are derived from our map of abandonment identified using the full annual time series. This includes the area in hectares (ha), in each age class (along with the number of pixels), at each of our 11 sites. This dataset is used to calculate the percent of cropland &quot;abandonment&quot; identified using the two-year method that is actually too &quot;young,&quot; i.e., less than 5 years old, and therefore not truly abandonment according to our five-year abandonment definition</li> <li><strong>abn_2yr_overestimation</strong> - Compares the area (in hectares) of cropland abandonment at each site identified with our full annual time series (and a five-year abandonment definition) and the &quot;abandonment&quot; identified using a two-timepoint method (2017-1987). This also includes the percent difference in area between the two methods, the Jaccard similarity of the areas identified as abandonment, and the percent of &quot;young&quot; (i.e., &lt;5-year-old) &quot;abandonment&quot; identified by the two-timepoint method.</li> </ul> </li> </ul> <p><strong>Input land cover maps:</strong></p> <p>As noted, the file &quot;input_rasters.zip&quot; contain the raw annual land cover maps for eleven sites generated by:</p> <blockquote> <p>Yin, H., A. Brand&atilde;o, J. Buchner, D. Helmers, B. G. Iuliano, N. E. Kimambo, K. E. Lewińska, E. Razenkova, A. Rizayeva, N. Rogova, S. A. Spawn, Y. Xie, and V. C. Radeloff. 2020. Monitoring cropland abandonment with Landsat time series. <em>Remote Sensing of Environment</em> 246:111873.&nbsp;https://doi.org/10.1016/j.rse.2020.111873</p> </blockquote> <p>These land cover maps served as raw inputs for this project and form the basis of the analysis.</p> <p>All land cover maps have a resolution of 30-m and exist for each year from 1987 through 2017. The exceptions are Nebraska / Wyoming (1986-2018) and Wisconsin (1987-2018); these additional years were excluded from our analysis of abandonment duration.</p> <p><strong>Land cover categories in these maps are coded as follows:</strong></p> <ol> <li>Non-vegetated area (e.g., water, urban, barren land)</li> <li>Woody vegetation (e.g., forests)</li> <li>Cropland</li> <li>Herbaceous vegetation (e.g., grassland)</li> </ol> <p><strong>Site file names correspond to the following geographic locations:</strong></p> <ul> <li>belarus = Vitebsk, Belarus / Smolensk, Russia</li> <li>bosnia_herzegovina = Bosnia &amp; Herzegovina</li> <li>chongqing = Chongqing, China</li> <li>goias = Goi&aacute;s, Brazil</li> <li>iraq = Iraq</li> <li>mato_grosso = Mato Grosso, Brazil</li> <li>nebraska = Nebraska / Wyoming, USA</li> <li>orenburg = Orenburg, Russia / Uralsk, Kazakhstan</li> <li>shaanxi = Shaanxi/Shanxi, China</li> <li>volgograd = Volgograd, Russia</li> <li>wisconsin = Wisconsin, USA</li> </ul> <p>This dataset is minimally altered from Yin et al. 2020.&nbsp; However, land cover codes were updated for five sites (Iraq, Nebraska/Wyoming, Orenburg/Uralsk, Volgograd, and Wisconsin) in order to maintain consistency in how land cover was coded across all sites. The original land cover codes (matching Yin et al. 2020) are described in the file &quot;site_df.csv&quot; and are as follows:</p> <ol> <li>Iraq: 1 Non-vegetated;&nbsp; 2 Cropland;&nbsp; 3 Woody;&nbsp; 4 Herbaceous</li> <li>Nebraska / Wyoming (USA):&nbsp; 1 Cropland;&nbsp; 2 Woody;&nbsp; 3 Non-vegetated;&nbsp; 4 Herbaceous</li> <li>Orenburg, Russia / Uralsk, Kazakhstan:&nbsp; 1 Non-vegetated;&nbsp; 2 Cropland;&nbsp; 3 Herbaceous;&nbsp; 4 Woody</li> <li>Volgograd (Russia):&nbsp; 1 Non-vegetated;&nbsp; 2 Cropland;&nbsp; 3 Herbaceous;&nbsp; 4 Woody</li> <li>Wisconsin (USA):&nbsp; 1 Cropland;&nbsp; 2 Herbaceous;&nbsp; 3 Woody;&nbsp; 4 Non-vegetated</li> </ol>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Sharkipedia: A Curated Open Access Database of Shark and Ray Life History Traits and Abundance Time-series

<p>This dataset represent the intial launch of Sharkipedia: a curated open access database of shark and ray life history traits and abundance time-series. A curated database of shark and ray biological data is increasingly necessary both to support fisheries management and conservation efforts, and to test the generality of hypotheses of vertebrate macroecology and macroevolution. Sharks and rays are one of the most charismatic, evolutionary distinct, and threatened lineages of vertebrates, comprising around 1,250 species. To accelerate shark and ray conservation and science, we developed Sharkipedia as a curated open-source database and research initiative to make all published biological traits and population trends accessible to everyone. Sharkipedia hosts information on 58 life history traits from 264 sources, for 170 species, from 39 families, and 12 orders related to length (n=9 traits), age (8), growth (12), reproduction (19), demography (5), and allometric relationships (5), as well as 871 population time-series from 202 species. Sharkipedia relies on the backbone taxonomy of the IUCN Red List and the bibliography of Shark-References. Sharkipedia has profound potential to support the rapidly growing data demands of fisheries management, international trade regulation as well as anchoring vertebrate macroecology and macroevolution.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

ERA-NUTS: meteorological time-series based on C3S ERA5 for European regions (1980-2021)

<p><strong># ERA-NUTS (1980-2021)</strong></p> <p>This dataset contains a set of time-series of meteorological variables based on <a href="https://climate.copernicus.eu/climate-reanalysis">Copernicus Climate Change Service (C3S) ERA5 reanalysis</a>. The data files can be downloaded from here while notebooks and other files can be found on the <a href="https://github.com/energy-modelling-toolkit/era-nuts-code">associated Github repository</a>.</p> <p>This data has been generated with the aim of providing hourly time-series of the <strong>meteorological variables</strong> commonly used for power system modelling and, more in general, studies on energy systems.</p> <p>An example of the analysis that can be performed with ERA-NUTS is shown <a href="https://youtu.be/zVeF8Dv6jlE">in this video</a>.</p> <p><strong>Important</strong>: <em>this dataset is still a work-in-progress, we will add more analysis and variables in the near-future. If you spot an error or something strange in the data please tell us <a href="mailto:matteo.de-felice@ec.europa.eu">sending an email</a> or opening an Issue in the <a href="https://github.com/energy-modelling-toolkit/era-nuts-code">associated Github repository</a>.</em></p> <p><strong>## Data</strong><br> The time-series have hourly/daily/monthly frequency and are aggregated following the <a href="https://ec.europa.eu/eurostat/web/nuts/background">NUTS&nbsp; 2016 classification</a>. NUTS (Nomenclature of Territorial Units for Statistics) is a European Union standard for referencing the subdivisions of countries (member states, candidate countries and EFTA countries).</p> <p>This dataset contains NUTS0/1/2 time-series for the following variables obtained from the <strong>ERA5 reanalysis data</strong> (in brackets the name of the variable on the Copernicus Data Store and its unit measure):</p> <p>&nbsp; - <strong>t2m</strong>: 2-meter temperature (`2m_temperature`, Celsius degrees)<br> &nbsp; - <strong>ssrd</strong>: Surface solar radiation (`surface_solar_radiation_downwards`, Watt per square meter)<br> &nbsp; - <strong>ssrdc</strong>: Surface solar radiation clear-sky (`surface_solar_radiation_downward_clear_sky`, Watt per square meter)<br> &nbsp; - <strong>ro</strong>: Runoff (`runoff`, millimeters)<br> &nbsp; -&nbsp;<strong>sd</strong>: Snow depth (`sd`, meters)<br> &nbsp;<br> There are also a set of derived variables:<br> &nbsp; - <strong>ws10</strong>: Wind speed at 10 meters (derived by `10m_u_component_of_wind` and `10m_v_component_of_wind`, meters per second)<br> &nbsp; - <strong>ws100</strong>: Wind speed at 100 meters (derived by `100m_u_component_of_wind` and `100m_v_component_of_wind`, meters per second)<br> &nbsp; - <strong>CS</strong>: Clear-Sky index (the ratio between the solar radiation and the solar radiation clear-sky)<br> &nbsp; - <strong>RH</strong>: Relative Humidity (computed following Lawrence, BAMS 2005 and Alduchov &amp; Eskridge, 1996)<br> &nbsp; - <strong>HDD</strong>/<strong>CDD</strong>: Heating/Cooling Degree days (derived by 2-meter temperature the <a href="https://ec.europa.eu/eurostat/cache/metadata/en/nrg_chdd_esms.htm">EUROSTAT definition</a>.</p> <p>For each variable we have <strong>367 440 hourly samples</strong> (from 01-01-1980 00:00:00 to 31-12-2021&nbsp;23:00:00) for <strong>34/115/309 regions</strong> (NUTS 0/1/2).<br> &nbsp;<br> The data is provided in two formats:</p> <p>&nbsp; - NetCDF version 4 (all the variables hourly and CDD/HDD daily). NOTE: the variables are stored as `int16` type using a `scale_factor` to minimise the size of the files.<br> &nbsp; - Comma Separated Value (&quot;single index&quot; format for all the variables and the time frequencies and &quot;stacked&quot; only for daily and monthly)<br> &nbsp;<br> All the CSV files are stored in a zipped file for each variable.</p> <p><strong>## Methodology</strong></p> <p>The time-series have been generated using the following workflow:</p> <p>&nbsp; 1. The NetCDF files are downloaded from the Copernicus Data Store from the <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels?tab=form">ERA5 hourly data on single levels from 1979 to present</a> dataset<br> &nbsp; 2. The data is read in R with the <a href="http://www.meteo.unican.es/climate4R">climate4r</a> packages and aggregated using the function `/get_ts_from_shp` from <a href="https://github.com/matteodefelice/panas">panas</a>. All the variables are aggregated at the NUTS boundaries using the average except for the runoff, which consists of the sum of all the grid points within the regional/national borders.<br> &nbsp; 3. The derived variables (wind speed, CDD/HDD, clear-sky) are computed and all the CSV files are generated using R<br> &nbsp; 4. The NetCDF are created using `xarray` in Python 3.8.</p> <p><strong>## Example notebooks</strong></p> <p>In the folder `notebooks` on the <a href="https://github.com/energy-modelling-toolkit/era-nuts-code">associated Github repository</a> there are two Jupyter notebooks which shows how to deal effectively with the NetCDF data in `xarray` and how to visualise them in several ways by using matplotlib or the <a href="https://github.com/kavvkon/enlopy">enlopy</a> package.</p> <p>There are currently two notebooks:</p> <p>&nbsp; - <strong>exploring-ERA-NUTS</strong>: it shows how to open the NetCDF files (with Dask), how to manipulate and visualise them.<br> &nbsp; - <strong>ERA-NUTS-explore-with-widget</strong>: explorer interactively the datasets with [<a href="https://jupyter.org/">jupyter</a>]() and <a href="https://ipywidgets.readthedocs.io/en/stable/">ipywidgets</a>.</p> <p>The notebook `exploring-ERA-NUTS` is also available rendered as HTML.<br> <br> <strong>## Additional files</strong></p> <p>In the folder `additional files`on the <a href="https://github.com/energy-modelling-toolkit/era-nuts-code">associated Github repository</a> there is a map showing the spatial resolution of the ERA5 reanalysis and a CSV file specifying the number of grid points with respect to each NUTS0/1/2 region.</p> <p><strong>## License</strong></p> <p>This dataset is released under <a href="https://creativecommons.org/licenses/by/4.0/">CC-BY-4.0 license</a>.</p> <p><strong>## Changelog</strong></p> <p><strong>2022-04-08 </strong>Added Relative Humidity (RH)<br> <strong>2022-03-07 </strong>Added the missing month in CDD/HDD&nbsp;<br> <strong>2022-02-08&nbsp;</strong>Updated the wind speed and temperature data due to missing months.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Enron Email Time-Series Network

<p>We use the&nbsp;<a href="https://www.kaggle.com/wcukierski/enron-email-dataset">Enron email dataset</a> to&nbsp;build a network of email addresses. It contains 614586 emails sent over the period from 6 January 1998 until 4 February 2004. During the pre-processing, we remove the periods of low activity and keep the emails from 1 January 1999 until 31 July 2002 which is 1448 days of email records in total. Also, we remove email addresses that sent less than three emails over that period. In total, the&nbsp;Enron email network contains 6 600 nodes and 50 897 edges.</p> <p>To build a graph <em>G = (V</em><em>, E</em><em>)</em>, we use email addresses as nodes <em>V</em>. Every node <em>v<sub>i</sub></em> has an attribute which is a time-varying signal that corresponds to the number of emails sent from this address during a day. We draw an edge <em>e</em><em><sub><em>ij</em></sub></em> between two nodes <em>i</em> and <em>j</em> if there is at least one email exchange between the corresponding addresses.</p> <p>Column <em>&#39;Count&#39;</em>&nbsp;in <em>&#39;edges.csv&#39;</em>&nbsp; file is the number of &#39;From&#39;-&gt;&#39;To&#39; email exchanges between the two&nbsp;addresses. This column can be used as an edge weight.</p> <p>The file <em>&#39;nodes.csv&#39;</em>&nbsp;contains a dictionary that is a compressed representation of time-series. The format of the dictionary is <em>Day-&gt;The Number Of Emails Sent By the Address During That Day.</em>&nbsp;The total number of days is 1448.</p> <p><em>&#39;id-email.csv&#39;</em>&nbsp;is a file containing the actual email addresses.</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Accelerometer-Based Multivariate Time-Series Dataset for Calf Behavior Classification

<p><strong>AcTBeCalf Dataset Description</strong></p> <p>The AcTBeCalf dataset is a comprehensive dataset designed to support the classification of pre-weaned calf behaviors from accelerometer data. It contains detailed accelerometer readings aligned with annotated behaviors, providing a valuable resource for research in multivariate time-series classification and animal behavior analysis. The dataset includes accelerometer data collected from 30 pre-weaned Holstein Friesian and Jersey calves, housed in group pens at the Teagasc Moorepark Research Farm, Ireland. Each calf was equipped with a 3D accelerometer sensor (AX3, Axivity Ltd, Newcastle, UK) sampling at 25 Hz and attached to a neck collar from one week of birth over 13 weeks.</p> <p>This dataset encompasses 27.4 hours of accelerometer data aligned with calf behaviors, including both prominent behaviors like lying, standing, and running, as well as less frequent behaviors such as grooming, social interaction, and abnormal behaviors.</p> <p>The dataset consists of a single CSV file with the following columns:</p> <ul> <li><strong>dateTime</strong>: Timestamp of the accelerometer reading, sampled at 25 Hz.</li> <li><strong>calfid</strong>: Identification number of the calf (1-30).</li> <li><strong>accX</strong>: Accelerometer reading for the X axis (top-bottom direction)*.</li> <li><strong>accY</strong>: Accelerometer reading for the Y axis (backward-forward direction)*.</li> <li><strong>accZ</strong>: Accelerometer reading for the Z axis (left-right direction)*.</li> <li><strong>behavior</strong>: Annotated behavior based on an ethogram of 23 behaviors.</li> <li><strong>segId</strong>: Segment identification number associated with each accelerometer reading/row, representing all readings of the same behavior segment.</li> </ul> <p>* the directions are mentioned in relation to the position of the accelerometer sensor on the calf.</p> <p><strong>Code Files Description</strong></p> <p>The dataset is accompanied by several code files to facilitate the preprocessing and analysis of the accelerometer data and to support the development and evaluation of machine learning models. The main code files included in the dataset repository are:</p> <ol> <li><strong>accelerometer_time_correction.ipynb</strong>: This script corrects the accelerometer time drift, ensuring the alignment of the accelerometer data with the reference time.</li> <li><strong>shake_pattern_detector.py</strong>: This script includes an algorithm to detect shake patterns in the accelerometer signal for aligning the accelerometer time series with reference times.</li> <li><strong>aligning_accelerometer_data_with_annotations.ipynb</strong>: This notebook aligns the accelerometer time series with the annotated behaviors based on timestamps.</li> <li><strong>manual_inspection_ts_validation.ipynb</strong>: This notebook provides a manual inspection process for ensuring the accurate alignment of the accelerometer data with the annotated behaviors.</li> <li><strong>additional_ts_generation.ipynb</strong>: This notebook generates additional time-series data from the original X, Y, and Z accelerometer readings, including Magnitude, ODBA (Overall Dynamic Body Acceleration), VeDBA (Vectorial Dynamic Body Acceleration), pitch, and roll.</li> <li><strong>genSplit.py:&nbsp;</strong>This script provides the logic used for the generalized subject separation for machine learning model training, validation and testing.</li> <li><strong>active_inactive_classification.ipynb</strong>: This notebook details the process of classifying behaviors into active and inactive categories using a RandomForest model, achieving a balanced accuracy of 92%.</li> <li><strong>four_behv_classification.ipynb</strong>: This notebook employs the mini-ROCKET feature derivation mechanism and a RidgeClassifierCV to classify behaviors into four categories: drinking milk, lying, running, and other, achieving a balanced accuracy of 84%.</li> </ol> <p>Kindly cite one of the following papers when using this data:</p> <p>Dissanayake, O., McPherson, S. E., Allyndr&eacute;e, J., Kennedy, E., Cunningham, P., &amp; Riaboff, L. (2024). <em>Evaluating ROCKET and Catch22 features for calf behaviour classification from accelerometer data using Machine Learning models</em>. arXiv preprint arXiv:2404.18159.</p> <p>Dissanayake, O., McPherson, S. E., Allyndr&eacute;e, J., Kennedy, E., Cunningham, P., &amp; Riaboff, L. (2024). <em>Development of a digital tool for monitoring the behaviour of pre-weaned calves using accelerometer neck-collars</em>. arXiv preprint arXiv:2406.17352</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Morphological phytoplankton counts for the SOMLIT-Astan time-series (2007-2017)

<p>The present text file includes morphologic microscopic counts&nbsp;for SOMLIT-Astan station (Western English Channel). Samples (250 mL) of natural seawater intended for the acquisition of microscopic counts were preserved with acid Lugol&rsquo;s iodine (Sournia, 1978, Guilloux et al. 2013), stored in the dark, and further processed between 15 days and up to 1 year after sampling.</p> <p>The morphological taxa contingency table was carefully examined to detect inconsistencies (e.g., abrupt changes in cell counts over the time series), and taxa for which identification was uncertain were grouped into broader taxonomic categories. For example, <em>Fragilaria</em> and <em>Brockmaniella</em> or <em>Cylindrotheca closterium</em> and <em>Nitzschia longissima</em> which are difficult to distinguished between each other, were considered in association in the same group of microscopic counts. The final morphological dataset consisted of counts of 146 taxonomical entities (taxa larger than 10&micro;m in size) across 185 dates from 2007 to 2017.</p> <p>Raw microscopic counts were regularly stored in a local MS-Access database and uploaded in the RESOMAR PELAGOS (<a href="http://abims.sb-roscoff.fr/pelagos/">http://abims.sb-roscoff.fr/pelagos/</a>) national database.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PSML: A Multi-scale Time-series Dataset for Machine Learning in Decarbonized Energy Grids (Dataset)

<p><strong>Abstract</strong></p> <p>The electric grid is a key enabling infrastructure for the ambitious transition towards carbon neutrality as we grapple with climate change. With deepening penetration of renewable energy resources and electrified transportation, the reliable and secure operation of the electric grid becomes increasingly challenging. In this paper, we present PSML, a first-of-its-kind open-access multi-scale time-series dataset, to aid in the development of data-driven machine learning (ML) based approaches towards reliable operation of future electric grids. The dataset is generated through a novel transmission + distribution (T+D) co-simulation designed to capture the increasingly important interactions and uncertainties of the grid dynamics, containing electric load, renewable generation, weather, voltage and current measurements at multiple spatio-temporal scales. Using PSML, we provide state-of-the-art ML baselines on three challenging use cases of critical importance to achieve: (i) early detection, accurate classification and localization of dynamic disturbance events; (ii) robust hierarchical forecasting of load and renewable energy with the presence of uncertainties and extreme events; and (iii) realistic synthetic generation of physical-law-constrained measurement time series. We envision that this dataset will enable advances for ML in dynamic systems, while simultaneously allowing ML researchers to contribute towards carbon-neutral electricity and mobility.&nbsp;</p> <p><strong>Data Navigation</strong></p> <p>Please download, unzip and put somewhere for later benchmark results reproduction and data loading and performance evaluation for proposed methods.</p> <pre><code>wget https://zenodo.org/record/5130612/files/PSML.zip?download=1 7z x 'PSML.zip?download=1' -o./ </code></pre> <p><strong>Minute-level Load and Renewable</strong></p> <ul> <li>File Name <ul> <li>ISO_zone_#.csv: `CAISO_zone_1.csv` contains minute-level load, renewable and weather data from 2018 to 2020 in the zone 1 of CAISO.</li> </ul> </li> <li>- Field Description <ul> <li>Field `<em>time</em>`: Time of minute resolution.</li> <li>Field `<em>load_power</em>`: Normalized load power.</li> <li>Field `<em>wind_power</em>`: Normalized wind turbine power.</li> <li>Field `<em>solar_power</em>`: Normalized solar PV power.</li> <li>Field `<em>DHI</em>`: Direct normal irradiance.</li> <li>Field `<em>DNI</em>`: Diffuse horizontal irradiance.</li> <li>Field `<em>GHI</em>`: Global horizontal irradiance.</li> <li>Field <em>`Dew Point</em>`: Dew point in degree Celsius.</li> <li>Field `<em>Solar Zeinth Angle</em>`: The angle between the sun&#39;s rays and the vertical direction in degree.</li> <li>Field `<em>Wind Speed</em>`: Wind speed (m/s).</li> <li>Field `<em>Relative Humidity</em>`: Relative humidity (%).</li> <li>Field `<em>Temperature</em>`: Temperature in degree Celsius.</li> </ul> </li> </ul> <p><strong>Minute-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>case #: The `case 0` folder contains all data of scenario setting #0. <ul> <li>pf_input_#.txt: Selected load, renewable and solar generation for the simulation.</li> <li>pf_result_#.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>Field <em>`time`</em>: Time of minute resolution.</li> <li>Field <em>`Vm_###`</em>: Voltage magnitude (p.u.) at the bus ### in the simulated model.</li> <li>Field <em>`Va_###`</em>: Voltage angle (rad) at the bus ### in the simulated model.</li> <li>Field <em>`P_#_#_#`</em>: `P_3_4_1` means the active power transferring in the #1 branch from the bus 3 to 4.</li> <li>Field <em>`Q_#_#_#`</em>: `Q_5_20_1` means the reactive power transferring in the #1 branch from the bus 5 to 20.</li> </ul> </li> </ul> <p><strong>Millisecond-level PMU Measurements</strong></p> <ul> <li>File Name <ul> <li>Forced Oscillation: The folder contains all forced oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li>&nbsp;info.csv: This file contains the start time, end time, location and type of the disturbance</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> <li>Natural Oscillation: The folder contains all natural oscillation cases. <ul> <li>row_#: The folder contains all data of the disturbance scenario #. <ul> <li>dist.csv: Three-phased voltage at nodes in the distribution system via T+D simualtion.</li> <li>info.csv: This file contains the start time, end time, location and type of the disturbance.</li> <li>trans.csv: Voltage at nodes and power on branches in the transmission system via T+D simualtion.</li> </ul> </li> </ul> </li> </ul> </li> <li>Filed Description <ul> <li>trans.csv <ul> <li>&nbsp; - Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li>&nbsp; - Field <em>`VOLT ###`</em>: Voltage magnitude (p.u.) at the bus ### in the transmission model.</li> <li>&nbsp; - Field <em>`POWR ### TO ### CKT #`</em>: `POWR 151 TO 152 CKT &#39;1 &#39;` means the active power transferring in the #1 branch from the bus 151 to 152.</li> <li>&nbsp; - Field <em>`VARS ### TO ### CKT #`</em>: `VARS 151 TO 152 CKT &#39;1 &#39;` means the reactive power transferring in the #1 branch from the bus 151 to 152.</li> </ul> </li> <li>dist.csv <ul> <li>Field <em>`Time(s)`</em>: Time of millisecond resolution.</li> <li>Field <em>`####.###.#`</em>: `3005.633.1` means per-unit voltage magnitude of the phase A at the bus 633 of the distribution grid, the one connecting to the bus 3005 in the transmission system.</li> </ul> </li> </ul> </li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Time-series of shoreline change for the Klamath River Littoral Cell (California)

<p>This repository contains 35&nbsp;years of tidally-corrected shoreline change data at the Klamath River Littoral Cell in northern California. This dataset was used&nbsp;in&nbsp;<em>Warrick et al. 2023, &quot;</em><strong>A Large Sediment Accretion Wave Along a Northern California Littoral Cell</strong>&quot;<em>,&nbsp;</em>to investigate and track the movement of a large sediment wave.</p> <p><em>CoastSat&nbsp;</em>was used to map shoreline changes on Landsat 5, Landsat 7 and Landsat 8 imagery between 1984 and 2022. The&nbsp;<em>Coastsat&nbsp;</em>toolbox is publicly available at&nbsp;https://github.com/kvos/CoastSat and described in&nbsp;<em>Vos et al. 2019,&nbsp;</em><a href="https://doi.org/10.1016/j.envsoft.2019.104528">https://doi.org/10.1016/j.envsoft.2019.104528</a>.&nbsp;The time-series of shoreline change were tidally-corrected along cross-shore transects using tide levels from a global tide model (FES2014) and a satellite-derived estimate of the beach slope (as described in&nbsp;<em>Vos et al. 2020, &quot;Beach slopes from satellite-derived shorelines&quot;,&nbsp;</em><a href="https://doi.org/10.1029/2020GL088365">https://doi.org/10.1029/2020GL088365</a><em>)</em>.</p> <p>The data is located in the <em>/shoreline_data</em> folder and structured as follows:</p> <ul> <li>The littoral cell is divided in 4 sections (kmt_01, kmt_02, kmt_03, kmt_04)</li> <li>For each section there is a&nbsp;folder with 4 CSV files: <ul> <li><em>time_series_tidally_corrected.csv</em>: this file contains the tidally-corrected time-series of shoreline change along each transect belonging to the site (e.g. kmt01-000, kmt01-001&nbsp;etc). This is the final product used for&nbsp;coastal change analyses.</li> <li><em>time_series_raw.csv</em>: this file contains the raw time-series of shoreline change, which have not be tidally-corrected. Note that each image is taken at a different stage of the tide.</li> <li><em>tide_levels_fes2014</em>: this file contains the tide levels at the time of image acquisition extracted from FES2014 (global tide model publicly available on AVISO+).</li> <li><em>transect_coordinates_and_beach_slopes.csv</em>: this file contains the coordinates (in WGS84 lat/lon coordinates) as well as the estimated beach slope for each transect.</li> </ul> </li> </ul> <p>In addition, there are 3 geospatial layers (.GEOJSON) which contain important spatial information. All the geospatial layers are in&nbsp;EPSG:2163 - US National Atlas Equal Area:</p> <ul> <li>&nbsp;<em>Klamath_polygons.geojson</em>: this layer contains the polygons that were used to run CoastSat for each section of the littoral cell.</li> <li><em>Klamath_shorelines.geojson</em>: this layer contains the sandy shorelines that were used to generate the cross-shore transects (also&nbsp;used as reference shorelines in CoastSat).</li> <li><em>transects.geojson</em>: this layer contains the cross-shore transects, which are spaced 100 m alongshore.</li> </ul> <p>Finally, in the<em> /animations</em> folder, there is a clip showing the mapped shorelines on the satellite imagery.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

EO4WildFires: An Earth Observation multi-sensor, time-series machine-learning-ready benchmark dataset for wildfire impact prediction

<p>This paper presents a benchmark dataset called EO4WildFires; a multi-sensor (multi spectral; Sentinel-2, Synthetic-Aperture Radar - SAR; Sentinel-1, meteorological parameters; NASA Power) time-series dataset that spans 45 countries, which can be used for developing machine learning and deep learning methods targeted for the estimation of the area that a forest wildfire might cover.</p> <p>This novel EO4WildFires dataset is annotated using EFFIS (European Forest Fire Information System) as forest fire detection and size estimation data source. A total of 31,742 wildfire events are gathered from 2018 to 2022. For each event, Sentinel-2 (multispectral), Sentinel-1 (SAR) and meteorological data are assembled into a single data cube. The meteorological parameters that are included in the data cube are: ratio of actual partial pressure of water vapor to the partial pressure at saturation, average temperature, bias corrected average total precipitation, average wind speed, fraction of land covered by snowfall, percent of root zone soil wetness, snow depth, snow precipitation, as well as percent of soil moisture.</p> <p>The main problem that this dataset is designed to address, is the severity forecasting before wildfires occur. The dataset is not used to predict wildfire events, but rather to predict the severity (size of area damaged by fire) of a wildfire event, if that happens in a specific place under the current and historical forest status, as recorded from multispectral and SAR images, and meteorological data.</p> <p>Using the data cube for the collected wildfire events, the EO4WildFires dataset is used to realize three (3) different preliminary experiments, in order to evaluate the contributing factors for wildfire severity prediction. The first experiment evaluates wildfire size using only the meteorological parameters, the second one utilizes both the multispectral and SAR parts of the dataset, while the third exploits all dataset parts. In each experiment, machine learning models are developed, and their accuracy is evaluated.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Pre-processed and modeled GNSS time-series after the 2011 Tohoku Earthquake

<p>The raw, pre-processed, and modeled&nbsp;GNSS time-series of the 213 GEONET sites in the Tohoku region, Japan, from Mar. 12, 2011 to Nov. 20, 2021, relative to the Okhotsk plate (Argus et al., 2011, <em><em>Geochemistry, Geophysics, Geosystems</em></em>).</p> <p>The original GNSS time-series are F5 solutions, which are distributed by&nbsp;Geospatial Information Authority of Japan (GSI,&nbsp;https://www.gsi.go.jp/). The details and availability of F5 solutions are written in Takamatsu et al. (2023, Earth, Planets, and Space)&nbsp;https://doi.org/10.1186/s40623-023-01787-7.</p> <p>The GNSS time-series processing was performed by Tomita (submitted), and the following signals were excluded from the raw time-series: seasonal variation, coseismic step, antenna maintenance offset, and common mode errors. Then, the pre-processed time-series were modeled by a trajectory modeling method considering&nbsp;postseismic deformation of the 2011 Tohoku earthquake, the Boso SSEs, and&nbsp;postseismic deformations due to aftershocks and L-ASE (long-term aseismic&nbsp;slip event) since late 2019.<br> <br> &quot;sitelist.txt&quot; - Site information file<br> column 1: Full site ID<br> column 2: 4digits site ID<br> column 3: Longitude [deg]<br> column 4: Latitude [deg]<br> column 5: Height [m]&nbsp;<br> <br> &quot;pre-process/xxxx.txt&quot; - Time-series at xxxx (4digits site ID) site<br> column 1: days from&nbsp;Mar. 12, 2011 (1 corresponds to Mar. 12, 2011)<br> column 2: raw East-West displacement [m]<br> column 3: raw North-South&nbsp;displacement [m]<br> column 4: raw Up-down&nbsp;displacement [m]<br> column 5: pre-processed&nbsp;East-West displacement [m]<br> column 6: pre-processed&nbsp;North-South&nbsp;displacement [m]<br> column 7: pre-processed&nbsp;Up-down&nbsp;displacement [m]</p> <p>&quot;model/xxxx/prediction_yy.txt&quot; - Time-series for yy&nbsp;component (yy=EW, NS, UD) at xxxx (4digits site ID) site<br> column 1: days from&nbsp;Mar. 12, 2011 (1 corresponds to Mar. 12, 2011)<br> column 2: modeled&nbsp;displacement excluding the Boso SSEs [m]<br> column 3: modeled&nbsp;displacement excluding the Boso SSEs and&nbsp;postseismic deformation due to aftershocks caused one year after the 2011 Tohoku Eq. [m]<br> column 4: modeled&nbsp;displacement excluding the Boso SSEs, postseismic deformation due to aftershocks caused one year after the 2011 Tohoku Eq. and the 2019 L-ASE&nbsp;[m]</p> <p><br> The displacement on Mar. 12, 2011 was initially set to be zero before the pre-processing, but the removal of the above factors provided some deviation from zero.</p> <p>The raw time-series excluded outliers from the original F5 solutions, and the raw time-series were transformed into the Okhotsk plate reference.</p> <p>Following the above trajectory modeling, the fully-relaxed postseismic displacement fields due to 2015 Feb. 17 Sanriku-oki earthquake (&quot;Table_displacement1.xlsx&quot;), the 2015&nbsp;May 13 Miyagi-oki earthquake (&quot;Table_displacement2.xlsx&quot;), and summation of the 2021 Feb. 13 Fukushima-oki, the 2021 Mar. 20 Miyagi-oki, and the 2021 May 1 earthquakes (&quot;Table_displacement2.xlsx&quot;) were calculated. Moreover, the cumulative displacement field due to the 2019 L-ASE since Nov. 25, 2019 was also calculated.&nbsp;In those files, the estimation errors are also shown as 1&sigma; standard deviation obtained from diagonal components of the model covariance matrices.&nbsp;</p> <p>&nbsp;</p> <p>The details of these data are introduced in the corresponding paper (Tomita, submitted).</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Ensemble Machine Learning Prediction of Potential FAPAR: Monthly time-series 2021 and Long-Term Comparison with Actual FAPAR

<p><strong>General Description</strong></p> <p>The dataset contains composites at 250 m spatial resolution of (1) &nbsp;monthly potential FAPAR for the year 2021 from ensemble ML model predictions, (2) the model deviance for each prediction, (3) the yearly average of potential FAPAR, (4) the yearly average of actual FAPAR and (5) the yearly average of the difference between actual and potential (actual minus potential) FAPAR. The dataset is based on the <a href="https://zenodo.org/record/8392976">95th percentile of the monthly aggregated FAPAR</a>&nbsp;derived from&nbsp;<a href="http://glass.umd.edu/Overview.html">250&thinsp;m 8&thinsp;d GLASS V6 FAPAR</a>. Potential FAPAR was predicted by fitting an ensemble ML model using globally distributed training points (cca 3 Mio) and a set of 52 biophysical covariates including several layers related to human pressure. The code for modeling potential FAPAR is openly available at <a href="http://github.com/Open-Earth-Monitor/Global_FAPAR_250m">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m</a>. The dataset can be used in many applications like land degradation modeling, land productivity mapping, and land potential mapping.&nbsp;</p> <p><strong>Data Details</strong></p> <ul> <li><strong>Time period:</strong> January 2021 - December 2021</li> <li><strong>Type of data: </strong>Fraction of Absorbed Photosynthetically Active Radiation (FAPAR)</li> <li><strong>How the data was collected or derived:</strong> Derived from 250m 8 d GLASS V6 FAPAR</li> <li><strong>Statistical methods used: </strong>Ensemble machine learning</li> <li><strong>Limitations or exclusions in the data: </strong>The dataset does not include data for Antarctica.</li> <li><strong>Coordinate reference system:</strong> EPSG:4326</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.0008094, 179.9999424, 87.37000)</li> <li><strong>Spatial resolution:</strong> 1/480 d.d. = 0.00208333 (250m)</li> <li><strong>Image size: </strong>172,800 x 71,698</li> <li><strong>File format: </strong>Cloud Optimized Geotiff (COG) format.</li> </ul> <p><strong>Support</strong></p> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: <a href="https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues">https://github.com/Open-Earth-Monitor/Global_FAPAR_250m/issues</a></p> <p><strong>Reference</strong></p> <p>Hackl&auml;nder, J., Parente, L., Ho, Y.-F., Hengl, T., Simoes, R., Consoli, D., Şahin, M., Tian, X., Herold, M., Jung, M., Duveiller, G., Weynants, M., Wheeler, I., (2023?) &quot;Land potential assessment and trend-analysis using 2000&ndash;2021 FAPAR monthly time-series at 250 m spatial resolution&quot;, submitted to PeerJ, preprint available at: <a href="https://doi.org/10.21203/rs.3.rs-3415685/v1">https://doi.org/10.21203/rs.3.rs-3415685/v1</a></p> <p>&nbsp;</p> <p><strong>Name convention</strong></p> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p> <ol> <li><strong>generic variable name:</strong> pot.fapar = Potential Fraction of Absorbed Photosynthetically Active Radiation</li> <li><strong>variable procedure combination: </strong>eml = ensemble machine learning</li> <li><strong>Position in the probability distribution / variable type:</strong> m = mean</li> <li><strong>Spatial support:</strong> 250m</li> <li><strong>Depth reference: </strong>s = surface</li> <li><strong>Time reference begin time:</strong> 20210101 = 2021-01-01</li> <li><strong>Time reference end time:</strong> 20211231 = 2021-12-31</li> <li><strong>Bounding box: </strong>go = global (without Antarctica)</li> <li><strong>EPSG code:</strong> epsg.4326 = EPSG:4326</li> <li><strong>Version code:</strong> v20230924 = 2023-09-24 (creation date)</li> </ol>

opencc-by-4.0Oct 2023View details →
edi44/100

SBC LTER: Ocean: Time-series: Mid-water SeaFET pH and CO2 system chemistry with surface and bottom Dissolved Oxygen at Arroyo Quemado Reef(ARQ), 2012-2017

Calibrated pH (Total scale, SeaFET sensor) an disoolved oxygen (miniDOT)data were collected from Arroyo Quemado Reef in the Santa Barbara Channel (site ID: ARQ). pH data are accompanied by in situ temperature and associated carbonate chemistry parameters. The SeaFET instrument is located about 4 meters from the surface, with other moored instruments. Associated carbonate chemistry parameters were calculated with the CO2calc programs from USGS, and include: partial pressure and fugosity of CO2, concentrations of bicarbonate, carbonate and hyrdroxide ion, Omega (saturation state) of calcite and aragonite. Dissolved oxygen sensors (miniDOT, PME) were added in 2014, and are mounted near the ocean surface and near the seafloor, and also report temperature. All data have been interpolated to a 20 minute time interval for compatibility with other SBC LTER moored instrument data. Data coverage is 2012-07-30 to 2017-03-10.All data from this site have been concatenated with the pH data from the other sites and merged into one data package: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-sbc&identifier=6005

openCC (other)Sep 2020View details →
edi44/100

SBC LTER: Ocean: Time-series: Mid-water SeaFET pH and CO2 system chemistry with surface and bottom Dissolved Oxygen at Mohawk Reef(MKO), 2012 - 2017

Calibrated pH (Total scale, SeaFET sensor) an disoolved oxygen (miniDOT)data were collected from Mohawk Reef in the Santa Barbara Channel (site ID: MKO). pH data are accompanied by in situ temperature and associated carbonate chemistry parameters. The SeaFET instrument is located about 4 meters from the surface, with other moored instruments. Associated carbonate chemistry parameters were calculated with the CO2calc programs from USGS, and include: partial pressure and fugosity of CO2, concentrations of bicarbonate, carbonate and hyrdroxide ion, Omega (saturation state) of calcite and aragonite. Dissolved oxygen sensors (miniDOT, PME) were added in 2014, and are mounted near the ocean surface and near the seafloor, and also report temperature. All data have been interpolated to a 20 minute time interval for compatibility with other SBC LTER moored instrument data. Data coverage is 2012-01-11 to 2017-12-19.The update of this dataset was terminated in 2019. All data from this site have been concatenated with the pH data from the other sites and merged into one data package: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-sbc&identifier=6005

openCC (other)Sep 2020View details →
edi44/100

SBC LTER: Ocean: Time-series: Mid-water SeaFET pH and CO2 system chemistry with surface and bottom Dissolved Oxygen at Santa Barbara Harbor/Stearns Wharf(SBH), 2012-2017

Calibrated pH (Total scale, SeaFET sensor) an disoolved oxygen (miniDOT)data were collected from Santa Barbara Harbor/Stearns Wharf in the Santa Barbara Channel (site ID: SBH). pH data are accompanied by in situ temperature and associated carbonate chemistry parameters. The SeaFET instrument is located about 4 meters from the surface, with other moored instruments. Associated carbonate chemistry parameters were calculated with the CO2calc programs from USGS, and include: partial pressure and fugosity of CO2, concentrations of bicarbonate, carbonate and hyrdroxide ion, Omega (saturation state) of calcite and aragonite. Dissolved oxygen sensors (miniDOT, PME) were added in 2014, and are mounted near the ocean surface and near the seafloor, and also report temperature. All data have been interpolated to a 20 minute time interval for compatibility with other SBC LTER moored instrument data. Data coverage is 2012-09-15 to 2016-09-14. The update of this dataset was terminated in 2019. All data from this site have been concatenated with the pH data from the other sites and merged into one data package: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-sbc&identifier=6005

openCC (other)Sep 2020View details →
zenodo40/100

Pre-eruption InSAR time-series at Kīlauea (Hawai`i, USA): COSMO-SkyMed Descending 2018

<p>InSAR time-series data for Kīlauea&nbsp;(Hawai`i, USA), between&nbsp;Jan 2010 and Sep 2011&nbsp;. Data were obtained by processing COSMO-SkyMed&nbsp;descending SAR data (track&nbsp;165). Data were processed using the JPL-developed InSAR Scientific Computing Environment (<code>ISCE</code>) open-source software package, and further time-series analysis was performed using the&nbsp;<code>MintPy</code>&nbsp;software toolbox (<a href="https://github.com/insarlab/MintPy">Miami INsar Time-series software in PYthon</a>), developed at the University of Miami.&nbsp;</p> <p>The following file&nbsp;is&nbsp;available in Hierarchical Data Format:</p> <p><code>geo_timeseries_tropHgt_demErr_cskDT165.h5</code>: Descending Track timeseries file.&nbsp;Dates available:</p> <p><code>[&#39;timeseries-20101001&#39;, &#39;timeseries-20101009&#39;, &#39;timeseries-20101017&#39;, &#39;timeseries-20101025&#39;, &#39;timeseries-20101102&#39;, &#39;timeseries-20101110&#39;, &#39;timeseries-20101118&#39;, &#39;timeseries-20101126&#39;, &#39;timeseries-20101204&#39;, &#39;timeseries-20101212&#39;, &#39;timeseries-20101220&#39;, &#39;timeseries-20110129&#39;, &#39;timeseries-20110206&#39;, &#39;timeseries-20110214&#39;, &#39;timeseries-20110222&#39;, &#39;timeseries-20110302&#39;, &#39;timeseries-20110303&#39;, &#39;timeseries-20110310&#39;, &#39;timeseries-20110318&#39;, &#39;timeseries-20110319&#39;, &#39;timeseries-20110322&#39;, &#39;timeseries-20110326&#39;, &#39;timeseries-20110403&#39;, &#39;timeseries-20110404&#39;, &#39;timeseries-20110407&#39;, &#39;timeseries-20110411&#39;, &#39;timeseries-20110419&#39;, &#39;timeseries-20110420&#39;, &#39;timeseries-20110423&#39;, &#39;timeseries-20110505&#39;, &#39;timeseries-20110506&#39;, &#39;timeseries-20110509&#39;, &#39;timeseries-20110513&#39;, &#39;timeseries-20110521&#39;, &#39;timeseries-20110522&#39;, &#39;timeseries-20110525&#39;, &#39;timeseries-20110529&#39;, &#39;timeseries-20110606&#39;, &#39;timeseries-20110607&#39;, &#39;timeseries-20110614&#39;, &#39;timeseries-20110622&#39;, &#39;timeseries-20110630&#39;, &#39;timeseries-20110708&#39;, &#39;timeseries-20110709&#39;, &#39;timeseries-20110716&#39;, &#39;timeseries-20110724&#39;, &#39;timeseries-20110725&#39;, &#39;timeseries-20110801&#39;, &#39;timeseries-20110809&#39;, &#39;timeseries-20110810&#39;, &#39;timeseries-20110817&#39;, &#39;timeseries-20110825&#39;, &#39;timeseries-20110826&#39;, &#39;timeseries-20110902&#39;, &#39;timeseries-20110910&#39;, &#39;timeseries-20110918&#39;]</code></p> <p>&nbsp;</p> <p>These data are supplemental to: Farquharson, J. I. and Amelung, F. [2020], &quot;<em>Extreme rainfall triggered the 2018 rift eruption at Kīlauea Volcano.</em>&quot;&nbsp;<a href="https://doi.org/10.1038/s41586-020-2172-5">https://doi.org/10.1038/s41586-020-2172-5</a></p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Association between meteorological factors and the number of tuberculosis notifications: a time-series study in Hong Kong

<p>&nbsp;Using a 22-year consecutive surveillance data in Hong Kong, including&nbsp; monthly averages of meteorological factors, air pollution concentrations , total number of TB cases notified,&nbsp;to analyze the association of monthly average temperature and relative humidity with temporal dynamics of monthly total number of TB cases notified.&nbsp;</p>

opencc-by-3.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record