Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,118
datasets available to search
ShareScore release 0.9.0
Dataset results
1,118 results for “Time series”
SBC LTER: Beach: Time-series of beach wrack cover and biomass, ongoing since 2008
These data describe the composition, cover, depth, and wet biomass of macrophyte wrack accumulated in the intertidal zone measured on five selected sandy beaches of the mainland coast of the Santa Barbara Channel. Data collection began in 2008 and this dataset is updated annually. Two data tables in this data package: 1. Wrack_Cover_All_Years. This is long-term time-series dataset starting in 2008 with fixed wrack species codes. 2. The Wrack_Cover_NON_FILLED_Partialkep dataset is added in 2019, to distinguish whether the fresh or old blade, stipe and holdfast was a fragment or attached to a whole plant.
SBC LTER: Beach: Time series of abundance of birds and stranded kelp on selected beaches, ongoing since 2008
Dataset contains the distribution, abundance and seasonal occurrence of birds, humans, dogs, and of freshly stranded giant kelp (Macrocystis pyrifera) plants and holdfasts present during monthly low-tide (<= 2.5 ft) weekday surveys of standard 1 km long transects on selected sandy beaches of the mainland coast of the Santa Barbara Channel. The abundance of humans, dogs, marine mammals and dead or oiled wildlife is also recorded during each survey. Sampling began in 2008 and is ongoing. See Methods for more information.
SBC LTER: Ocean: Time-series: Mid-water SeaFET and CO2 system chemistry at Alegria (ALE), 2011-2014
Calibrated pH (Total scale, SeaFET sensor) data was collected from Alegria in the Santa Barbara Channel (site ID: ALE) along with in situ temperature and associated carbonate chemistry parameters. The SeaFET instrument is located about 4 meters from the surface, with other moored instruments. Associated carbonate chemistry parameters were calculated with the CO2calc programs from USGS, and include: partial pressure and fugosity of CO2, concentrations of bicarbonate, carbonate and hyrdroxide ion, Omega (saturation state) of calcite and aragonite. Data have been interpolated to a 20 minute time interval for compatibility with other SBC LTER moored instrument data. Data coverage is 2011-06-21 to 2014-01-07.The update of this dataset was terminated in 2019. All data from this site have been concatenated with the pH data from the other sites and merged into one data package: https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-sbc&identifier=6005
SBC LTER: Ocean: Time-series: nearshore dissolved oxygen and temperature outside of reefs, ongoing since 2011
Dissolved oxygen (miniDOT sensors) data was collected from 6 reefs in the Santa Barbara Channel along with in situ temperature. Most dissolved oxygen sensors are deployed together with SBC long-term mooring instruments. Data collection intervals and miniDOT sensor depths vary based on the site location.
SBC LTER: pH time series: Water-sample pH and CO2 system chemistry, ongoing since 2011
Data are a time-series of pH and carbonate chemistry in manually-collected sea water samples from sixteen near-shore locations and 20 locations along the Santa Barbara Channel, California, USA, and intended to benchmark data from moored pH instruments. Data collection began in June 2011 and is ongoing. Time-series for pH sensors deployed in the field (e.g., SeaFET) are available elsewhere. Discrete seawater samples were collected by hand or using a Niskin bottle. In the laboratory, samples were analyzed for pH (total scale) and total alkalinity. Two parameters were used to calculate the full suite of carbonate chemistry parameters using the CO2calc algorithms. Program outputs are included for laboratory conditions as input, and adjusted for situ temperature. See Methods for more information, e.g., Lunden, et al. Data were collected by a consortium of research groups with the intent to leverage their collective assets and expertise. These groups (and sponsors) include Santa Barbara Coastal LTER (NSF), OMEGAS Project (NSF), Passow Lab (Calif. State Water Board), Hofmann Lab (NSF, CINMS, US NPS).
SBC LTER: Beach: Time-series of beach wrack consumers, ongoing since 2011
These data contain the composition, count, and wet biomass of macroinvertebrates along intertidal transects at five selected sandy beaches of the mainland coast of the Santa Barbara Channel. Data collection began in 2012 and is ongoing. This dataset is updated annually. Macroinvertebrates were collected from uniformly spaced cores along the shore-normal transects also used for physical measurements and macrophyte wrack sampling. Cores were pooled and animals and organic material sieved. In the laboratory, all macroinvertebrate species were identified, enumerated, blotted dry and weighed.
A Spatially Variable Time Series of Sea Level Change Due to Artificial Water Impoundment
<p>This database contains a series of gravitational, rotational, and deformational (GRD) "fingerprints"—the spatial response of sea level—corresponding to redistribution of water mass because of impoundment of water in artificial reservoirs, as reported in Hawley <em>et al</em>. (2020). Fingerprints for the GRanD database (Lehner <em>et al</em>.; 2011) are for individual years, noted in the file name.</p> <p>Three additional files come from the dataset provided by Zarfl <em>et al</em>. (2015), as described in Hawley <em>et al.</em> (2020). "Const" includes the fingerprint for all reservoirs under construction in their database; "Plan" includes the fingerprint for all reservoirs in the planning phase. "Zarfl" includes the fingerprint for all reservoirs in "Const," with 15 years of seepage, as well as all reservoirs for "Plan" with 5 years of seepage, as described in Hawley <em>et al</em>. (2020).</p> <p>Each fingerprint has 525,825 points, which fill out a global grid of 513 x 1025 [lat x lon] points. Each node in latitude and longitude is evenly spaced. The first point represents the northernmost point at 0 [deg] longitude, and increase first to the east, then to the south.</p>
Overview of the time series in the PALMOD 130k marine palaeoclimate data synthesis
<p>Palaeoclimate time series in the PALMOD 130k marine palaeoclimate data synthesis v1.0.1. This table lists the site names and location, parameters including additional information as well as the source of the data and the original publications where the data were presented.</p>
A 2-minute rainfall (12 locations) and discharge time series at the Vallon de Nant catchment, Switzerland, for 2018 summer seasons
<p>The data set contains rainfall time series within the experimental 13.4 km² Vallon de Nant catchment, Switzerland (Michelon et al., 2020), from June 30th to September 23rd 2018 at 12 locations. A network of <em>Pluvimate</em> drop-counting raingauges (www.driptych.com) measured continuously the rainfall intensity at a 2-minute resolution. Operation and characteristics of the raingauges are detailed in Benoit et al. (2018) and Michelon et al. (2020), and the rating curve is described by Ceperley et al. (2018).</p> <p>Description of the files:</p> <ul> <li><em><strong>data.csv</strong></em> contain the rainfall intensities for the observation period, along with the main river discharge measured at the <a href="https://map.geo.admin.ch/?lang=fr&topic=ech&bgLayer=ch.swisstopo.pixelkarte-farbe&layers=ch.swisstopo.zeitreihen,ch.bfs.gebaeude_wohnungs_register,ch.bav.haltestellen-oev,ch.swisstopo.swisstlm3d-wanderwege,KML%7C%7Chttps:%2F%2Fpublic.geo.admin.ch%2FaLKDanGXRPGMpB_D51f2Tg&layers_visibility=false,false,false,false,true&layers_timestamp=18641231,,,,&E=2574619.27&N=1122462.26&zoom=8">outlet</a> over the same 2-minutes time step as the rainfall intensity. We also provide areal rainfall intensity aggregated over the whole catchment:<br> Columns: <ul> <li>year [-]</li> <li>month [-]</li> <li>day [-]</li> <li>hour [-]</li> <li>minute [-]</li> <li>specific discharge 95% inf. [mm/day]: inferior values of the specific discharge (with 95% of confidence interval) over 2 minutes</li> <li>specific discharge 95% sup. [mm/day]: superior values of the specific discharge (with 95% of confidence interval) over 2 minutes</li> <li>specific discharge mean [mm/day]: mean value of the specific discharge over 2 minutes</li> <li>specific discharge median [mm/day]: median value of the specific discharge over 2 minutes</li> <li>P St. #X [mm]: rainfall amount measured at the station X over 2 minutes</li> <li>P stochastic mean [mm/h]: rainfall amount interpolated over the whole catchment over 2 minutes</li> <li>P stochastic std [mm/h]: standard deviation of the stochastic rainfall interpolation, over 2 minutes</li> </ul> </li> <li><strong><em>stations.csv</em></strong> describes the raingauge locations.<br> Columns: <ul> <li>Station ID [-]</li> <li>lon [WGS84]: decimal longitude of the station into WGS84</li> <li>lat [WGS84]: decimal latitude of the station into WGS84</li> <li>E [CH1903]: east coordinate into Swiss Coordinate System</li> <li>N [CH1903]: north coordinate into Swiss Coordinate System</li> <li>elevation [m asl]: altitude of the station in meters above the sea level</li> <li>data in 2017 [-]: flag if the station was working over the 2017 observation period</li> <li>data in 2018 [-]: flag if the station was working over the 2018 observation period</li> </ul> </li> <li><em><strong>rainfall_viewer.m</strong></em> is a <em>MatLab</em> script (created with <em>MatLab 2017b</em>) which allows the joint visualization of the rainfall intensities and river discharge. It produces a composite figure with the following plots: <ul> <li>On top the general hydrograph over the whole observation period [mm/day]. The red dashed lines mark out period that the other plots are focus on. The shaded orange areas correspond to when the river stage data was not available.</li> <li>Below, the zoomed hydrogram show a detailed view of the river discharge (and uncertainty). In case a river reaction is associated, the discharge event is marked out by red dashed lines. Between these vertical lines is drawn a line joining the initial and final baseflow, separating the discharge amount fed by the baseflow (under the line) to the fast runoff (over the line). The red square shows the center of mass of the fast runoff part.</li> <li>In the middle a zoomed magnification of the hydrograph that shows a detailed view of the discharge in the river [mm/day]. When a river response is associated, the discharge event is marked with dashed red lines. Between these vertical lines a line joining the initial and final baseflow is drawn, separating the discharge amount fed by the baseflow (under the line) to the fast runoff (over the line). The red square shows the center of mass of the fast runoff.</li> <li>At the bottom are shown the rainfall recorded by each of the 12 rain gauges (the y-axis scale between 2 stations is about 20 mm/h). The rainfall event is marked out by green dashed lines.</li> <li>Above is shown the rainfall amount (and uncertainty) interpolated over the catchment using the stochastic method. The rainfall event is marked out by green dashed lines.</li> <li>On the left, a map with the 12 raingauge locations show the total amount of rainfall recorded by each station during the event (a red cross shows missing data).<br> <br> It is possible to zoom in the plots by clicking with the left and right mouse buttons to define respectively the starting and ending of the visualization window. The middle button defines a third time reference used to identify rainfall intensity peaks or discharge peaks. Statistics concerning the visualization period are displayed on the MatLab console.<br> Pressing [enter] will save the figure into a PNG file named with the starting and ending dates of the visualization window.</li> </ul> </li> <li><strong><em>Q_stats.m </em></strong>is a MatLab function used by the main code rainfall_viewer.m</li> <li><strong><em>print_figure.m </em></strong>is a MatLab function used by the main code rainfall_viewer.m</li> <li><strong>data.mat</strong> is a MatLab data file with all data required by the main code rainfall_viewer.m</li> </ul>
Antarctic time series of temperature, precipitation, and stable isotopes in precipitation from the ECHAM5/MPI-OM-wiso past1000 climate model simulation
<p>This data set contains time series of two-metre air temperature (tas), surface temperature (ts), total precipitation (pr), oxygen-18 isotopic composition in precipitation (oxy), and deuterium isotopic composition in precipitation (dtr) from the past-millennium (800-1999 CE) simulation of the fully coupled ECHAM5/MPI-OM-wiso atmosphere-ocean general circulation model equipped with stable isotope diagnostics (Sjolte et al., 2018, Werner et al., 2016) used in the publication of Münch et al. (2021).</p> <p>The data here are provided for the Antarctic region, i.e., all model grid cells south of 60° S. The model's atmospheric component was run with a T31 spectral resolution (3.75° x 3.75°) and with 19 vertical levels, resulting in a total of N = 768 model grid cells covered by this data set. Note, however, that all time series off the continent of Antarctica have been set to NA values, so that the effectively available number of model grid cells is N<sub>eff</sub> = 442.</p> <p>Time series are provided at the original monthly resolution of the model output and on annual resolution obtained from the monthly resolution data. At annual resolution, the temperature and isotopic composition data are available as normal time averages and as precipitation-weighted time averages. In addition to the time series, the spatial field of time-invariant means is supplied, also as normal and precipitation-weighted time averages.</p> <p>Data are available as netcdf files and as R data files. In addition, processing code (bash and R scripts) are provided to reproduce the processing from monthly to annnual and time-invariant resolution and to read the data into the R data format. To process the R data, you will need the CRAN packages "ncdf4" and "lubridate", and the package "pfields" available on GitHub (see References).</p>
TIDMAD: Time Series Dataset for Discovering Dark Matter with AI Denoising
<p>TIDMAD is the first dataset and benchmark from a dark matter physics experiment, providing ultra-long time series data and comprehensive tools that enable machine learning models to directly advance the fundamental physics search for dark matter.</p> <p>This data is availble for download via <code>download_data.py</code>. Metadata for this dataset is specified in <code>TIDMAD_croissant.json</code>. The file names are listed in <code>filelist.dat</code>. For furhter information and publically available code, please see the associated <a href="https://github.com/jessicafry/TIDMAD" target="_blank" rel="noopener">GitHub repository</a>. For more information on this dataset and benchmark, please reference our TIDMAD paper.</p>
Monthly time series of rainfall, potential evapotranspiration and streamflow for 201 catchments in South-East Australia
<p>The data set contains data for 201 catchments located in South-Eastern Australia. The data was extracted from the datasets collated by Lerat, Thyer et al. (2020) including rainfall and potential-evapotranspiration data obtained from the Bureau of Meteorology Australian Water Outlook website (Frost, Ramchurn et al. 2016) and streamflow data obtained from the Bureau of Meteorology Water Data Online website (Bureau of Meteorology 2019). The data was collected over the period from 1980 to 2018, split into the two sub-periods 1980-1999 (Period 1) and 1999-2018 (Period 2).</p><p> </p><p>Bureau of Meteorology. (2019). "Water Data Online." from <a href="http://www.bom.gov.au/waterdata">http://www.bom.gov.au/waterdata</a>.</p><p>Frost, A. J., A. Ramchurn and A. Smith (2016). "The bureau's operational AWRA landscape (AWRA-L) Model." Bureau of Meteorology Technical Report.</p><p>Lerat, J., M. Thyer, D. McInerney, D. Kavetski, F. Woldemeskel, C. Pickett-Heaps, D. Shin and P. Feikema (2020). "A robust approach for calibrating a daily rainfall-runoff model to monthly streamflow data." Journal of Hydrology<strong>591</strong>: 125129.</p>
Labeled high-resolution orthoimagery time-series of an alluvial river corridor; Elwha River, Washington, USA.
<h2>Labeled high-resolution orthoimagery time-series of an alluvial river corridor; Elwha River, Washington, USA.</h2><h4>Daniel Buscombe, Marda Science LLC</h4><p>There are two datasets in this data release:</p><p>1. <strong>Model training dataset</strong>. A manually (or semi-manually) labeled image dataset that was used to train and evaluate a machine (deep) learning model designed to identify subaerial accumulations of large wood, alluvial sediment, water, and vegetation in orthoimagery of alluvial river corridors in forested catchments. </p><p>2. <strong>Model output dataset</strong>. A labeled image dataset that uses the aforementioned model to estimate subaerial accumulations of large wood, alluvial sediment, water, and vegetation in a larger orthoimagery dataset of alluvial river corridors in forested catchments. </p><p>All of these label data are derived from raw gridded data that originate from the U.S. Geological Survey (<i>Ritchie et al., 2018</i>). That dataset consists of 14 orthoimages of the Middle Reach (MR, in between the former Aldwell and Mills reservoirs) and 14 corresponding Lower Reach (LR, downstream of the former Mills reservoir) of the Elwha River, Washington, collected between the period 2012-04-07 and 2017-09-22. That orthoimagery was generated using SfM photogrammetry (following <i>Over et al., 2021</i>) using a photographic camera mounted to an aircraft wing. The imagery capture channel change as it evolved under a ~20 Mt sediment pulse initiated by the removal of the two dams. The two reaches are the ~8 km long Middle Reach (MR) and the lower-gradient ~7 km long Lower Reach (LR). </p><p>The orthoimagery have been labeled (pixelwise, either manually or by an automated process) according to the following classes (inter class in the label data in parentheses):</p><p>1. vegetation / other (0)</p><p>2. water (1)</p><p>3. sediment (2)</p><p>4. large wood (3)</p><h3>1. Model training dataset.</h3><p>Imagery was labeled using a combination of the open-source software Doodler (<i>Buscombe et al., 2021</i>; <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler</a>) and hand-digitization using QGIS at 1:300 scale, rasterizeing the polygons, and gridded and clipped in the same way as all other gridded data. Doodler facilitates relatively labor-free dense multiclass labeling of natural imagery, enabling relatively rapid training dataset creation. The final training dataset consists of 4382 images and corresponding labels, each 1024 x 1024 pixels and representing just over 5% of the total data set. The training data are sampled approximately equally in time and in space among both reaches. All training and validation samples purposefully included all four label classes, to avoid model training and evaluation problems associated with class imbalance (<i>Buscombe and Goldstein, 2022</i>). </p><p>Data are provided in geoTIFF format. The imagery and label grids (imagery) are reprojected to be co-located in the NAD83(2011) / UTM zone 10N projection, and to consist of 0.125 x 0.125m pixels.</p><p>Pixel-wise labels measurements such as these facilitate development and evaluation of image segmentation, image classification, object-based image-analysis (OBIA), and object-in-image detection models, and numerous potential other machine learning models for the general purposes of river corridor classification, description, enumeration, inventory, and process or state quantification. For example this dataset may serve in transfer learning contexts for application in different river or coastal environments or for different tasks or class ontologies.</p><h4>Files:</h4><p>1. Labels_used_for_model_training_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 63 MB, label tiffs</p><p>2. Model_<i>training_</i> images1of4.zip, 1.5 GB, imagery tiffs</p><p>3. Model_<i>training_</i> images2of4.zip, 1.5 GB, imagery tiffs</p><p>4. Model_<i>training_</i> images3of4.zip, 1.7 GB, imagery tiffs</p><p>5. Model_<i>training_</i> images4of4.zip, 1.6 GB, imagery tiffs</p><h3>2. Model output dataset.</h3><p>Imagery was labeled using a deep-learning based semantic segmentation model (<i>Buscombe, 2023</i>) trained specifically for the task using the Segmentation Gym (<i>Buscombe and Goldstein, 2022</i>) modeling suite. We use the software package Segmentation Gym (<i>Buscombe and Goldstein, 2022</i>) to fine-tune a Segformer (<i>Xie et al., 2021</i>) deep learning model for semantic image segmentation. We take the instance (i.e. model architecture and trained weights) of the model of <i>Xie et al. (2021)</i>, itself fine-tuned on ADE20k dataset (<i>Zhou et al., 2019</i>) at resolution 512x512 pixels, and fine-tune it on our 1024x1024 pixel training data consisting of 4-class label images.</p><p>The spatial extent of the imagery in the MR is [455157.2494695878122002,5316532.9804129302501678 : 457076.1244695878122002,5323771.7304129302501678] (NAD83(2011) / UTM zone 10N). Imagery width is 15351 pixels and imagery height is 57910 pixels. The spatial extent of the imagery in the LR is [457704.9227139975992031,5326631.3750646486878395 : 459241.6727139975992031,5333311.0000646486878395] (NAD83(2011) / UTM zone 10N). Imagery width is 12294 pixels and imagery height is 53437 pixels. Data are provided in Cloud-Optimzed geoTIFF (COG) format. The imagery and label grids (imagery) are reprojected to be co-located in the NAD83(2011) / UTM zone 10N projection, and to consist of 0.125 x 0.125m pixels. All grids have been clipped to the union of extents of active channel margins during the period of interest.</p><p>Reach-wide pixel-wise measurements such as these facilitate comparison of wood and sediment storage at any scale or location. These data may be useful for studying the morphodynamics of wood-sediment interactions in other geomorphically complex channels, wood storage in channels, the role of wood in ecosystems and conservation or restoration efforts. </p><h4>Files:</h4><p>1. Elwha_MR_labels_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 9.67 MB, label COGs from Elwha River Middle Reach (MR)</p><p>2. Elwha<i>MR_ imagery_ part1_ of</i>_<i> </i>2.zip, 566 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>3. Elwha<i>MR_ imagery_ part2_ of</i>_<i> </i>2.zip, 618 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>3. Elwha_LR_labels_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 10.96 MB, label COGs from Elwha River Lower Reach (LR)</p><p>4. ElwhaL<i>R_ imagery_ part1_ of</i>_<i> </i>2.zip, 622 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>5. ElwhaL<i>R_ imagery_ part2_ of</i>_<i> </i>2.zip, 617 MB, imagery COGs from Elwha River Middle Reach (MR)<br> </p><p>This dataset was created using open-source tools of the Doodleverse, a software ecosystem for geoscientific image segmentation, by Daniel Buscombe (<a href="https://github.com/dbuscombe-usgs">https://github.com/dbuscombe-usgs</a>) and Evan Goldstein (<a href="https://github.com/ebgoldstein">https://github.com/ebgoldstein</a>). Thanks to the contributors of the Doodleverse!. Thanks especially Sharon Fitzpatrick (<a href="https://github.com/2320sharon">https://github.com/2320sharon</a>) and Jaycee Favela for contributing labels. </p><h3>References</h3><p>• Buscombe, D. (2023). <strong>Doodleverse/Segmentation Gym SegFormer models for 4-class (other, water, sediment, wood) segmentation of RGB aerial orthomosaic imagery (v1.0)</strong> [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8172858">https://doi.org/10.5281/zenodo.8172858</a></p><p>• Buscombe, D., Goldstein, E. B., Sherwood, C. R., Bodine, C., Brown, J. A., Favela, J., et al. (2021).<strong> Human-in-the-loop segmentation of Earth surface imagery</strong>. Earth and Space Science, 9, e2021EA002085. <a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a></p><p>• Buscombe, D., & Goldstein, E. B. (2022). <strong>A reproducible and reusable pipeline for segmentation of geoscientific imagery.</strong> Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p><p>• Over, J.R., Ritchie, A.C., Kranenburg, C.J., Brown, J.A., Buscombe, D., Noble, T., Sherwood, C.R., Warrick, J.A., and Wernette, P.A., 2021, <strong>Processing coastal imagery with Agisoft Metashape Professional Edition, version 1.6—Structure from motion workflow documentation</strong>: U.S. Geological Survey Open-File Report 2021–1039, 46 p., <a href="https://doi.org/10.3133/ofr20211039">https://doi.org/10.3133/ofr20211039</a>.</p><p>• Ritchie, A.C., Curran, C.A., Magirl, C.S., Bountry, J.A., Hilldale, R.C., Randle, T.J., and Duda, J.J., 2018, <strong>Data in support of 5-year sediment budget and morphodynamic analysis of Elwha River following dam removals</strong>: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/F7PG1QWC">https://doi.org/10.5066/F7PG1QWC</a>.</p><p>• Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M. and Luo, P., 2021. <strong>SegFormer: Simple and efficient design for semantic segmentation with transformers</strong>. Advances in Neural Information Processing Systems, 34, pp.12077-12090.</p><p>• Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A. and Torralba, A., 2019. <strong>Semantic understanding of scenes through the ade20k dataset</strong>. International Journal of Computer Vision, 127, pp.302-321.</p><p><br> </p>
Energy System Time Series Suite (ESTSS) - Data Archive
<h2>Energy System Time Series Suite - Data Archive</h2> <p> </p> <p>This archive contains variously sized sets of declustered time series within the context of energy systems. These series demonstrate low discrepancy and high heterogeneity in feature space, resulting in a roughly uniform distribution within this space.</p> <p>For detailed information, please refer to the corresponding GitHub project:<br><a href="https://github.com/s-guenther/estss/">https://github.com/s-guenther/estss/</a></p> <p>For associated research, see<br><a href="https://doi.org/10.1186/s42162-024-00304-8">https://doi.org/10.1186/s42162-024-00304-8</a></p> <p>Data is provided in .csv format. The GitHub project includes a Python function to load this data as a dictionary of pandas data frames.</p> <p>Should you utilize this data, kindly also cite the associated research paper. For any queries, please feel free to reach out to us through GitHub or the contact details provided at the end of this readme file.</p> <p> </p> <h3>Folder Content</h3> <ul> <li>`ts_*.csv`: Contains declustered load profile time series in tabular format. <ul> <li>Size: `(n+1) x (m+1)`, with `n` representing time steps (1000 per series) and `m` the number of series.</li> <li>Includes a header row and index column. Headers indicate series id, and the index column numbers each time step, starting from `0`.</li> <li>The first half of the series `(m/2)` consistently display a constant sign (negative). They are sequentially numbered from 0.</li> <li>The second half `(m/2)` display varying signs. Numbering starts from `1,000,000`.</li> </ul> </li> <li>`features_*.csv`: Tabulates features corresponding to the time series. <ul> <li>Size: `(m+1) x (f+1)`, where `m` is the number of time series and `f` is the number of features</li> <li>Includes a header row and index column. Indexes represent time series id (matching `ts_*.csv` headers), and headers name the features.</li> </ul> </li> <li>`norm_space_*.csv`: Shows feature vectors in normalized feature space where time series are declustered. Provided for completeness; typically not needed by users. <ul> <li>Size: `(m+1) x (g+1)`, where `m` is the number of timer series and `g` is the number of selected features space features. (a subset of `f` from `features_*.csv`).</li> <li>Format matches `features_*.csv`.</li> </ul> </li> <li>`info_*.csv`: Maps declustered datasets to the manifolded dataset. Provided for completeness; typically not needed by users. <ul> <li>Size: `(m+1) x 2`, with `m` as series count. Columns contain manifolded set time series ids.</li> <li>Includes an index column and a header. The index holds the remapped id of declustered series. Header `0` is non-significant.</li> </ul> </li> </ul> <p>Each `ts_*.csv`, `features_*.csv`, `norm_space_*.csv`, and `info_*.csv` file comes in four versions to accommodate various set sizes:</p> <ul> <li>`*_4096.csv`</li> <li>`*_1024.csv`</li> <li>`*_256.csv`</li> <li>`*_64.csv`</li> </ul> <p>These represent sets with 4096, 1024, 256, and 64 time series, respectively,offering different densities in feature space population. The objective is to balance computational load and resolution for individual research needs.</p> <p> </p> <h3>Contact</h3> <p>ESTSS - Energy System Time Series Suite<br>Copyright (C) 2023<br>Sebastian Günther<br>sebastian.guenther@ifes.uni-hannover.de</p> <p>Leibniz Universität Hannover<br>Institut für Elektrische Energiesysteme<br>Fachgebiet für Elektrische Energiespeichersysteme</p> <p>Leibniz University Hannover<br>Institute of Electric Power Systems<br>Electric Energy Storage Systems Section</p> <p><a href="https://www.ifes.uni-hannover.de/ees.html">https://www.ifes.uni-hannover.de/ees.html</a></p>
Sentinel-5P Methane Density at 2 km from 2021-12 to 2023-11 Monthly Aggregation Time-series Reconstructed
<p><strong>General Description</strong></p><p>The <i>monthly aggregated Methane Volume Mixing Ratio </i>dataset is derived from Sentinel-5P to generate a time-series reconstructed monthly aggregated map. The dataset time spans from December 2021 to November 2023 and provides data that covers the entire globe. The mission is still underway and expected to update periodically.</p><p>For more info about the s5p Methane product see: <a href="">https://maps.s5p-pal.com/ch4/</a>.</p><p>The dataset can be used in many applications like emission tracing, livestock monitor, and greenhouse gas monitor.</p><ul><li><strong>Monthly time-series:</strong></li></ul><p>Methane monthly average value December 2021 – November 2023. Derived using the <a href="https://eumap.readthedocs.io/en/latest/">eumap</a> and <a href="https://github.com/openlandmap/scikit-map">scikitmap</a> package in Python . We derived three standard statistics: (1) 10th percentile (p10), median (p50), and 90th percentile (p90).</p><p><strong>Data Details</strong></p><ul><li><strong>Time period:</strong> December 2021 – November 2023</li><li><strong>Type of data:</strong> Methane Volume Mixing Ratio (Unit: ppbv)</li><li><strong>How the data was collected or derived:</strong> Derived from 2km Sentinel-5P Menthane using Python running in a local HPC. The time-series analysis were computed using the <a href="https://github.com/scikit-map/scikit-map">Scikit-map</a> and <a href="https://eumap.readthedocs.io/en/latest/">eumap </a>Python package.</li><li><strong>Statistical methods used:</strong> percentiles 10, 50, and 90.</li><li><strong>Limitations or exclusions in the data:</strong> The dataset is not completed gap-filled. Certain areas have no data in the whole time series</li><li><strong>Coordinate reference system:</strong> EPSG:4326</li><li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -61.9966697, 180.0000072, 87.37000)</li><li><strong>Spatial resolution:</strong> 1/60 d.d. = 0.016666667 (2km)</li><li><strong>Image size:</strong> 21,600 x 8,962</li><li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li></ul><p><strong>Support</strong></p><p>If you discover a bug, artifact or inconsistency, or if you have a question please use some of the following channels:</p><ul><li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/-/issues">https://gitlab.com/openlandmap/global-layers/-/issues</a></li><li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li></ul><p><strong>Name convention</strong></p><p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p><ol><li><strong>generic variable name:</strong> ch4.vmr = methane density methane volume mixing ratio</li><li><strong>variable procedure combination:</strong> m.seacov = monthly aggregated and gap filled by seasonal convolution</li><li><strong>Position in the probability distribution / variable type:</strong> p10/p50/p90 = 10th/50th/90th percentile</li><li><strong>Spatial support:</strong> 2km</li><li><strong>Depth reference:</strong> a = above surface</li><li><strong>Time reference begin time:</strong> 20211201 = 2021-12-01</li><li><strong>Time reference end time:</strong> 20231131 = 2023-11-31</li><li><strong>Bounding box:</strong> go = global (without Antarctica)</li><li><strong>EPSG code:</strong> epsg.4326 = EPSG:4326</li><li><strong>Version code:</strong> v20230628 = 2023-12-08 (creation date)</li></ol>
Dataset of pomegranate tree (Punica granatum L. 'Wonderful') image times series
<p>Dataset of pomegranate tree (Punica granatum L. ‘Wonderful’) image times series. The pictures were collected by means of Raspberry Pi cameras with OV5647 sensor (5 MP, f2.9). Sensors were installed on fixed platforms for continuous measurement with zenithal orientation at a distance of approximately 1 metre from the canopy. Images were captured daily at 9 a.m. (GMT+2) from July to mid-October in 2021 and 2022. The resolution of the images is 640x480 pixels.</p>
Videos of the processed microscope images and time series of the petrophysical parameters from image processing and geochemical simulation and of the measured induced polarisation [Video][Dataset]
<p>Supporting Information for the manuscript <em>Microfluidics and spectral induced polarization for direct observation and petrophysical modeling of calcite dissolution</em> published in Geophysical Research Letters</p> <ul> <li><strong>Data Set S1.</strong> Porosity, water saturation, and calcite sample perimeter from image<br>processing.</li> <li><strong>Data Set S2.</strong> Porosity, water conductivity, and pH from geochemical simulation.</li> <li><strong>Data Set S3.</strong> Real and imaginary components of the complex electrical conductivity at<br>2.5 Hz and CEC from petrophysical modeling.</li> <li><strong>Movie S1.</strong> Dissolution of the calcite sample with the detected contour superimposed in<br>white on the grayscale images. Time, length scale, and flow direction are indicated. In<br>case of problems launching the file, we recommend using VLC Media Player software.</li> <li><strong>Movie S2.</strong> Segmented images of the CO2 bubbles produced by the calcite dissolution.<br>Time, length scale, and flow direction are indicated. In case of problems launching the<br>file, we recommend using VLC Media Player software.</li> </ul>
Synthetic time series data generation for edge analytics
<p>In this research, we create synthetic data with features that are like data from IoT devices. We use an existing air quality dataset that includes temperature and gas sensor measurements. This real-time dataset includes component values for the Air Quality Index (AQI) and ppm concentrations for various polluting gas concentrations. We build a JavaScript Object Notation (JSON) model to capture the distribution of variables and structure of this real dataset to generate the synthetic data. Based on the synthetic dataset and original dataset, we create a comparative predictive model. Analysis of synthetic dataset predictive model shows that it can be successfully used for edge analytics purposes, replacing real-world datasets. There is no significant difference between the real-world dataset compared the synthetic dataset. The generated synthetic data requires no modification to suit the edge computing requirements. The framework can generate correct synthetic datasets based on JSON schema attributes. The accuracy, precision, and recall values for the real and synthetic datasets indicate that the logistic regression model is capable of successfully classifying data</p>
DeepOrchidSeries: A Sentinel-2 Dataset to inform convolutional SDMs with twelve-month Sentinel-2 image time-series, Orchid family
<p><strong>Deep Species Distribution Modelling from Sentinel-2 Image Time-series: a Global Scale Analysis on the Orchid Family</strong> </p> <ul> <li><strong><em>DeepOrchidSeries</em></strong> dataset gathers Sentinel-2 image time-series around geolocated orchid occurrences. Seasonal evolutions of the habitats are captured in the twelve-month RGB/IR time-series with 640x640m spatial resolution. It allows novel Species Distribution Models (SDMs) coupled with convolutional networks to take advantage of both spatial and temporal information.</li> <li>Our <strong>associated article</strong> is describing the modeling choices made to shape this ambitious dataset. It is submitted to <a href="https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence">https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence</a>. We believe such global data, methods and scripts are valuable to the conservation ecology community and especially deep-SDMs users. To our knowledge, no similar ready-to-use dataset is available. In the article, the dataset's temporal dimension is proven to significantly improve SDMs performances.</li> <li><strong><em>sen2patch</em></strong> is the gitlab project gathering the code to create such dataset. It is available at <a href="https://gitlab.inria.fr/jestopin/sen2patch">https://gitlab.inria.fr/jestopin/sen2patch</a>.</li> <li><strong><em>DeepOrchidSeries.csv</em></strong> contains all occurrences-level information. <ul> <li>We advice to load it with: <pre><code class="language-python">import pandas as pd df = pd.read_csv("path/to/DeepOrchidSeries.csv", sep=';') df.columns ['gbifid', 'canonical_name', 'decimallatitude', 'decimallongitude', 'speciesKey', 'cell_index', 'bot_country', 'bot_code', 'lvl2_code', 'continent_code']</code></pre> <ul> <li>'gbifid' is the occurrences GBIF ID</li> <li>'canonical_name', is the species canonical name</li> <li>'decimallatitude', 'decimallongitude' are the species coordinates in decimal degrees</li> <li>'speciesKey' is the species GBIF unique identifier</li> <li>'cell_index' is a unique cell ID in a 0.0025° lon/lat grid partitioning the Earth (used to stratify train/val/test set by geographic blocks)</li> <li>'bot_country', 'bot_code', 'lvl2_code', 'continent_code' are geographic subdivisions defined in <a href="https://github.com/tdwg/wgsrpd">https://github.com/tdwg/wgsrpd</a> (code and string for WGSRPD level 1, the botanical countries)</li> </ul> </li> </ul> </li> <li> <p>Initial <a href="https://www.gbif.org/">GBIF</a> query DOI is <a href="http://https://doi.org/10.15468/dl.4bijtu">https://doi.org/10.15468/dl.4bijtu</a> (26 August 2019).</p> </li> <li><strong><em>DeepOrchidSeries.tar</em></strong> file contains the satellite image time-series and is available at <a href="https://lab.plantnet.org/deeporchidseries/">https://lab.plantnet.org/deeporchidseries/</a> <ul> <li><em>.tar</em> archive measure 286 GB and extends to 432 GB once decompressed.</li> <li>Image time-series relative tree paths are constructed from the occurrences unique GBIF IDs.</li> <li>For a given occurence <em>gbifid</em>, matching patches are located in: <em>final_dataset_by_gbifid/gbifid[-2:]/gbifid[-4:-2]</em>, <em>i.e.</em> in a first folder named with the <em>gbifid</em> last two numbers and a subfolder with the previous two ones. Example: the time-series files matching occurrence 2236837714 are located at <em>final_dataset_by_gbifid/14/77/</em>. </li> <li>Image time-series are composed of twelve 16 bits RGB <em>.png</em> and twelve 16 bits IR <em>.png</em> files containing data identical to the original L1C products, no lossy compression was made. There are one RGB and one IR .png file per month.</li> <li>Patches from month MM/YYYY of occurrence <em>gbifid</em> are named<em> </em><em>RGB_YYYY_MM_gbifid_.png</em> and <em>IR0_YYYY_MM_gbifid_.png</em>.</li> </ul> </li> <li><em><strong>models.zip</strong></em> is the archive containing the four PyTorch models weights described in our article and<strong><em> </em></strong><em><strong>inception_env.py</strong></em> the used Inception V3 architecture. <em><strong>index.json</strong></em> contains the dictionnary linking the models class indexes from 0 to 14128 with our labels <em>speciesKey</em>: {"class_index":speciesKey}.</li> </ul> <p> </p> <ul> <li><strong>ACKNOWLEDGMENTS</strong>: We warmly thank Alexander Zizka et al. for providing us the geographically and taxonomically curated set of Orchids occurrences. This dataset contains modified Copernicus Sentinel data and Copernicus Service information (2018). Sentinel-2 MSI data used were available at no cost from ESA Sentinels Scientific Data Hub.</li> </ul>
Global GFED-based monthly burned area time series (1996-2016) at 1 km and ESA CCI MODIS-based long-term monthly P90 burned area occurrence at 500 m
<p>Contains two separate datasets:</p> <ol> <li>Global <a href="https://www.globalfiredata.org/data.html">GFED-based monthly burned area</a> (in ha) <a href="https://youtu.be/kBJcP8mL2Qs">time series (1996-2016)</a> at 1 km (downscaled using cubic-splines from 25 km);</li> <li>Global burned area long term (2000-2012) P90 (quantile probability = 0.9) based on the <a href="http://maps.elie.ucl.ac.be/CCI/viewer/index.php">ESA CCI burned area accumulated weekly product</a>;</li> </ol> <p>Original GFED monthly data is provided as HDF4 files (ftp.fuoco.geog.umd.edu/data/GFED/GFED4). Dataset is described in detail in <a href="https://doi.org/10.1002/jgrg.20042">Giglio et al. (2013)</a>. Processing steps are available <a href="https://gitlab.com/openlandmap/global-layers/tree/master/input_layers/GFED"><strong>here</strong></a>. Antarctica is not included.</p> <p>To access and visualize global datasets use: <a href="https://openlandmap.org"><strong>https://openlandmap.org</strong></a> or watch <a href="https://youtu.be/kBJcP8mL2Qs"><strong>this video</strong></a>.</p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> </ul> <p>All files provided as Cloud-Optimized GeoTIFFs / internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>nhz = theme: natural hazards,</li> <li>monthly.burned.ha = variable: estimated monthly burned area in ha,</li> <li>gfed = data source GFED data,</li> <li>m = mean value,</li> <li>1km = spatial resolution / block support: 1 km,</li> <li>s0..0cm = vertical reference: land surface,</li> <li>2000.02 = time reference aggregated: month Feb of year 2000,</li> <li>v4 = version number: GFEDv4,</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.