Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
403
datasets available to search
ShareScore release 0.7.1
Dataset results
403 results for “satellite data”
Data from: Satellite-based Lagrangian model reveals how upwelling and oceanic circulation shape krill hotspots in the California Current System [updated]
<p><strong>Abstract</strong></p> <p>In the California Current System, wind-driven nutrient supply and primary production, computed from satellite data, provide a synoptic view of how phytoplankton production is coupled to upwelling. In contrast, linking upwelling to zooplankton populations is difficult due to relatively scarce observations and the inherent patchiness of zooplankton. While phytoplankton respond quickly to environmental forcing, zooplankton grow slower and tend to aggregate into mesoscale “hotspot” regions spatially decoupled from upwelling centers. To better understand mechanisms controlling the formation of zooplankton hotspots, we use a satellite-based Lagrangian method where variables from a plankton model, forced by wind-driven nutrient supply, are advected by near-surface currents following upwelling events. Modeled zooplankton distribution reproduces published accounts of euphausiid (krill) hotspots, including the location of major hotspots and their interannual variability. This satellite-based modeling tool is used to analyze the variability and drivers of krill hotspots in the California Current System, and to investigate how water masses of different origin and history converge to form predictable biological hotspots. The Lagrangian framework suggests that two conditions are necessary for a hotspot to form: a convergence of coastal water masses, and above average nutrient supply where these water masses originated from. The results highlight the role of upwelling, oceanic circulation, and plankton temporal dynamics in shaping krill mesoscale distribution, seasonal northward propagation, and interannual variability.</p> <p><strong>Data set description</strong></p> <p>This data set includes 2 files:</p> <ul> <li>a satellite-based 1993-2023 monthly retrospective of krill concentrations (Zbig) modeled using the growth-advection method in the California Current upwelling system. Inputs include the nitrate supply product described below and GlobCurrent 15 m oceanic currents. This dataset is updated monthly (using NRT data) at https://www.mbari.org/science/upper-ocean-systems/biological-oceanography/krill-hotspots-in-the-california-current/.</li> <li>a satellite-based 1993-2023 monthly retrospective of wind-driven nitrate supply estimated in a 150 km coastal band at 0.125° latitudinal resolution. Nitrate supply was calculated based primarily on CCMP v3.1 winds, AVISO geostrophic currents, and a climatology of in situ nitrate at 60m. This dataset is updated monthly (using NRT data) at https://www.mbari.org/science/upper-ocean-systems/biological-oceanography/nitrate-supply-estimates-in-upwelling-systems/.</li> </ul> <p>See details regarding data sources and calculations in <a href="https://doi.org/10.3389/fmars.2022.835813">Messié et al. (2022)</a>.</p> <p>[IMPORTANT NOTE:] There is an error in the Ekman pumping fields (trans_pump, Nsupply_pump, Nsupply_total) that will be corrected soon (those fields are not used in publications where only coastal transport was considered). Please contact me if you need Ekman pumping fields before this is fixed.</p>
Seasonal sea ice indices including the timing of ice-edge advance and ice-edge retreat (in year day), the ice season duration (in days) and number of actual ice days (versus open water days) within the ice season, extracted for various PAL LTER sub-regions West of the Antarctic Peninsula and derived from passive microwave satellite data for 1979/80 to 2023/24 ice seasons.
Seasonal sea ice indices including the timing of ice-edge advance and ice-edge retreat (in year day), the ice season duration (in days) and number of actual ice days (versus open water days) within the ice season, extracted for various PAL LTER sub-regions West of the Antarctic Peninsula and derived from passive microwave satellite data for 1979/80 to 2023/24 ice seasons. The ice season duration is defined as the time elapsed between day of ice-edge advance and day of ice-edge retreat within a given sea ice year, which begins mid-February (mean minimum of summer sea ice extent for the Southern Ocean) and ends the following mid-February. See Stammerjohn et al (2008, JGR) for further details.
Average monthly sea ice coverage for various PAL LTER sub-regions West of the Antarctic Peninsula derived from passive microwave satellite data, 1978 - June 2024.
Monthly sea ice coverage derived from passive microwave satellite measurements and extracted for various PAL LTER subregions. Several different sea-ice metrics are provided including (1) monthly sea-ice extent, sea-ice area, and open-water area (km^2) extracted for the greater WAP region (from the AP to 80W) and for 3 PAL LTER grid regions: the original PAL grid (000-900 lines), the PAL 'DSR' grid (200-600 lines), and the PAL 'new' grid (-200 to 600 lines); and (2) monthly sea-ice concentration (%) extracted for small PAL subregions, including the nominal penguin foraging areas (~200km by ~200km) southwest of King George Island (KGI), Anvers, Avian and Charcot islands, as well as for Marguerite Bay (~140km by ~140km) (inland of the area defined for Avian).
Average yearly sea ice coverage for various PAL LTER sub-regions West of the Antarctic Peninsula derived from passive microwave satellite data, 1979 - 2023.
Annual sea ice coverage derived from passive microwave satellite measurements and extracted for various PAL LTER subregions. Several different sea-ice metrics are provided including: monthly sea-ice extent, sea-ice area, and open-water area (km^2) extracted for the greater WAP region (from the AP to 80W) and for 3 PAL LTER grid regions: the original PAL grid (000-900 lines), the PAL 'DSR' grid (200-600 lines), and the PAL 'new' grid (-200 to 600 lines).
PEATCLSM(Tb): A land surface data assimilation product for peatlands using PEATCLSM and brightness temperature (Tb) satellite observations (Northern Hemisphere output)
<p>The datasets archived here include simulation results shown in the paper, “Improved Groundwater Table and L-band Brightness Temperature Estimates for Northern Hemisphere Peatlands Using New Model Physics and SMOS Observations in a Global Data Assimilation Framework”, published in Remote Sensing of Environment Journal (Bechtold et al., 2020). The output was produced by combining peatland-specific land surface modeling (Bechtold et al., 2019b) embedded in the NASA Catchment Land Surface Model (CLSM) with L-band brightness temperature (Tb) observations (SMOS), applying the data assimilation framework of the SMAP Level‐4 Soil Moisture product (Reichle et al., 2019). We provide netcdf files (9-km resolution EASEv2 grid, period Jan 2010 – Nov 2019, and between 45°N and 70°N, NE Asia excluded) of the four experiments of the manuscript: model-only (open-loop, OL) and data assimilation (DA) for each land model version, that is CLSM without and with the use of the PEATCLSM modules. The highest accuracy is provided by the DA product using PEATCLSM and Tb observations. When referring to the latter product use the name ‘PEATCLSM(Tb)’. We provide three types of netcdf files:<br> • daily_images_*.nc: Daily land states and fluxes (Table 1), provided as netCDF image-chunked image stack<br> • ObsFcstAna_images_*.nc: Brightness temperature observations, forecasts and analysis (Table 2), provided as netCDF image-chunked image stack<br> • incr_timeseries_*.nc: Data assimilation increments (Table 3), provided as netCDF timeseries-chunked image stack</p> <p>The file content is described in the file PEATCLSM_Tb_Documentation_20200505.pdf</p> <p>Please contact Michel Bechtold (michel.bechtold@kuleuven.be) for any questions.</p> <p>Data usage statement:<br> This work is licensed under a Creative Commons Attribution 4.0 International License: https://creativecommons.org/licenses/by/4.0/<br> If you decide to work with this data, we kindly ask to be informed at the outset of the nature of this work. If the data are essential to the work, or if an important result or conclusion depends on the PEATCLSM(Tb) data product, we would appreciate that you discuss these findings with us to ensure correct use and interpretation of the PEATCLSM(Tb) product. Furthermore, we are continuously improving the data assimilation product, a discussion of your work at an early stage may (i) help us to improve our product, and (ii) allow us to provide you with a newer version. Thanks!</p> <p>References:</p> <p>Bechtold, M., De Lannoy, G. J. M., Reichle, R. H., & Koster, R. D. (2019a). PEAT-CLSM simulation output (Northern Peatlands) version 1. https://doi.org/10.17605/OSF.IO/E58YM</p> <p>Bechtold, M. et al. (2019b). PEAT‐CLSM: A Specific Treatment of Peatland Hydrology in the NASA Catchment Land Surface Model. <em>Journal of Advances in Modeling Earth Systems</em>, <em>11</em>(7), 2130–2162. https://doi.org/10.1029/2018MS001574</p> <p>Bechtold, M., De Lannoy, G. J. M., Reichle, R. H., Roose, D., Balliston, N., Burdun, I., Devito, K., Kurbatova, J., Strack, M., & Zarov, E. A. (2020). Improved Groundwater Table and L-band Brightness Temperature Estimates for Northern Hemisphere Peatlands Using New Model Physics and SMOS Observations in a Global Data Assimilation Framework. <em>Remote Sensing of Environment</em>. https://doi.org/10.1016/j.rse.2020.111805</p> <p>Reichle, R. H., Liu, Q., Koster, R. D., Crow, W. T., De Lannoy, G. J. M., Kimball, J. S., Ardizzone, J. V., Bosch, D., Colliander, A., Cosh, M., Kolassa, J., Mahanama, S. P., Prueger, J., Starks, P., & Walker, J. P. (2019). Version 4 of the SMAP Level-4 Soil Moisture Algorithm and Data Product. <em>Journal of Advances in Modeling Earth Systems</em>, <em>11</em>(10), 3106–3130. https://doi.org/10.1029/2019MS001729</p>
Enrichment index related to seamounts and islands in the South West Indian Ocean from chlorophyll-a satellite remote sensing data
<p>This data set is the result of the calculation of an original “enrichment index” (EI) from chlorophyll-a (chl-a) remote sensing data (MODIS-Aqua sensor) and initially dedicated to highlight localized chl-a enrichments associated to isolated seamounts and islands in the South West Indian Ocean, in order to estimate their contribution in increasing the local primary productivity. Details and results are described in the DSR-II paper entitled “Satellite observations of phytoplankton enrichments around seamounts in the South West Indian Ocean with a special focus on the Walters Shoal” from Demarcq et al. 2020.<br> 1. Initial data used<br> We used daily L3 data chl-a and sea surface temperature (SST) collected by the MODIS (Moderate-resolution Imaging Spectroradiometer) sensor on board the Aqua platform (downloaded from https://oceancolor.gsfc.nasa.gov/) from January 2003 to December 2018. This has a spatial resolution of 1/24° (ca. 4.5–5 km). The data covers the region (45°S – 10°S / 25°W – 80°W).<br> 2. The calculation method<br> The calculations were done at the pixel level. The EI is the difference (expressed in %) between the value of each ‘candidate pixel’ and its medium range surrounding, defined as the average value of all chl-a values around the candidate pixel between a fix range of distance between 30 and 90 km, the R1 and R2 terms of the equation enclosed.<br> 3. Data sets<br> The data set contains two files:<br> - the monthly climatology (12 frames) of the EI from January to December (2003 to 2018 average), in an internally compressed netCDF-4 format (NC-compliant or almost)<br> - the yearly average of the EI (period 01/2003 - 12/2018)<br> <br> Two images are joined with this data set:<br> - a "technical view" of the yearly average of the index for the full region sub-region (45°S – 10°S / 25°W – 80°W)<br> (file: indsw4_modis_p100_4km_16y_20030101_20181231.R2018.0.enrichment-index.dist-30-90km.png).</p> <p> - a slightly improved view of the yearly average of the index for the sub-region (40°S – 10°S / 30°W – 70°W).<br> (file: Figure-enrichment-index.pdf)<br> <br> An improved version of this index will be available in a near future.</p>
Rating curves based on satellite altimetry and in-situ discharge data
<h1>Context: </h1> <p>The ESA river discharge Climate Change Initiative (CCI) project is a precursor study. It aims to derive long term climate data records (at least over 20-years) of river discharge for some selected river basins (and some locations in the river network) using satellite remote sensing observations (altimetry and multispectral images) and ancillary data. It aims to provide a proof-of-concept for the feasibility for a potential River Discharge ECV product to meet the requirements for the <a href="https://gcos.wmo.int/en/essential-climate-variables/rivers/" target="_blank" rel="noopener">Global Climate Observing System</a>. This project covers precursor activities towards the production of data products that address the GCOS-defined requirements for the River Discharge ECV.</p> <h1>Data description :</h1> <p>Just as in-situ stage measurements can be used to gauge river discharge, altimetry-derived water surface elevation (WSE) can serve as an alternative means of estimating river discharge when discharge time series data is available. Several methodologies have been documented for deriving discharge time series from multimission altimetry observations and supplementary data (Biancamaria et al., 2024). At least two approaches will be used, depending on the available in situ discharge and altimetry water surface elevation (WSE) time series:</p> <p>⋅ <strong><em>Method 1</em>: </strong>The preferred approach relies on the altimetry water surface elevation time series and in situ discharge time series to create a rating curve (RC) characterized by a power relationship between these two variables following a Bayesian approach (Rantz et al., 1982). However, this method necessitates a significant overlap period between discharge data and radar altimetry measurements (e.g., Biancamaria et al., 2011; Papa et al., 2012), or it requires the assumption that the rating curve remains valid and consistent when discharge data is only available prior to the altimetry observation period.</p> <p>⋅ <em><strong>Method 2:</strong></em> The final option, in cases where there is no temporal overlap between in-situ or simulated discharge and water surface elevation data, assumes that the validity and stability of the rating curve persist across the various time periods covered by the two datasets. Both of these time periods should be sufficiently long to encompass a wide range of events. With this assumption, Tourian et al. (2013, 2017) introduced a method for calculating the rating curve, not based on the time series of discharge and water surface elevation, but on the distribution of their quantiles. This method has been adopted by a limited number of recent studies (e.g., Belloni et al., 2021). However, it’s important to note that this methodology naturally introduces higher errors when compared to the preferred approach. For this reason, this methodology will be validated over some stations with various hydrological dynamics and satisfying previous methods (overlap period exists between WSE and Q).</p> <h1>Approaches to derive Rating Curve (RC) :</h1> <h2>Bayesian Approach :</h2> <p>The Bayesian method is a robust statistical approach used for constructing a rating curve, frequently applied in the field of hydrology when the goal is to estimate unknown parameters from observed data, while taking into consideration the associated uncertainty in these estimates. </p> <p>According to this, the estimation of the rating curve using the Bayesian method involves several steps:</p> <ul> <li>The initial step entails defining a probabilistic model that describes the relationship between observed data and the parameters we aim to estimate. In many hydrological applications, the relationship between discharge data (Q) and water surface elevation data (WSE) is often expressed as a power function:</li> </ul> <p><em> Q = a⋅(WSE-z</em><em>0</em><em>)</em><sup><em>b</em></sup></p> <p>Here, <em>a, z0</em> and <em>b</em> are the parameters of the rating curve. <em>a,</em> is a scaling coefficient governing the magnitude of the Q-WSE relationship, <em>b,</em> characterizes the nature of this relationship, and <em>z0</em>, represents the height of the free surface above the reference point, corresponding to the river bottom's altitude. The power relationship is especially pertinent due to its consistency with numerous hydrodynamic phenomena. The exponent b within the equation allows for the representation of distinctive flow characteristics, including factors like roughness and channel geometry. Moreover, it offers adaptability in modelling to accommodate variations in flow characteristics, whether they are turbulent or laminar. This relationship, despite its mathematical simplicity, facilitates the fine-tuning of model adjustments in accordance with observed data (Chow, 1959).</p> <ul> <li>The second step involves the use of prior normal distributions, reflecting our prior knowledge about these parameters. These distributions can either be informative or uninformative, depending on our level of knowledge. The limits and ranges for a, z0 and b can vary depending on the specific context of the study, the dataset used, and the characteristics of the river or channel being analysed.</li> </ul> <p><u>- Coefficient “a”</u>: adjustment parameter for the rating curve representing the scaling factor for discharge. Its value can significantly fluctuate based on various factors such as the characteristics of the river or channel, hydraulic conditions, and other influencing factors. Consequently, "a" must be non-negative and constrained within a sensible range specific to the system under study. Following the Manning equation, “a” must be equal to W/n*S<sup>1/2</sup> (Chow et al., 1988) where W is the river’s width (m), n the Manning’s roughness coefficient and S the slope (m/m). Given the considerable variability in river width and slope across different stations, a feasible range for this coefficient can be considered as:</p> <p> a ∈ [0; 3000]</p> <p><u>- Coefficient “b”</u>: adjustment parameter representing the exponent of the rating curve and indicating the hydraulic condition of the study site. Like "a," this value must comply with physical constraints and cannot be negative. Following the Manning equation, “b” must be equal to 5/3 for reference hydraulic condition (Rantz et al., 1982). To accommodate the variability in system characteristics across sites, the following range values can be considered for this coefficient:</p> <p> b ∈ [0; 5]</p> <p><u>- Coefficient “z0”</u>: offset or the elevation at which discharge begins. It should be within the range of elevations relevant to your study. For this <em>reason, the value</em> cannot exceed the minimum value of water surface elevation (WSE) and the range value need to consider of the variability in term of water depth over the sites. A feasible range for this coefficient can be considered as:</p> <p> z0 ∈ [min(WSE)-30; min(WSE)]</p> <ul> <li>The final step involves parameter estimation. The posterior distribution of the parameters yields probabilistic estimates of the rating curve parameters in the form of mean values (optimal values) and credibility intervals (95th percentiles). This accounts for the uncertainty associated with these parameters and is achieved through Markov Chain Monte Carlo (MCMC) sampling from the posterior distribution. Two commonly employed MCMC algorithms are "NUTS" (No-U-Turn Sampler) and "Metropolis-Hastings." The Metropolis-Hasting sampler "MH" algorithm, which is relatively simple and efficient where a balance between exploration and exploitation is desired. This algorithm can be adapted to sample from discrete state spaces.</li> </ul> <h2>Quantile approach : </h2> <p>The Quantile approach employs statistical modelling using quantile functions to create a rating curve, eliminating the necessity for overlapping measurements. This algorithmic method enables the estimation of river discharge using satellite altimetry, even in instances where there are no in situ measurements within the altimeter's timeframe. This approach has undergone application and validation in diverse river basins spanning different climatic zones, such as the Amazon, Brahmaputra, Danube, Niger, and Ob (Tourian et al., 2013).</p> <p>Assuming a stationary flow behaviour and no modification in the river bathymetry both at the altimetry virtual station and at the in-situ gage, this approach ensures the utilization of historical in situ data in current applications. This method computes the quantile functions of the altimetry water surface elevation on one hand and of the discharge time series on the other hand. Then a scatter plot of these in-situ discharge quantiles versus altimetry water surface elevation quantiles is computed to establish the rating curve using the bayesian approach described previously.</p> <h1>File description :</h1> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>basin-station</td> <td>Basin name in capital letters and Station name in capital letters separated by "_" and where spaces have been replaced by "-".</td> </tr> <tr> <td>lon</td> <td>Longitude in decimal degrees [-180,180] with 4 decimals - corresponding to the insitu discharge station.</td> </tr> <tr> <td>lat</td> <td>Latitude in decimal degrees [-90,90] with 4 decimals – corresponding to the insitu discharge station.</td> </tr> <tr> <td>a</td> <td>Adjustment parameter for the rating curve representing the scaling factor for discharge. Number with 3 decimals.</td> </tr> <tr> <td>b</td> <td>Adjustment parameter representing the exponent of the RC and indicating the hydraulic condition of the study site. Number with 3 decimals.</td> </tr> <tr> <td>z0</td> <td>Offset of the elevation at which discharge begins. Number with 3 decimals.</td> </tr> <tr> <td>a_sd</td> <td>Standard deviation of the coefficient "a". Number with 3 decimals.</td> </tr> <tr> <td>b_sd</td> <td>Standard deviation of the coefficient "b". Number with 3 decimals.</td> </tr> <tr> <td>z0_sd</td> <td>Standard deviation of the coefficient "z0". Number with 3 decimals.</td> </tr> <tr> <td>period</td> <td>Period used to compute the rating curve under the format %Y-%m-%d where the start and the end dates are separated by ":"</td> </tr> <tr> <td>nb</td> <td>Number of overlap dates to compute the rating curve.</td> </tr> <tr> <td>Methodology</td> <td>Methodology used to compute the rating curve. The first part describes the approach used to compute the RC and the second part, separated by “_”, describes the algorithm used. To avoid any issue for the reader the spaces have been replaced by “-”. At the end 2 approaches has been used: “Overlap-approach” or “Quantile-approach” and 2 algorithms: “Bayesian-algorithm” or “Multiple-algorithms” designed for Arctic rivers experiencing frozen periods. </td> </tr> <tr> <td>Source</td> <td>In-situ data sources to compute the rating curve. If multiple sources has been used, the sources are separate by "/"</td> </tr> </tbody> </table> <p>---------</p> <p><em>THE DATASET IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR </em><em>IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,</em><br><em>FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE </em><em>AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER </em><em>LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, </em><em>OUT OF OR IN CONNECTION WITH THE DATASET OR THE USE OR OTHER DEALINGS IN THE </em><em>DATASET.</em></p>
Training data for bathymetry estimation via EO satellite - Hel Peninsula
<p>This dataset contains satellite image from Sentiel-2A (bands B2, B3, B4, B8) and reference sonar based bathymetry measurements.</p> <p>The reference data was acquired from Polish Maritime Administration (htttp://www.um.gdy.pl) and is publicly available. File reference_data_34.csv contains in-situ measrements aquired at northen shore of Hel Peninsula. Points coordinates are expressed in UTM34N coordinate system.</p> <p>The satellite and the reference datasets were preprocessed by the authors for adjust them for Machine Learning algorithms used in the research.</p> <p> </p>
A Novel Framework to Harmonise Satellite Data Series for Climate Applications: Matchups, Calibration Parameters and Residuals
<p>The datasets included with this archive supplement the journal article:</p> <p>Giering, R.; Quast, R.; Mittaz, J.P.D.; Hunt, S.E.; Harris, P.M.; Woolliams, E.R.; Merchant, C.J. A Novel Framework to Harmonise Satellite Data Series for Climate Applications. <em>Remote Sens. 2019</em>, <strong>11</strong>, 1002. doi:<a href="https://doi.org/10.3390/rs11091002">10.3390/rs11091002</a>.</p> <p>The archive includes a README file with further explanations.</p>
Maximum Independent Set Satellite Scheduling World Cities Data Set
<h1>Satellite Scheduling World Cities Data Set</h1> <p>The Satellite Scheduling World Cities Data Set is the a set of cities treated as point locations used to simulate a set of image collection tasking requests for AIAA paper "A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations". It provides an open reference and benchmark for the satellite task scheduling problem. This could also be considered as a sparse Maximum Independent Set problem for a generic graph. The requests represent point collects, from which we can compute multiple distinct collection opportunities. The tasking problem is then to select a subset of these collects that it is possible for the spacecraft to feasibly collect in a given time period, subject to constraints on the spacecraft's agility and constraints on only collecting a single collect per request (no duplication of effort).<br><br>The data set is hosted on both <a href="https://github.com/duncaneddy/aiaa-mis-satellite-scheduling-dataset">Github</a> and <a href="../">Zenodo</a>. The Github repository contains the original source data, the associated requests generated from the source data, and scripts to reproduce the scenario files. Zenodo (DOI 10.5281/zenodo) hosts copies of the output Metis graph files and collect data files. Due to the large size of produced files these are not included in the Github repository.</p> <h2>Notes</h2> <p><strong>Notes</strong><br><br>Please note that while the source data and generation methods are identical to the satellite task planning paper it was created for. The specific generated problems do not exactly reproduce the scenario in the paper. Since the original reproduction, updates in upstream software dependencies have changed the output of the generation process (specifically, Earth orientaiton parameter handling libraries). This can be determined by considering the cardinality of the generated collect set. However, these differences are generally small and since the constriant rate is similar, the results should be comparable.</p> <table> <tbody> <tr> <td>Spacecraft Count</td> <td>Orignial Publication Collect Count</td> <td>Reproduction Collect Count</td> </tr> <tr> <td>4</td> <td>59356</td> <td>59624</td> </tr> <tr> <td>6</td> <td>90777</td> <td>91204</td> </tr> <tr> <td>12</td> <td>180008</td> <td>180939</td> </tr> <tr> <td>24</td> <td>359170</td> <td>361519</td> </tr> </tbody> </table> <p><br>This repository also adds additional scenarios for 1, 2, and 36 satellites. Note, the provided scenarios represent the largest 10,000 request data set. Should a smaller request set be desired, the requests should be filtered to the top `x` request based on city population and any collects not associated with those requests should be discarded.</p> <p>Note the Zenodo repository excludes the collect and graph files for the 1 and 2 satellite scenarios to avoid the file limits. These can still be reproduced from the Github source code.</p> <h2>Acknolwedgement</h2> <p>If this data set is used in your research, please cite the following paper</p> <p><a href="https://arc.aiaa.org/doi/abs/10.2514/1.A34931">A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations</a></p> <blockquote> <pre><code>@article{eddy2021maximum, title={A Maximum Independent Set Method for Scheduling Earth-Observing Satellite Constellations}, author={Eddy, Duncan and Kochenderfer, Mykel J}, journal={Journal of Spacecraft and Rockets}, volume={58}, number={5}, pages={1416--1429}, year={2021}, publisher={American Institute of Aeronautics and Astronautics} }</code></pre> </blockquote> <h2>Licensing</h2> <p>The source of the world cities data is from the <a href="https://simplemaps.com/data/world-cities">simplemaps.com</a> website,<br>licensed under the <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License </a>with the specific license found at `./data/worldcities_license.txt`.</p>
WINTERC-G: a global upper mantle thermochemical model from coupled geophysical–petrological inversion of seismic waveforms, heat flow, surface elevation and gravity satellite data
<p>WINTERC-G: A global, temperature and compositional model of the lithosphere<br> and upper mantle.<br> Version: v5.4, December 2020, J. Fullea, S. Lebedev, Z. Martinec, N. Celli<br> <br> Contact: Javier Fullea (jfullea@ucm.es)<br> Facultad de Fisica,<br> Universidad Complutense de Madrid (UCM),<br> Spain<br> ////////<br> Geophysics Section,<br> Dublin Institute for Advanced Studies<br> Dublin, Ireland<br> </p> <p>TYPE:<br> This contains files with:<br> i) the model directly on the triangular grid solved for in the surface wave inversion.</p> <p> ii) an interpolated grid at 0.5 deg lateral resolution for the density and density discontinuities used in the gravity field data inversion<br> </p> <p>If you have any questions regarding the methodology or the construction<br> of the model, please contact the authors. If you use the model, we would<br> request that you cite the reference indicated below, and appreciate<br> your feedback regarding the model and its application.</p> <p>Citation:</p> <p>Fullea, J., Lebedev, S., Martinec, Z., & Celli, N. L. (2021). WINTERC-G: mapping the upper mantle thermochemical heterogeneity from coupled geophysical–petrological inversion of seismic waveforms, heat flow, surface elevation and gravity satellite data. Geophysical Journal International, 226(1), 146-191.</p> <p>*******************************<br> Summary: construction of the model.<br> WINTERC-G is a Waveform tomography and Gravity (geoid and gravity anomalies and gradiometric measurements<br> from ESA's GOCE mission) INversion model of the TEmpeRature and Composition of the lithosphere and upper mantle at<br> global scale. WINTERC-G is based on upon the integrated geophysical-petrological<br> approach LitMod (Afonso et al., 2008; Fullea et al. 2009) and, hence, all<br> relevant mantle rock physical properties modelled (seismic velocities and density) are<br> computed within a thermodynamically self-consistent framework allowing for a direct<br> parameterization in terms of the temperature and composition of the lithosphere-upper<br> mantle. The inversion is a two-step procedure. In a first step, we invert surface-wave, Rayleigh and Love<br> fundamental mode dispersion curves from a high resolution global dataset measured using waveform inversion,<br> along with surface heat flow and elevation (isostasy) for temperature and crustal structure<br> using a point-wise, non-linear, gradient-search inversion<br> over a triangular grid with an average 225 km lateral inter-knot spacing. In a second step we<br> use a fully parallelized spherical harmonic formalism to invert satellite gravity field data in<br> order to refine the initial crustal density and mantle composition distributions from the step 1<br> for a fixed temperature field.</p> <p>The parameter space in step 1 includes crust (densities and S-wave velocities for a three-layered crust)<br> and mantle variables (the depth of the thermal Lithosphere-Athenosphere-Boundary,<br> the thickness of the sublithospheric thermal buffer, the sublithospheric temperatures at 3 different<br> equispaced nodes down to 400 km, the lithospheric and sublithospheric mantle compositon, and<br> the the radial anisotropy at the 3 crustal layers and at 56, 80, 110, 150, 200, 260, 330,<br> and 400 km depths.</p> <p>The parameter space in step 2 is defined by the average crustal density, and the<br> mantle composition in the lithosphere and sublithosphere.<br> We use the output crustal density from step 1 as the<br> initial value in step 2 inversion. Mantle densities are derived based on the output temperature<br> field from step 1 (kept fixed) and the bulk mantle composition inversion variables.</p> <p> </p> <p> </p> <p>*******************************</p> <p>This archive contains the following files:<br> README (this file)<br> WINTERC-G_Vp-Vs.lis (triangular grid)<br> WINTERC-G_rad_anis_Vs.lis (triangular grid)<br> WINTERC-G_Temperature.lis (triangular grid)<br> WINTERC-G_Density.lis (triangular grid)<br> WINTERC-G_LAB.lis (triangular grid)<br> WINTERC_T_rho_1D.z (1D average model of temperature and density)<br> rho_*_out.xyz (0.5 deg egular grid for gravity field)<br> ETOPO2_km_continental.xyz (0.5 deg egular grid for gravity field)<br> ETOPO2_km_depth_Ice.xyz (0.5 deg egular grid for gravity field)<br> ETOPO2_km_depth_Bed.xyz (0.5 deg egular grid for gravity field)<br> Global_Moho_WINTERC-G.xyz (0.5 deg egular grid for gravity field)</p> <p><br> Files in the triangular grid with an average 225 km lateral inter-knot spacing (12232 grid points):</p> <p>* WINTERC-G_Vp-Vs.lis: Vp and Vs (in km/s) in all model columns with a vertical grid step of 2 km<br> Format for each column:<br> #Column number longitude latitude depth(km, <0 downwards) Vp (km/s) Vs(km/s)<br> 5640 93.72 4.135 -5.0 3.91 2.11</p> <p><br> * WINTERC-G_rad_anis_Vs.lis: radial anisotropy, (Vsh-Vsv)/Vs_iso (in %) in all model columns with a vertical grid step of 2 km<br> Format for each column:<br> #Column number longitude latitude depth(km, <0 downwards) anisotropy (%)</p> <p>* WINTERC-G_Temperature.lis: temperature (in ºC) in all model columns with a vertical grid step of 2 km<br> Format for each column:<br> #Column number longitude latitude depth (km, <0 downwards) T (ºC) dT (%) dT(K) <br> 6437 297.20 -2.524 -259.000 1431.9 -1.91 -27.9<br> The anomalies dT are in % and K with respect to the 1D model in WINTERC_T_rho_1D.z (column 2).</p> <p>* WINTERC-G_Density.lis: density (in kg/m3) in all model columns with a vertical grid step of 2 km<br> Format for each column:<br> #Column number longitude latitude depth(km, <0 downwards) rho (kg/m3) drho(%) drho(kg/m3)<br> The anomalies drho are in % and kg/m3 with respect to the 1D model in WINTERC_T_rho_1D.z (column 3).</p> <p>* WINTERC_T_rho_1D.z: 1D average model of temperature (column 2 in ºC) and density (column 3 in kg/m3) with a vertical grid step of 2 km <br> 5.00000000 0.0000000000000000 6.0259973839110526<br> 3.00000000 0.0000000000000000 38.960571309690394<br> 1.00000000 0.33634006819423840 174.42296045978722<br> -1.00000000 3.8888495253719624 1692.8437489147236<br> -3.00000000 23.974111923225379 1863.8834351235944<br> -5.00000000 47.727920701943034 2568.2414495590924<br> -7.00000000 89.633398074381162 2819.8386016341910<br> -9.00000000 137.01489361657013 2839.5325893195904<br> -11.0000000 182.35233447017222 2897.6600872935287<br> -13.0000000 224.46247069572485 2945.2036923862997<br> -15.0000000 260.63395547331390 3069.6809340323475<br> -17.0000000 292.28175449521456 3132.4574175461721<br> -19.0000000 322.29571965406632 3145.5747337463940<br> -21.0000000 351.58698283375054 3157.2401512748038<br> -23.0000000 380.30002225705056 3177.0000420059773<br> -25.0000000 408.50259805632055 3183.6651032398490<br> -27.0000000 436.22632217636487 3190.9586785996116<br> -29.0000000 463.48733903170023 3198.9369509456310<br> -31.0000000 490.29841705549831 3209.7229872383764<br> -33.0000000 516.71149258457456 3221.6329506091679<br> ...</p> <p>Files in the interpolated regular grid at 0.5 deg lateral resolution used for gravity field data inversion:</p> <p> * rho_c_out.xyz: average crustal density<br> * rho_submoho_out.xyz: mantle density below the Moho discontinuity<br> * rho_*_out.xyz: mantle density defined at different model depths: 20, 35, 56, 80, 110, 150, 200, 260, 330 and 400 km.</p> <p> Format for the density files:<br> # longitude latitude density (kg/m3)<br> <br> Files containing layer discontinuities:</p> <p> * ETOPO2_km_continental.xyz: surface elevation including ice sheet and 0 in marine areas (km, <0 upwards)</p> <p> * ETOPO2_km_depth_Ice.xyz: surface elevation including ice sheet (km, >0 downwards, <0 above sea level)</p> <p> * ETOPO2_km_depth_Bed.xyz: bedrock surface elevation without ice sheet (km, >0 downwards, <0 above sea level)</p> <p> * Global_Moho_WINTERC-G.xyz: crust-mantle discontinuity depth (km, >0 downwards)</p> <p> Format for the discontinuity files:<br> # longitude latitude depth (km)<br> <br> <br> The gravity field in WINTERC-G is computed using an spherical harmonic formalism and a model discretization<br> in 13 layers with laterally varying density. The first 7 layers are characterized by top and bottom boundaries with laterally varying radius whereas the last 6 layers are defined by top and bottom boundaries with constant radius:</p> <p>1/ Water: from ETOPO2_km_continental.xyz to ETOPO2_km_depth_Ice.xyz with rho=1030 kg/m3 (constant vertically)</p> <p>2/ Ice: from ETOPO2_km_depth_Ice.xyz to ETOPO2_km_depth_Bed.xyz with rho=910 kg/m3 (constant vertically)</p> <p>3/ Crust: from ETOPO2_km_depth_Bed to Global_Moho_WINTERC-G.xyz with rho=rho_c_out.xyz (constant vertically)</p> <p>4/ submoho-20km: from Global_Moho_WINTERC-G.xyz to z_20km (file with 20 km everywhere except where z_moho>20km) with rho=rho_submoho_out.xyz (top) and rho=rho_20km_out.xyz (bottom)</p> <p>5/ 20km-36km: from z_20km (file with 20 km everywhere except where z_moho>20km) to z_36km (file with 36 km everywhere except where z_moho>36km) with rho=rho_20km_out.xyz (top) and rho=rho_36km_out.xyz (bottom)</p> <p>6/ 36km-56km: from z_36km (file with 36 km everywhere except where z_moho>36km) to z_56km (file with 56 km everywhere except where z_moho>56km) with rho=rho_36km_out.xyz (top) and rho=rho_56km_out.xyz (bottom)</p> <p>7/ 56km-80km: from z_56km (file with 56 km everywhere except where z_moho>56km) to 80 km depth with rho=rho_56km_out.xyz (top) and rho=rho_80km_out.xyz (bottom)</p> <p>The next 6 layers are computed using the constant radius option:</p> <p>8/ 80km-110km: from z=80km to z=110 km with rho=rho_80km_out.xyz (top) and rho=rho_110km_out.xyz (bottom)</p> <p>9/ 110km-150km: from z=110km to z=150 km with rho=rho_110km_out.xyz (top) and rho=rho_150km_out.xyz (bottom)</p> <p>10/ 150km-200km: from z=150km to z=200 km with rho=rho_150km_out.xyz (top) and rho=rho_200km_out.xyz (bottom)</p> <p>11/ 200km-260km: from z=200km to z=260 km with rho=rho_200km_out.xyz (top) and rho=rho_260km_out.xyz (bottom)</p> <p>12/ 260km-330km: from z=260km to z=330 km with rho=rho_260km_out.xyz (top) and rho=rho_330km_out.xyz (bottom)</p> <p>13/ 330km-400km: from z=330km to z=400 km with rho=rho_330km_out.xyz (top) and rho=rho_400km_out.xyz (bottom)</p> <p> </p> <p> </p>
Sea ice satellite and model data between October 2014 and December 2015
<p>These images contain monthly averages for Arctic sea ice concentration and thickness between October 2010 and December 2015 from two simulations, one with the TOPAZ4 ocean-sea ice data assimilation system (Sakov et al., 2012 and https://resources.marine.copernicus.eu/?option=com_csw&view=details&product_id=ARCTIC_REANALYSIS_PHYS_002_003) and another with the S4K Arctic regional model. Image names are self-explanatory. Images with S4K results include a rectangle delimiting S4K domain, inserted in the larger TOPAZ4 domain. This was done to emphasize the transition between the two models, where the latter provides the boundary conditions to the former. </p> <p> </p>
Data and code used in "Satellite magnetic data reveal interannual waves in Earth's core"
<p>Eigen mode solutions and code to obtain them for the results presented in <a href="https://doi.org/10.1073/pnas.2115258119">Satellite magnetic data reveal interannual waves in Earth's core</a>. The package uses the freely available code <a href="https://github.com/fgerick/Mire.jl">Mire.jl</a>.</p> <p><strong>Prerequisites</strong></p> <p>Installed python3 with matplotlib ≥v2.1, cmocean and cartopy. A working Julia ≥v1.7.</p> <p><strong>Run</strong></p> <p>In the project folder run</p> <pre><code>julia --project=.</code></pre> <p><br> Then, from within the Julia REPL run</p> <pre><code>]instantiate</code></pre> <p>at first time, to install all dependencies.</p> <p>After that, to compute all plots, run</p> <pre><code>using QGMCSat allfigs()</code></pre> <p>They're automatically saved in the "figs" subfolder of the repository.</p> <p>If loading QGMCSat fails, due to a missing cartopy or cmocean in the python version. Run (within Julia)<br> </p> <pre><code>ENV["PYTHON"] = "python" #this should point to the python version that has cartopy installed ]build PyCall</code></pre> <p><br> To calculate all data, run</p> <pre><code>using QGMCSat calculate_data()</code></pre> <p>This will take several hours/days depending on the machine (needs enough memory).</p> <p>Individual data can be accessed directly through the .jld2 files from Julia. You can check out the individual figure functions to get an idea where which data is stored.</p> <p>If there are any issues or questions, please don't hesitate to get in touch!</p>
Satellite tracking data of white sharks in the southwest Indian Ocean (2012-2014)
<p>These data comprise locations and individual metadata from 34 white sharks (<em>Carcharodon carcharias</em>) instrumented March-May 2012 with telemetry devices along the coast of South Africa. These devices were SPOT5 transmitters (SPOT-257, SPOT-258; Wildlife Computers) which transmit locations via ARGOS CLS. All research methods were approved and conducted under the South African Department of Environmental Affairs: Oceans and Coasts permitting authority.</p> <p>This dataset is linked to the manuscript Kock et al. 2021 "Sex and size influence the spatiotemporal distribution of white sharks, with implications for interactions with fisheries and spatial management in the southwest Indian Ocean".</p> <p>The data are structured in long format, so that each row in the dataset represents an observation. The columns in the data are as follows.</p> <p>DeployID: This a factor variable identifying each individual shark. It has 34 levels.</p> <p>SPOT: This is a numeric variable identifying the tag number unique to each shark.</p> <p>Date: This is a date variable (POSIXct) that gives the date and time of a geographic location record in UTC time.</p> <p>Type: This is a character variable identifying the type of location record.</p> <p>Quality: This is a character variable made up of numbers and letters giving the location error associated with each location as provided by ARGOS.</p> <p>Latitude: This is a numeric variable and gives the latitude of the shark at the time of each record.</p> <p>Longitude: This is a numeric variable and gives the longitude of the shark at the time of each record.</p> <p>Area_tagged: This is a character variable that gives the area where the shark was tagged.</p> <p>Sex: This is a character variable identifying the sex of the shark, either "F" or "M" for female and male.</p> <p>TL: This is a numeric variable giving the total length of the shark in centimetres.</p> <p>Maturity: This is a character variable giving the maturity of the shark based on its total length following Malcolm et al. 2001: juveniles (male and female: 175-300 cm TL), sub-adults (male: >300-360 cm TL; females: >300-480 cm TL) and adults (male: >360 cm TL; female: >480 cm TL).</p> <p> </p>
Pop-up satellite archival tagging data of Atlantic bluefin tuna in the Gulf of Lions, Northwestern Mediterranean Sea
<p>24 Atlantic bluefin tuna (Thunnus thynnus) individuals (117–158 cm fork length) were tagged with pop-up archival tags in the Gulf of Lion, NW-Mediterranean Sea between 2015 and 2016.</p> <p><strong>Tag programming and data</strong></p> <p>The tags applied (miniPATs by Wildlife Computers, https://wildlifecomputers.com) can record depth and temperature time series (denoted hereafter as DepthTS and TempTS, respectively) at a temporal resolution of 3–5 s (depending on the predefined deployment duration) and a vertical resolution of 0.5 m. Based on these data, the tag calculates and stores additional data products such as PAT-style Depth–Temperature profiles (PDT), time at depth data, and time at temperature. After pop-up, the tags transmit user-defined data products and subsets from the recorded data sets. All our tags were configured to transmit the following data products: daily light curves, DepthTS, and PDT. In order to maximize data coverage of the transmitted datasets, we decreased the temporal resolution of the DepthTS and PDT data after the first tagging campaign in 2015 from 150 to 600 s and 6 to 24 h, respectively. For both years, deployment durations were set to 150 and 90 d during spring (April–May) and summer (August–September), respectively. A description of the electronic tagging procedure can be found in <a href="https://doi.org/10.1093/icesjms/fsaa083">Bauer et al. (2020)</a>.</p> <p>Seven tags were physically recovered, providing the complete archived time series data at a resolution of 3–5 s. Nineteen tags provided more than 7 d of complete DepthTS data (i.e. without transmission gaps). Three tags from 2016 had deployment durations of <1 week (#15P0983, #15P0985, and #15P0986) because of hardware failure.</p> <p>Provided files contain raw tag data (transmitted and recovered datasets) from the Wildlife Computers Data Portal as well as related GPE3 model runs (geolocation estimates).</p> <p>We thank the crews of the Cyngali and Roussillon Fishing recreational fishing vessels for their cooperation during the tagging cruises. This tagging study was part of the BLUEMED project and funded by the French National Research Agency (ANR; Project-ID ANR-14-ACHN-0002).</p> <p> </p>
Data associated with the publication "Interannual variability in the Australian carbon cycle over 2015-2019, based on assimilation of OCO-2 satellite data".
<p>This dataset refers to the publication "Interannual variability in the Australian carbon cycle over 2015-2019, based on assimilation of OCO-2 satellite data". https://doi.org/10.5194/acp-2022-15.</p> <p> </p>
Atmospheric Distribution of HCN from Satellite Observations and 3-D Model Simulations - TOMCAT data
<p>This repository contains the model data from the paper "Atmospheric Distribution of HCN from Satellite<br> Observations and 3-D Model Simulations" submitted to ACP.</p> <p>The files contains the monthly mean hydrogen cyanide (HCN) mixing ratios modelled using the TOMCAT 3-D offline chemical transport model with a horizontal resolution of 2.8° × 2.8° with 60 hybrid σ-pressure levels from the surface to ~60 km.</p>
TOMCAT model data & IASI/GOME-2B satellite data of European ozone between 2008 - 2023
<p>Daily mean data of ozone (O3) from the TOMCAT 3D chemical transport model (Chipperfield, 2006) and two satellite products, the Infrared Atmospheric Sounding Interferometer (IASI) on the MetOp-A & B satellites and the Global Ozone Monitoring Experiment-2 (GOME-2) on the MetOp-B satellite. The satellite observations are retrieved using schemes developed by the Rutherford Appleton Laboratory (RAL) see Miles et al. (2015) and Pope et al. (2021). The TOMCAT model data is available for 2017 - 2021, the IASI data is available for 2008 - 2023 and the GOME-2 data for 2015 - 2020. </p>
Extra data to accompany code in GitHub burntfields_punjab, both used in Walker et. al. 2022, Detecting crop burning in India using satellite data
<p>Supplementary data files to accompany GitHub code 'burntfields_punjab' supporting Walker et. al. (2022) Detecting crop burning in India using satellite data [<a href="https://arxiv.org/abs/2209.10148">available here</a>] and Jack et. al. (2024) Money (not) to burn: Payments for ecosystem services to reduce crop residue burning).</p> <p>Includes custom Sentinel-2 cloud masks and data from Sentinel-2 Spectral Mixture Analysis to highlight Char (burning) based on general concept and methods from Daldegan et. al (2019). Spectral mixture analysis in Google Earth Engine to model and delineate fire scars over a large extent and a long time-series in a rainforest-savanna transition zone. Remote Sensing of Environment 232, 111340. </p> <p>Note: Bands in weekly BASMA layers are: 0 = green vegetation, 1 = Non-productive vegetation and bare soil, 2 = Char (burned).</p> <p>further details are provided at: <a href="https://github.com/klwalker-sb/burntfields_punjab">https://github.com/klwalker-sb/burntfields_punjab</a> (archived at: <a href="https://doi.org/10.5281/zenodo.11225292" target="_blank" rel="noopener">DOI: 10.5281/zenodo.11225292</a>)</p>
Solar Asset Mapper: A continuously-updated global inventory of solar energy facilities built with satellite data and machine learning
<p><strong>TransitionZero’s Solar Asset Mapper is a global, satellite-derived dataset of utility-scale solar farms generated with a combination of machine learning and human annotation. Our Q1 2024 dataset contains the location and shape of 63,616 assets, along with estimated capacities. We estimate the construction date for over 80% of these assets. The dataset contains over 19,100 square kilometres of solar farms across 183 countries, with a total estimated capacity of 711 GW.</strong></p> <p>Download the dataset, read the explainer and explore our polygon browser UI at <a href="https://www.transitionzero.org/products/solar-asset-mapper" target="_blank" rel="noopener">TransitionZero.org.</a></p> <p><a href="https://blog.transitionzero.org/hubfs/Data%20Products/TZ-SAM/tz-sam-scientific-methodology-Q12024.pdf" target="_blank" rel="noopener">Download our methodology paper here </a></p> <h1><strong>1. Dataset Description</strong></h1> <p>We publish six files.</p> <ul> <li><em>analysis_polygons.gpkg:</em> our “analysis-ready” dataset containing geometries, capacity estimates and construction date estimates.</li> <li><em>analysis_polygons.csv:</em> a version of analysis_polygons.gpkg containing a central latitude and longitude in place of a geometry, to allow parsing without geospatial software.</li> <li><em>sources.csv</em>: a table mapping the IDs of our analysis-ready dataset to the raw geometries that make them up.</li> <li><em>raw_polygons.gpkg:</em> the raw geometries used to compose analysis_polygons.gpkg.</li> <li><em>TZ Solar Asset Mapper Q1 2024.xlsx</em>: an Excel formatted version of the analysis_polygons.csv file.</li> <li><em>tz-sam_scientific_data.pdf</em>: A pre-print aricle that explains the methodology in detail.</li> </ul> <h2><strong>1.1 Analysis-level datasets</strong></h2> <p>Our analysis-level dataset comprises our most complete view of global asset-level solar installations, incorporating our own detections as well as known solar farm geometries from other datasets.</p> <p>The geospatial dataset contains the following fields:</p> <ul> <li>id: unique ID for the asset</li> <li>geometry: Polygon or MultiPolygon defining the asset</li> <li>capacity_mw: estimated capacity of the asset in megawatts</li> <li>constructed_before: upper bound for construction date (estimated date of the image in which the solar plant was first seen in a constructed state)</li> <li>constructed_after: lower bound for construction date (estimated date of the image in which construction began for the solar plant)</li> </ul> <p>The CSV version replaces the Geometry column with:</p> <ul> <li>latitude: the latitude of the centroid of the asset</li> <li>longitude: the longitude of the centroid of the asset</li> <li>country: administrative country name</li> </ul> <h2><strong>1.2 Raw datasets and sources</strong></h2> <p>The analysis-level datasets hide some complexity in the underlying data that we expose in the <em>raw_polygons</em> and <em>sources</em> file.</p> <ul> <li>We produce new sets of polygons for each run. Often these overlap, sometimes in complicated ways.</li> <li>We cluster together overlapping and nearby geometries from both our detections and external sources. Currently these sources are:</li> <li>Large solar farms scraped from OpenStreetMap (OSM)</li> <li>Validated geometries from <a href="../records/5005868">Kruitwagen et. al., A global inventory of solar photovoltaic generating units</a>.</li> </ul> <p>Each cluster comprises one row in the analysis-level dataset. In order to enable tracking raw detections from run to run, as well as to provide detailed sourcing information, we provide all of these raw polygons, along with a source file that lists all of the raw polygons contained in each analysis-level polygon.</p> <p>raw_polygons.gpkg contains the following fields:</p> <ul> <li>id: ID of the raw source polygon</li> <li>geometry<strong>: </strong>Polygon or MultiPolygon defining the asset</li> <li>source: either “solar asset mapper”, “osm” or “2019_global_pv”.</li> <li>acquisition_date: for solar asset mapper polygons, this is the date of the inference run that produced the polygon; for OSM polygons it is the date that the polygon was scraped from OSM; for 2019_global_pv it is 2019-01-01, the approximate detection date of that dataset.</li> </ul> <p>Sources.csv contains the following fields:</p> <ul> <li>cluster_id: ID of the corresponding item in the analysis-level dataset</li> <li>source_id: ID of the raw source polygon</li> <li>source: either “solar asset mapper”, “osm” or “2019_global_pv”.</li> <li>acquisition_date: for solar asset mapper polygons, this is the date of the inference run that produced the polygon; for OSM polygons it is the date that the polygon was scraped from OSM; for 2019_global_pv it is 2019-01-01, the approximate detection date of that dataset.</li> </ul> <h2><strong>1.3 Caveats and limitations</strong></h2> <h3><strong>1.3.1 Capacity Estimates</strong></h3> <p>While we have made every effort to remove false positives from the published dataset, some will remain due to the difficulty of manually validating detections in 10-metre satellite imagery. To estimate false positive prevalence throughout the data a subset of approximately 2000 detections were selected at random from our positively labelled solar assets. Each of these were validated through a higher degree of scrutiny utilising high-resolution imagery. This analysis yielded an expected rate of false positives of around 1%.</p> <h3>1.3.2 Plant Shapes</h3> <p>Our plant outlines are not perfect. They will occasionally be much smaller or larger than the underlying plant. Our tests show that on average, these effects average out.</p> <h3>1.3.3 Capacity Updates</h3> <p>Our capacity estimation model should produce relatively unbiased country-level aggregates, since it is trained to learn the typical ground coverage ratio of plants by country. The model has no way to distinguish between a very dense and a very sparse (e.g. dual-axis-tracking) plant in the same country. Plants with unusually high or low ground coverage ratios will not have accurate capacity estimates.</p> <h3>1.3.4 Construction Date Estimates</h3> <p>We are not able to directly estimate the construction date of a plant. We estimate an upper bound (the date of the image in which the plant was first seen in constructed state) and a lower bound (the date of the image in which the plant was last seen in an unconstructed state). For plants that were constructed before the launch date of Sentinel-2 in 2017, we produce only an upper bound.</p> <p>We leave it to consumers of the data to interpret these bounds and/or estimate likely grid connection dates.</p> <p><strong>2. Attribution</strong></p> <p>TZ-SAM is made available under a Creative Commons Attribution Non-Commercial 4.0 International License (CC-BY-NC-4.0). Attribution to TransitionZero is required. You must also clearly indicate if you have made any changes to the TZ-SAM dataset and what these are. Please refer to the suggested citation formats:</p> <ul> <li>“TransitionZero Solar Asset Mapper, TransitionZero, May 2024 release.”</li> <li>“TZ-SAM, TransitionZero, May 2024 release.”</li> <li>“TransitionZero (2024) Solar Asset Mapper.”</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.