Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
242
datasets available to search
ShareScore release 0.9.0
Dataset results
242 results for “Spatial Dataset”
Temporal and spatial heat exposure in Colombo - dataset
<p>The dataset consists of two shapefiles that provide information on outdoor heat stress and anthropogenic heat flux in Colombo, Sri Lanka in 500 m resolution.</p> <p>1. "heat_stress_indicator.shp" contains the outdoor heat stress indicators, including heat index (HI), humidex (HD), and discomfort index (DI). These indicators were calculated using the modelled variables of urban land surface model SUEWS, and the values represent the averages at 7:00, 14:00, 19:00, and 23:00 during a heatwave (23-28, 2020) in Colombo, Sri Lanka. The results are for each 500 m grid across Colombo.</p> <p>2. "QF.shp" contains the calculated anthropogenic heat flux for each 500 m grid in Colombo, Sri Lanka. Published as: Blunn, L., Xie, X., Grimmond, S., Luo, Z., Sun, T., Perera, N., Ratnayake, R. and Emmanuel, R., 2024. Spatial and temporal variation of anthropogenic heat emissions in Colombo, Sri Lanka. Urban Climate, 54, p.101828. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.uclim.2024.101828" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.uclim.2024.101828</span></span></a></p> <p> </p>
Spatial datasets associated with decontamination and remediation operations following the Fukushima nuclear accident, Japan (2011–2023)
<p>At the onset of the full reopening in Spring 2023 of the Difficult-to-Return Zone of Northeastern Japan following the Fukushima Daiichi Nuclear Power Plant (FDNPP) accident that took place in March 2011, several spatial layers were regrouped and compiled to facilitate environmental studies dealing with the redistribution of radiocesium fallout across landscapes.</p> <p><strong>The current dataset is composed of 23 shapefiles including those of the delineations of different spatial zones (Intensive Contamination Survey Areas – ICAs, Special Decontamination Zones – SDZ, Difficult-to-Return Zone –</strong> <strong>DTRZ, and FNDPP location) (Evrard et al. 2019), municipalities where mushroom consumption restrictions were enforced (restricted and partially lifted restrictions), river hydrographic networks and their respective drainage areas (Mano, Niida, Ota, Takase, and Ukedo), dam reservoirs and drainage areas (Mano, Ogaki, Takanokura, and Yokokawa), multiple administrative delineations in Japan (whole Japan administrative boundaries, Prefectures, and municipalities) (GIS, 2016), and one raster file of the reconstruction of initial <sup>137</sup>Cs fallout across eastern Japan (from Kato et al., 2019).</strong></p> <p><strong>The current dataset provides a support to a publication submitted to the SOIL journal:<br></strong></p> <div> <div><strong>Evrard, O., Chalaux-Clergue, T., Chaboche, P.-A., Wakiyama, Y., and Thiry Y. (2023). Research and Management Challenges Following Soil and Landscape Decontamination at the Onset of the Reopening of the Difficult-To-Return Zone, Fukushima (Japan)’. <em>SOIL</em> 9: 479–97. <a href="https://doi.org/10.5194/soil-9-479-2023">https://doi.org/10.5194/soil-9-479-2023</a>. </strong></div> <div> </div> </div> <p>All map processing was carried out using QGIS 3.26.0 (QGIS, 2022) and under the EPSG:WGS 84 projection system.</p> <p>The <sup>137</sup>Cs fallout raster (in Bq m<sup>-2</sup>, decay-corrected to July 2011) was generated from the point grid of Kato et al. (2019). A total of 126 tiles (0.25 x 0.25 degree) were generated by Inverse Distance Weighted (IDW) interpolation using the '<em>IDW interpolation'</em> tool with the following settings: distance coefficient P = 1.0 and pixel size (x and y) = 0.0015 degree. Tiles were then merged into a single tile using the raster<em> 'Merge'</em> tool. The initial point grid footprint was manually delineated to define the spatial applicability zone of the airborne survey. A buffer zone corresponding to half plus 10% of the longest distance between two airborne points (x = 0.002, y = 0.003), i.e. 0.0017 degree, was generated using the '<em>buffer'</em> tool. The single tile was then cut according to the footprint of the buffer zone using the <em>'clip a raster by a mask layer'</em> tool. A <em>single-band pseudo-colour </em>scale is provided and displays pixels with a value above 1000 Bq m<sup>-2</sup> (eq. global background).</p>
SPVPANELEX: Dataset containing aerial orthoimages (covering 257.93 km2 of the Spanish territory, with a spatial resolution of 0.5 m) labelled with photovoltaic panel information for binary recognition and semantic segmentation
<p>The data have been generated using scripts developed in Python with Open-Source libraries (GDAL/OGR and MapScript) to rasterize of vector cartography representing the photovoltaic (PV) panels instalations in urban, industrial, and rural areas. This PV panels cartography has been generated by manual digitalizing the PV panels found latest aerial orthofotographs available on June 1, 2021 from Plano Nacional de Ortofotografía Aérea (PNOA), produced by the National Geographic Institute of Spain, using the Web Map Service PNOA-MA.<br> <br> The dataset consists of 239,680 images of 256 × 256 pixels in size, in png format, labelled with Class_1: “Contains PV panel” and Class_2: “Does not contain PV panel”, that were pre-divided with a split criterion of 70:10:20%. in train, validation and test folders, respectively.<br> <br> The structure of the data is as follows:<br> 1-Panels-Ortho and 1-Panels-Mask contain the images featuring PV panels and their corresponding ground truth mask for training the semantic segmentation networks.<br> 1-Panels-Ortho and 2-NoPanels-Ortho contain images containing and not containing PV panels, for the training of binary recognition models of PV panels.<br> <br> Moreover, in each folder the structure is the same: train, test, validation containing 70%, 10% and 20% of the total images and masks of each type.<br> <br> 1-Panels-Ortho<br> |----Train<br> |----Test<br> -----Validation<br> <br> 1-Panels-Mask<br> |----Train<br> |----Test<br> -----Validation<br> <br> 2-NoPanels-Ortho<br> |----Train<br> |----Test<br> -----Validation</p>
Multiscale Spatial Patterns in Giant Dike Swarms Identified through Objective Feature Extraction Datasets
<p>S1 - Linked dike clusters for the Columbia River Flood Basalt group including the four identified subswarms: Chief Joseph, Monument, Ice Harbor, and Steens as compiled in Morriss et al., 2020. This dataset uses the a UTM Zone 11N projection (EPSG:26911).</p> <p>S2 - Linked dike clusters for the Deccan Traps including the four identified subswarms: Saurashtra, Narmada-Tapi, Central and Coastal. Due to their overlap Central and Coastal Swarms have been combined in this dataset into the Central Swarm. This dataset uses the a WGS 84 projection (EPSG:3857). </p> <p>S3 - Dike segment data for Spanish Peaks and Dike Mountain located in the Rio Grande Rift of Colorado. This dataset was digitized using QGIS based on the map by Johnson (1961). This dataset uses the a UTM Zone13N projection (EPSG:32613). The file includes the start, end points, and midpoints of the dikes; segment length; calculated $\rho$ and $\theta$ for the Hough Transform; the origin used for the Hough Transform which is different for each subswarm (xc,yc); dike rock type if known; and a unique identification calculated based on the start and endpoints. This dataset has been preprocessed to remove curving dikes and is the data set used to produce later products (Data set S4). </p> <p>S4 - Linked dike clusters for the Spanish Peaks and Dike Mountain. This dataset was produced using the Agglomerative Clustering algorithms using the parameters set in Table 1. This dataset uses the a UTM Zone 13N projection (EPSG:32613). </p> <p> </p> <p> </p> <p>These datasets were produced using the Agglomerative Clustering algorithms using the parameters set in Table 1. The datasets are in the format of a CSV file but can be read into GIS programs using Well Known Text (WKT) linestring. TThe file includes the start and end points of the average line in the cluster and it's mid points, cluster length and width (Xstart, Xend, Xmid, Ymid, in meters and UTM coordinates, Dike Cluster Width or R\_Width, Dike Cluster Length or R\_Length all in meters); calculated average $\rho$ and $\theta$ for the Hough Transform $\rho$ units measured in meters, $\theta$ units measured in degrees, unless otherwise stated); the origin used for the Hough Transform which is different for each subswarm ($xc$,$yc$, meters in UTM coordinates); average slope and intercept (AvgSlope, AvgIntercept meters); range and standard deviation for $\rho$ and $\theta$ for all objects in the cluster ($\rho$ units measured in meters, $\theta$ units measured in degrees); cluster size (Size); sum of segment lengths in a cluster (SegmentLSum, meters); whether the cluster crosses between negative and positive values (ClusterCrossesZero, boolean); overlap as calculated in the main text where the length of overlap is normalized by the sum of segment lengths in a cluster; maximum number of overlapping segments (nOverlapingSegments); twist angle which is the difference in angle betweeen the average cluster line and the average line formed by cluster midpoints (EnEchelonAngleDiff, degrees); the p-value for the midpoint line fit of the segments where $p<0.05$ is considered to be a significant fit (EEPValue); the maximum, median, and minimum segment nearest neighbors distances in the cluster which is calculated using the cartesian midpoints of each segment and normalized by the Cluster Length (MaxSegNNDist, MedianSegNNDist, MinSegNNDist); characterization of each cluster as filtered or not, filtered clusters are of size greater than $3$ and have a MaxSegNNDist of less than $0.5$ (TrustFilter, boolean); the date edited (Date\_Changed), and the clustering parameters used for each cluster (Rho\_Threshold in meters, Theta\_Threshold in degrees) and a unique identification calculated based on the start and endpoints (ClusterHash). </p>
Gene expression dataset of the Spatially Resolved Single-cell Translatomics at Molecular Resolution
<p>Here are the gene expression datasets of RIBOmap included in "<strong>Spatially Resolved Single-cell Translatomics at Molecular Resolution</strong>" from Zeng et al. Please refer to the README file for more detailed information. </p> <p> </p> <p><strong>Abstract</strong></p> <p>The precise control of mRNA translation is a crucial step in post-transcriptional gene regulation of cellular physiology. However, it remains a major challenge to systematically study mRNA translation at the transcriptomic scale with spatial and single-cell resolution. Here, we report the development of RIBOmap, a three-dimensional (3D) in situ profiling method to detect mRNA translation of thousands of genes simultaneously in intact cells and tissues. By applying RIBOmap to 981 genes in HeLa cells, we revealed a remarkable dependency of translation on cell-cycle stages and subcellular localization. Furthermore, we profiled single-cell translatomes of 5,413 genes in adult mouse brain tissues yielding a spatial cell atlas of 119,173 cells. The pairwise spatial mapping of single-cell translatome and transcriptome in two adjacent mouse brain slices revealed cell-type and brain-region-dependent translational regulation and suggested a translation remodeling during oligodendrocyte lineage maturation. The spatial translatome profiling detected widespread patterns of localized translation in neuronal and glial cells in intact brain tissue networks. Together, RIBOmap presents the first spatially resolved single-cell translatomics technology, accelerating our understanding of protein synthesis in the context of subcellular architecture, cell types, and tissue anatomy.</p>
Output raster datasets from an application of a fine resolution spatially explicit forest water yield model in Florida's panhandle
<p>These raster datasets are the output results for a spatial water yield model applied to an 11 county area in the state of Florida panhandle. The water yield model is adapted from Acharya, et al. 2022 and the spatial modelling process is detailed in this datasets associated publication. All data are in the WGS 1984 UTM Zone 16N coordinate system and have 10m horizontal spatial resolution. </p> <p>The output raster datasets contained here are water yield estimate informed with 2018 pine basal area, binary depth to water table, and average aridity index input rasters. These rasters have 10m spatial resolution, the raster extent covers 11 counties in the panhandle of Florida, the units are in centimeters of water yield per year. The water yield outputs consist of ten rasters representing the current water yield using the mean aridity raster, the water yield expected from the three pine tree thinning scenarios: 7m/hectare ba, 11 m/hectare, and 18 m/hectare, taken from the mean aridity index. Then rasters representing the water yield expected from the three thinning scenarios under maximum, and minimum aridity indexes.</p> <p> </p> <p>These ten outputs are listed here:</p> <p>"wy_current_mean" Based on 2018 BA conditions; Mean ARID</p> <p>"wy_18_mean" BA reduced to 18m2ha-1; Mean ARID</p> <p>"wy_11_mean" BA reduced to 11m2ha-1; Mean ARID</p> <p>"wy_7_mean" BA reduced to 7m2ha-1; Mean ARID</p> <p>"wy_ 18_max" BA reduced to 18m2ha-1; Maximum ARID</p> <p>"wy_ 11_max" BA reduced to 11m2ha-1; Maximum ARID</p> <p>"wy_7_max" BA reduced to 7m2ha-1; Maximum ARID</p> <p>"wy_ 18_min" BA reduced to 18m2ha-1; Minimum ARID</p> <p>"wy_ 11_min" BA reduced to 11m2ha-1; Minimum ARID</p> <p>"wy_ 7_min" BA reduced to 7m2ha-1; Minimum ARID</p> <p> </p> <p>Water yield was estimated for 2018 using the following datasets to inform the model in the Current Water Yield Calculation tool:</p> <ul> <li>Leaf area index modeled from a 2018 pine species basal area raster,</li> <li>Depth to water table data provided by Florida Geological Survey and reclassified as a binary raster,</li> <li>Average aridity index raster generated with precipitation data from PRISM Climate Group and MODIS PET data.</li> </ul> <p>For detailed information on how the above inputs were developed, please see the associated publication:</p> <p>Vernon, J., St. Peter, J., Crandall, C., Awowale, O.E., Medley, P., Drake, J., & Ibeanusi, V. (2023). Spatial application of southern pine water yield for prioritizing forest management activities. ISPRS International Journal of Geo-Information, 12(2), 34. <a href="https://doi.org/10.3390/ijgi12020034">https://doi.org/10.3390/ijgi12020034</a> </p>
Input raster datasets for an application of a fine resolution spatially explicit forest water yield model in Florida's panhandle
<p>These raster datasets are the inputs for a spatial water yield model applied to an 11 county area in the state of Florida's panhandle. The water yield model is adapted from Acharya, et al. 2022 and the spatial modelling process is detailed in the associated publication. The five input datasets required for this water yield analysis are: 1) a model of pine species basal area, named "ARSA_PineBA_10m" 2) a binary depth to water table raster named "DTW_cm_binary2" , and 3) three spatial aridity index raster dataset named "Aridity_Min", "Aridity_Max" and "Aridity_Mean", created from potential evapotranspiration, and precipitation raster datasets. The min max and mean codifiers relate to the range of aridity values found in our dataset of 7 year temporal range, from MODIS PET and PRISM percipitation yearly data. All input and output data are in the WGS 1984 UTM Zone 16N coordinate system and have 10m horizontal spatial resolution. </p>
mapspamc_db: a database with global spatial datasets to support the implementation of the mapspamc R package.
<p>This repository contains the mapspamc database (mapspamc_db), a collection of global spatial datasets to support the implementation of the <a href="https://github.com/michielvandijk/mapspamc">mapspamc</a> R package. The database also includes subnational crop statistics and matching country shapefiles for several country examples. For more information on how to use the mapspamc package in combination with mapspamc_db, see the <a href="https://michielvandijk.github.io/mapspamc/">mapspamc documentation</a>. Detailed information on the contents of mapspam_db, such as the sources of information and pre-processing is described in the mapspamc_db documentation (pdf file) that is part of the repository.</p>
High spatial resolution dataset of grapevine yield components at the within-field level
<p>This dataset comprises a comprehensive mapping of vine yield at the plant scale over two vine fields located in the southern region of France. Both vine fields were planted with the Vitis vinifera : cv. Syrah. The first field (Field 1) occupies 0.8 ha and data were collected in 2022, while the second field (Field 2) has an area of 0.5 ha and data were collected in 2008. Throughout the growing season, information regarding unproductive vines, inflorescence number, and bunch weight was collected for both vine fields. For both fields, at the flowering stage, the location of each productive and unproductive vines (dead and missing vines) was georeferenced, and the number of inflorescences was manually counted for all productive vines. For Field 1, at harvest, all bunches of the field were manually weighed with an accuracy of ±1 gram and georeferenced precisely (one point per vine). For each vine, total yield (grams per vine) was then computed as as the sum of the weight of its bunches. For Field 2, at harvest, the total yield per vine was estimated based on the weighing of representative bunches obtained from several regularly spaced set of 5 vines. In addition to the yield data, two ancillary data, including soil apparent resistivity measurements and common vegetative index derived from remote sensed imagery, are provided for both vine fields. Overall, the dataset consists of 3644 vines, with 2151 being productive, along with a total count of 33354 inflorescences and 19635 manually weighed bunches at harvest.</p> <p>Raw data includes 9 shapefiles (.shp), one per data type and per field.</p> <ul> <li>“Field1_Dead_Missing_Vines.shp” (Figure 1.A) contains the location of missing and dead vine identified in Field 1;</li> <li>“Field1_Inflorescences.shp” and “Field2_Inflorescences.shp” (Figure 1.B and Figure 1.C) both contain the location and the number of inflorescences per vine counted during flowering;</li> <li>“Field1_Final_Yield.shp” and “Field2_ Final_Yield.shp” (Figure 1.D and Figure 1.E) both contain location and measured values of yield weight per vine at harvest. For Field 1, the list of the bunch weight is also available;</li> <li>“Field1_Soil_Resistivity.shp” and “Field2_ Soil_Resistivity.shp” (Figure 1.F and Figure 1.G) contain the electrical resistivity measurements of the soil on each field;</li> <li>“Field1_Vegetation_Index.shp” and “Field2_ Vegetation_Index.shp” (Figure 1.H and Figure 1.I) contain vegetation index values, NDVI (Normalized Difference Vegetation Index, without unit) for Field 1 and FCover (Fraction of vegetation Cover, in %) for Field 2.</li> </ul> <p>Aggregated data are composed of three csv files, two for Field 1 and one for Field 2. “Field1_Yield.csv” and “Field2_Yield.csv” aggregate all available yield data for each vine plant (either productive or other types). In both these csv files, each line represents a planted vine. Another file named "Field1_Bunches.csv" contains another representation of the data for Field 1. In this file, each row corresponds to a weighed bunch.</p> <p>A Data in Brief article is associated to this dataset.</p>
High spatial resolution dataset of downscaled LUH2 land use scenarios for Belgium (10 m and 100 m)
<p>This dataset comprises high-resolution land use data downscaled from LUH2 scenarios for Belgium at both 10 m and 100 m resolutions. These datasets were generated based on research conducted by Rashidi et al. in 2023 and published in the Land Journal. We employed the GLOBIO land allocation routine to downscale fractional land use data, originally at a 0.25° resolution (approximately 25 km), into discrete land use maps at 10 m and 100 m resolutions. This process utilized three distinct reference land cover maps: ESA WorldCover at 10 m resolution, ESA WorldCover upscaled to 100 m resolution, and CORINE land cover at 100 m resolution.</p> <p>During the downsizing process, we considered three SSP-RCP scenarios to model land use trends for both the present and the year 2050 on a national scale in Belgium. Key components of the model included regional land use demand, an assessment of grid cells' suitability for various land use types, and a reference land cover map. It's important to note that the classification system used in the reference maps differs from that of LUH2. To ensure comparability for land use simulations, we conducted a reclassification process following the methodologies outlined by Pérez-Hoyos et al. (2012), Dong et al. (2018), and Liao et al. (2020). This reclassification consolidated land use classes, except for water, into seven general categories: 1) urban, 2) cropland, 3) pasture, 4) forestry, 5) secondary vegetation, 6) undefined, and 7) natural.</p> <p>The raw data consists of three folders corresponding to the three reference maps, each containing four TIFF files (.tif), one for each scenario type.</p>
Associated Dataset for Genome-wide DNA methylation patterns in bumble bee (Bombus vosnesenskii) populations from spatial-environmental range extremes
<p>The dataset contains the final methylation call set (n=14,627,533), variant calling file for population genomics analyses, analysis codes/scripts, and other associated files related to the research (Constitutive and variable patterns of genome-wide DNA methylation in populations from spatial-environmental range extremes of the bumble bee <em>Bombus vosnesenskii)</em>. Raw WGBS reads generated in this study have been deposited and are currently available at the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA) under NCBI BioProject PRJNA956115.</p>
Soil texture dataset from the publication: "Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula'
<p>Clay, silt and sand distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The coefficient of variation and quantile data represent the spatial uncertainty of the predictions. For more information about the methodology used, users are referred to the article: </p> <p>Siqueira, R.G., Moquedace, C.M., Francelino, M.R., Schaefer, C.E.G.R., Fernandes-Filho, E.I., 2023. Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula. Geoderma 432, 116405. https://doi.org/10.1016/j.geoderma.2023.116405</p> <p>The .zip file has the following folders:</p> <p>1) soil_texture_antarctica: soil texture information containing clay, silt and sand contents</p> <p>2) soil_texture_coefficient_variation: uncertainty from the coefficient of variation of the soil texture prediction</p> <p>3) soil_texture_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil texture prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil texture prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil texture prediction</p>
Dataset for Task-dependent spatial processing in the visual cortex
<p>The dataset contains the mean ERPs values of each participant after the audiovidual stimulus (S2 for the spatial bisection and S for the spatial localization), divided as follows:</p> <p>- condition (i.e., 1sc: short distance between S1 and S2, 1sl: long distance between S1 and S2),</p> <p>- task (i.e., spatial bisection or spatial localization),</p> <p>- time window (i.e., 50-90 ms or 110-160 ms post stimulus), </p> <p>- roi (i.e., O1, O2, C1, C2, T7 or T8 electrodes). </p>
Global dataset of plant diversity and the spatial variability of grassland biomass from NutNet
While there is strong evidence of diversity effects on temporal variability of productivity, whether this mechanism extends to variability across space remains elusive. Here, we present data from Nutrien Network (www.nutnet.org) that were used to determine the relationship between diferent scales of plant diversity and spatial variability of productivity in 83 grasslands worldwide, and to quantify the effect of experimentally increased spatial heterogeneity in environmental conditions on this relationship. There are two data sets, one for the pre-treatment (observational_data.csv) data, and other for the experimentally increased heterogeneity (increased heterogeneity.csv). In these data sets, study sites contained at least three replicates that originated from blocks each composed of ten 5 m × 5 m plots. In addition, pre-treatment data has a subset of sites in where soil conditions where measured (observational_data_soil.csv) and a version in which data for each site are sumarized and site-level climatic variables obtained from WorldClim (www.worldclim.org) are added. (observational_data_site_climate.csv). If you need any clarification or further information, please contact us.
Wyoming Oil and Gas Development Spatial Datasets
<p>This file contains a file geodatabase with spatial datasets that accompany the scientific paper titled: "Recent Greater Sage Grouse (Centrocercus urophasianus)<br /> Population Dynamics in Wyoming Are Primarily Driven by<br /> Climate, not Oil and Gas Development" (Ramey, Thorley, Ivey 2015).</p>
Temporal and spatial evaluation of satellite-based rainfall estimates across the complex topographical and climatic gradients of Chile (datasets)
<p>This file contains the <strong>dataset</strong> (both raw observed precipitation data and figures obtained as output of the analysis) accompanying the manuscript '<strong>hess-2016-453'</strong> submitted to the HESS journal (http://www.hydrology-and-earth-system-sciences.net/).</p> <p> </p> <p><strong>Title</strong>: "Temporal and spatial evaluation of satellite-based rainfall estimates across the complex topographical and climatic gradients of Chile"</p> <p> </p> <p><strong>Abstract</strong></p> <p>Accurate representation of the real spatio-temporal variability of catchment rainfall inputs is currently severely limited. Moreover, spatially interpolated catchment precipitation is subject to large uncertainties, particularly in developing countries and regions which are difficult to access (e.g., high elevation zones). Recently, satellite-based rainfall estimates (SRE) provide an unprecedented opportunity for a wide range of hydrological applications, from water resources modelling to monitoring of extreme events such as droughts and floods. </p> <p>This study attempts to exhaustively evaluate -for the first time- the suitability of seven state-of-the-art SRE products (TMPA 3B42v7, CHIRPSv2, CMORPH, PERSIANN-CDR, PERSIAN-CCS-adj, MSWEPv1.1 and PGFv3) over the complex topography and diverse climatic gradients of Chile. Different temporal scales (daily, monthly, seasonal, annual) are used in a point to-pixel comparison between precipitation time series measured at 366 stations (from sea level to 4600 m a.s.l. in the Andean Plateau) and the corresponding grid cell of each SRE. The modified Kling-Gupta efficiency was used to identify possible sources of systematic errors in each SRE. In addition, several categorical indices were used to assess the ability of each SRE to correctly identify different precipitation intensities.</p> <p><br> Results revealed that most SRE products performed better for the humid South (36.4-43.7ºS) and Central Chile (32.18-36.4ºS), in particular at low- and mid-elevation zones (0-1000 m a.s.l.) compared to the arid northern regions and the Far South. Seasonally, all products performed best during the wet seasons (MAM-JJA) compared to summer (DJF) and autumn (SON). In addition, all SREs were able to correctly identify the occurrence of no rain events, but they presented a low skill in classifying precipitation intensities during rainy days. Overall, PGFv3 exhibited the best performance everywhere and for all time scales, which can be clearly attributed to its bias-correction procedure using 213 stations from Chile. Good results were also obtained by CHIRPSv2, TMPA 3B42v7 and MSWEPv1.1, while CMORPH, PERSIANN-CDR and PERSIANN-CCS-adj were not able to represent observed rainfall. While PGFv3 (currently available up to 2010) might be used in Chile for historical analyses and calibration of hydrological models, the high spatial resolution, low latency and long data records of CHIRPS and TMPA 3B42v7 (in transition to IMERG) show promising potential to be used in meteorological studies and water resources assessments. We finally conclude that despite improvements of most SRE products, a site-specific calibration is still needed before any use in catchment-scale hydrological studies.</p>
Dataset from Reese et al.: "Local Mixing Determines Spatial Structure of Diahaline Exchange Flow in a Mesotidal Estuary: A Study of Extreme Runoff Conditions" - PART 1
<p>Model data from the numerical setup of the tidal Elbe presented in Reese et al. (2023): "Local Mixing Determines Spatial Structure of Diahaline Exchange Flow in a Mesotidal Estuary: A Study of Extreme Runoff Conditions" [1]</p><p>PART 1</p><p>Each file contains data for a full month, as given through the file naming convention: description.YYYYMMDD.nc4</p><p>The numerical model uses terrain-following sigma coordniates, with sigma level 0 being the bottommost layer.</p><p>Certain variables are also given in salinity class bins of dimension salt_s instead of vertical coordinates.</p><p>Explanation of each data type:</p><ul><li> 2D_elv_all: Spatially resolved simulated surface elevation from 08/2012 to 12/2013: Tidal analysis Fig. 4, Table 1 (simulated surface elevation vs. time)<ul><li>5 min snapshots</li></ul></li><li>Elbe_dia_getm_all: Diahaline analysis Fig. 8, 11: on-line GETM computation of u_dia,z^S in September 2012 and June 2013<ul><li>44700s temporal resolution (M2 tidal period); averaged over each period</li></ul></li><li>Elbe_TEF_mean_all: Total Exchange Flow analysis in September 2012 and June 2013, Fig.s 7, 8<ul><li>1-hourly averages</li></ul></li><li>Mixing_mean_all: Physical and numerical Mixing from 08/2012 to 12/2013. Fig. 7, 8, 9, 10, 11<ul><li>44700s temporal resolution (M2 tidal period); averaged over each period</li></ul></li><li>ST_stations: Surface elevation at given location for comparison with observational data at named station from 08/2012 to 12/2013<ul><li>5 min snapshots</li></ul></li><li>SST_stations: Salinity and temperature at given location for comparison with observational data at named station from 08/2012 to 12/2013; Fig. 3, Fig. 5, Table 2<ul><li>30-min snapshots</li></ul></li></ul><p> </p><p>[1] L. Reese, U. Graewe, K. Klingbeil, X. Li, M. Lorenz, H. Burchard, 2023:</p><p> Local mixing determines spatial structure of diahaline exchange flow in a</p><p> mesotidal estuary – a study of extreme runoff conditions.</p><p> J. Phys. Oceanogr., in press.</p>
Dataset from Reese et al.: "Local Mixing Determines Spatial Structure of Diahaline Exchange Flow in a Mesotidal Estuary: A Study of Extreme Runoff Conditions" - PART 2
<p>Model data from the numerical setup of the tidal Elbe presented in Reese et al. (2023): "Local Mixing Determines Spatial Structure of Diahaline Exchange Flow in a Mesotidal Estuary: A Study of Extreme Runoff Conditions" [1]</p><p>PART 2</p><p>Each file contains data for a full month, as given through the file naming convention: description.YYYYMMDD.nc4</p><p>The numerical model uses terrain-following sigma coordniates, with sigma level 0 being the bottommost layer.</p><p>Certain variables are also given in salinity class bins of dimension salt_s instead of vertical coordinates.</p><p>Explanation of each data type:</p><ul><li>Mean_all: Spatially resolved, temporally varying salt distribution in the Elbe estuary: Fig. 6, 10<ul><li>1-hourly averages</li></ul></li></ul><p> </p><p>[1] L. Reese, U. Graewe, K. Klingbeil, X. Li, M. Lorenz, H. Burchard, 2023:</p><p> Local mixing determines spatial structure of diahaline exchange flow in a</p><p> mesotidal estuary – a study of extreme runoff conditions.</p><p> J. Phys. Oceanogr., in press.</p>
The global distribution of plants used by humans datasets: list of utilised species, occurrence data and model outputs at 10 arc-minutes spatial resolution
<p>Datasets and model outputs used to map the global distribution of utilised plants by humans. The folder is composed of two subfolders <em>raw_data</em> and <em>processed_data</em> containing respectively the list of utilised plant species modelled -<em>utilised_plants_species_list.csv</em>-, and their occurrence data -<em>occurrence_data.zip-</em> and predicted distribution -<em>species_proba_per_cell.rds-.</em></p> <p> </p> <ul> <li>The file <em>utilised_plants_species_list.csv</em> in the <em>raw_data</em> folder contains a<strong> </strong>list of 35687 plant species (and hybrids) used by humans and 10 plant use categories with the following 14 fields:</li> </ul> <p><strong>plant_ID:<em> </em></strong>plant identifier number ranging from between 1-35687</p> <p><strong>binomial_acc_name:</strong> binomial accepted name of the plant species</p> <p><strong>author_acc_name</strong>: name of the author(s)</p> <p><strong>is_hybrid:</strong> logical TRUE or FALSE indicating whether the species is an hybrid or not.</p> <p><strong>AnimalFood:</strong> forage and fodder for vertebrate animals only.</p> <p><strong>EnvironmentalUses:</strong> examples include intercrops and nurse crops, ornamentals, barrier hedges, shade plants, windbreaks, soil improvers, plants for revegetation and erosion control, wastewater purifiers, indicators of the presence of metals, pollution, or underground water.</p> <p><strong>Fuels:</strong> charcoal, petroleum substitutes, fuel alcohols, etc. Given the importance of energy plants for people, those were distinguished from Materials.</p> <p><strong>GeneSources:</strong> wild relatives of major crops which may possess traits associated with biotic or abiotic resistance and may be valuable for breeding programs.</p> <p><strong>HumanFood:</strong> food for humans only, including beverages and food additives.</p> <p><strong>InvertebrateFood:</strong> plants consumed by invertebrates used by humans, such as bees, silkworms, lac insects and edible grubs.</p> <p><strong>Materials:</strong> woods, fibers, cork, cane, tannins, latex, resins, gums, waxes, oils, lipids, etc. and their derived products.</p> <p><strong>Medicines:</strong> both human and veterinary.</p> <p><strong>Poisons:</strong> plants which are poisonous to both vertebrates and invertebrates, both accidentally and intentionally, e.g., for hunting and fishing, molluscicides, herbicides, insecticides.</p> <p><strong>SocialsUses:</strong> plants used for social purposes, which cannot be defined as food or medicine, for instance, masticatories, smoking materials, narcotics, hallucinogens and psychoactive drugs, and plants with ritual or religious significance.</p> <p><strong>Totals:</strong> total number of uses recorded for a species</p> <p> </p> <ul> <li>The zipfile <em>occurrence_data.zip</em> in the <em>processed_data</em> folder contains 35687 Comma Separated Values (CSV) files, one for each species, containing curated geographic occurrence records used to build species distribution models with the following 14 fields:</li> </ul> <p><strong>Species:</strong> the binomial accepted name of the species</p> <p><strong>Fullname:</strong> same as species</p> <p><strong>decimalLongitude:</strong> the geographic longitude of the occurrence records of the species in decimal degrees</p> <p><strong>decimalLatitude:</strong> the geographic latitude of the occurrence records of the species in decimal degrees</p> <p><strong>countryCode:</strong> a three-letter standard abbreviation for the country of the occurrence locality</p> <p><strong>coordinateUncertaintyinMeters</strong>: indicator for the accuracy of the coordinate location, described as the radius of a circle around the stated point location</p> <p><strong>year:</strong> year of the observation of the occurrence record of the species</p> <p><strong>individualCount:</strong> the number of individuals present at the time of the observation</p> <p><strong>gbifID:</strong> unique identifier number for the occurrence from the original database</p> <p><strong>basisOfRecords:</strong> the type of the individual record, e.g. observation, physical specimen, fossil, living ex-situ, culture collection specimen</p> <p><strong>institutionCode</strong>: the name of the institution or organization listed as the data publisher on GBIF</p> <p><strong>establishmentMeans:</strong> statement about whether an organism has been introduced to a given place and time through the direct or indirect activity of modern humans</p> <p><strong>is_cultivated_observation:</strong> whether or not an organism is cultivated</p> <p><strong>sourceID:</strong> name of the source database</p> <p> </p> <ul> <li>The file <em>species_proba_per_cell.rds</em> in the <em>processed_data</em> folder is<em> a R Data Serialization </em>(RDS) file containing a data.table object with the following 3 fields:</li> </ul> <p><strong>plant_ID:</strong><em> </em>plant identifier number ranging from between 1-35687</p> <p><strong>proba:</strong> species occurrence probability</p> <p><strong>cell:</strong><em> </em>raster grid cell number between 1-2251762</p> <p>This object can be used in combination with a raster layer to reconstruct the modelled distribution of each species or retrieve species richness and endemism.</p>
Spatial partitioning of terrestrial precipitation and corresponding dataset agreement
<p>The study of the water cycle at planetary scale is crucial for our understanding of large-scale climatic processes. However, very little is known about how terrestrial precipitation is distributed across different environments. In this study, we address this gap by employing a 17-dataset ensemble to provide, for the first time, precipitation estimates over a suite of land cover types, biomes, elevation zones, and precipitation intensity classes. We estimate annual terrestrial precipitation at approximately 114,000 ± 9,400 km3, with about 70% falling over tropical, subtropical and temperate regions. Our results highlight substantial inconsistencies, mainly, over the arid and the mountainous areas. To quantify the overall discrepancies, we utilize the concept of dataset agreement and then explore the pairwise relationships among the datasets in terms of "genealogy", concurrency, and distance. The resulting uncertainty-based partitioning demonstrates how precipitation is distributed over a wide range of environments and improves our understanding on how their conditions influence observational fidelity.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.