Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8,038

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

8,038 results for “Validation”

Learn how ShareScore rates datasets ↗
edi60/100

Tree Inventories for Validating Terrestrial Lidar Measurements at Harvard Forest 2007-2014

Our objective is to improve the measurements of canopy structure and biomass of a forest stand and detect their annual changes via a ground-based laser scanning technology, also known as terrestrial lidar (TLS). A TLS instrument utilizes lasers to scan an environment, measure 3D locations of objects encountered by lasers and detect intensities of laser lights scattered by those objects back to the TLS instrument. TLS have shown abilities and is being further explored to retrieve stem diameter, stem count density, stand height, leaf area index, foliage profile, foliage area volume density, aboveground biomass and other useful forest structural parameters rapidly and accurately. Three TLS instruments used in this project include: (1) the Echidna (R) Validation Instrument (EVI), built by CSIRO Australia; (2) Dual-Wavelength Echidna® Lidar (DWEL), built by Boston University, University of Massachusetts, Lowell, University of Massachusetts, Boston and CSIRO Australia; (3) Compact Biomass Lidar (CBL), built by University of Massachusetts, Boston. To validate the forest structural parameters retrieved using these TLS instruments, we set up a one-ha (100 m by 100 m) forest site and collected tree inventory data including: tree location, tree species, DBH, tree height and crown dimension since 2007 with a two-year gap of 2008 and 2009. Lidar data are available from the ORNL DAAC (http://dx.doi.org/10.3334/ORNLDAAC/1045).

openCC0Dec 2023View details →
zenodo52/100

Sample of HYPERNETS Hyperspectral Surface Reflectance Measurements for Satellite Validation from the Gobabeb Site in Namibia

<p>The HYPERNETS project (www.hypernets.eu; Ruddick et al. 2024) has the overall aim to ensure that high quality in situ measurements are available to support the (VNIR/SWIR) optical satellite products. Therefore, it established a new autonomous hyperspectral spectroradiometer (HYPSTAR&reg; - www.hypstar.eu; Kuusk et al. 2024) dedicated to land and water surface reflectance validation with instrument pointing capabilities. This instrument has been deployed over various sites covering a range of water and land types and a range of climatic and logistic conditions. Here, we provide the first fully quality-checked data for the Gobabeb HYPERNETS site in Namibia (GHNA). The HYPERNETS data products were processed using the HYPERNETS_processor (De Vis et al. 2024b).</p> <p>The provided&nbsp;NetCDF files are the L2B hypernets products with surface reflectances, their associated uncertainties and error-correlation information. The reflectance in these products is&nbsp;the Hemispherical-conical Reflectance Factor (HCRF) defined as: HCRF = &pi; L / E where L is the conical upwelling radiance (with field of view of 5&nbsp;degrees) and E is the (hemispherical)&nbsp;downwelling irradiance (i.e. including both direct solar and diffuse sky irradiance). These reflectances have dimensions of wavelength and series, where each series is a set of measurements for a given geometry (combination of viewing zenith and azimuth angle). In addition to variables for&nbsp;wavelength and bandwidth, the files also contain variables that provide for each series the acquisition time, viewing and solar angles, number of valid scans used, and quality flags (typically no flags are set in the data provided in this dataset).&nbsp;These NetCDF files also contain further relevant metadata as attributes. See&nbsp;https://hypernets-processor.readthedocs.io/ for further info.</p> <p>The GHNA site has minimal daily variation in surface cover and weather conditions and is an ideal location for sustained, homogeneous measurements. The site is well characterised as it is very close to an instrument already recognised as a radiometric calibration site (GONA) as part of the RadCalNet network (Bialek et al. 2016).&nbsp;The HYPERNETS site itself&nbsp;(23.60153 degrees&nbsp;S, 15.12589 degrees E)&nbsp;is 650 m from the RadCalNet site, and is located on a gravel plain near a dry riverbed which separates it from the neighbouring dune sea. The HYPSTAR&reg;-XR sensor was installed May 2022 at the top of a 9m mast on an extended 1&nbsp;m horizontal boom to minimise interruption of the field of view. Data are collected every 30 minutes between 9am and 6pm local time (UTC+02) between viewing zenith angles of 0 and 60 degrees. No measurements are taken at 2pm and 2:30pm local time to avoid the hottest part of the day.</p> <p>The HYPSTAR&reg;-XR (eXtended Range) instruments deployed at each land HYPERNETS site consist of&nbsp;a VNIR and a SWIR sensor and autonomously collect data between 380-1700 nm at various viewing&nbsp;geometries and send it to a central server for quality control and processing. The VNIR sensor spans&nbsp;1330 channels between 380 and 1000 nm with a FWHM of 3 nm and the SWIR sensor has 220 channels&nbsp;between 1000 and 1700 nm with a FWHM of 10 nm. The hypernets_processor (De Vis et al.&nbsp;2024b)&nbsp;automatically processes all this data into various products, including the&nbsp;L2A surface&nbsp;reflectance product provided here. All of the products have associated uncertainties (divided into random and systematic uncertainties, including error-correlation information) which were propagated using the CoMet toolkit (www.comet-toolkit.org).&nbsp;For an example using these data for satellite vicarious calibration, see De Vis et al. 2024a).</p> <p>To obtain this dataset, we start from the full GHNA data record and omit data that does not pass the relevant quality checks (QC). Some QC are performed during the near-real time processing done by the hypernets_processor (see https://hypernets-processor.readthedocs.io/en/latest/content/atbd/processing/quality_checks.html) to produce the L2A files. Then, a number of site-specific QC are performed as post-processing to produce the L2B files. These site-specific QC cover things such as removing flags in the L2A data, avoiding periods with bad deployment conditions, removing unsuitable viewing and solar angles, as well as poorly performing wavelength ranges and individual sequences. Any potential misalignment of the sensor is also corrected, affecting L1D irradiances, and L2B reflectances. These corrected data are then used in a more stringent clear sky check, and in a check that verifies the reflectances are within realistic ranges for a given angle and time of year for the given site.&nbsp;</p> <p>There was a rain event in Gobabeb in March 2025, resulting in the growth of grass at the site. We expect the site will be back to its normal surface cover in the near future. Since the rain event, less data passed the site-specific QC. A dedicated QC will be developed for this period, as the data with grass surface cover will still be useful for satellite validation. These updated data will be made available in the future.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

WILLOW - Norther: data set for the full-scale validation of model-based virtual sensing methods for an operational offshore wind turbine

<h1><em><strong>1. General description&nbsp;</strong></em></h1> <p>This data set contains as-build design information, as well as full-scale vibration response measurements from an operational offshore wind-turbine. The turbine is part of the Norther wind farm which is located in the Belgian North Sea<em> </em>and includes a total of 44 Vestas V164 (8.4MW) wind turbines on monopile foundations, see <a href="../api/records/11093262/draft/files/Fig1_Norther_locaction.png/content" target="_blank" rel="noopener noreferrer">Fig1_Norther_locaction.png</a>. This data set is intended to verify and validate model-based virtual sensing algorithms, using data as well as modeling information from a real turbine.&nbsp;</p> <h2><em><strong>1.1 Summary of the shared structural information</strong></em></h2> <p>The included information entails a detailed description of the geometric properties of the monopile and transition piece, distributed and lumped structural masses&nbsp;. All information shared in this record is conform the as-designed documentation.&nbsp;An example of the lumped masses considered in the model input files is presented in "<a href="../api/records/11093262/draft/files/Fig2_Sensor_Network.png/content" target="_blank" rel="noopener">Fig2_Sensor_Network.png"</a></p> <h2><em><strong>1.2 Summary of the shared geotechnical information</strong></em></h2> <p>Monopiles are distinguished by the significant role of soil-structure interaction. Ground reaction is most typically included in the structural model as non-linear p-y curves. Different p-y curves are available for a certain number of soils in the standards applicable to offshore structures (API RP 2GEO, 2011, and ISO 19901-4:2016(E), 2016).</p> <p>The required soil properties to define p-y curves according to the API framework are given in the soil profile provided in a separate Excel. Rather than symbols, the name of the soil properties is generally used as column header (e.g.,&nbsp;<em>Undrained shear strength</em>). Therefore, it is straightforward to identify each soil parameter. The only soil parameter that might lead to confusion is:</p> <ul> <li><em>"epsilon50 [-]"&nbsp;</em>represents&nbsp;the vertical strain at half the maximum principal stress difference in a static undrained triaxial compression test on an undisturbed soil sample.</li> </ul> <p>It's worthy to note that estimates for the small shear strain stiffness, referred to as Gmax, are also included. Despite not being required as an input to define the API p-y curves, this parameter remains a key input for other soil reaction frameworks than the API (e.g., PISA).&nbsp;</p> <h2><em><strong>1.3 Summary of the shared measurement data</strong></em></h2> <p>Two sets of measurement data have been curated for validation purposes; the first interval has been collected during parked conditions, whereas the second interval has been collected during rated operational conditions. Both records have a length of 2 hours, and are subdivided into 10-minute data sets. Furthermore 1Hz SCADA data has been made available for the selected intervals. All different data sources are time synchronized and have been subjected to several internal quality checks.&nbsp;</p> <p>The sensor network on NRT-WTG is illustrated in in <strong>Fig. 2, </strong>whereas a description of the sensor types is presented in&nbsp;<strong>Tab.1.</strong> The acceleration sensors are installed in the horizontal plane, and measure tangential (Y) and orthogonal (X) to the wall, where the positive Y direction is pointing clockwise and the positive X direction is pointing inwards. All strain sensors are installed vertically and are located on the inside of the wall.</p> <table> <tbody> <tr> <td><strong>Data type&nbsp;</strong></td> <td><strong>Sensor type</strong></td> <td><strong>Fs (Hz)</strong></td> <td> <p><strong>Level mLAT (m)</strong></p> </td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>Acceleration (g)&nbsp;&nbsp;</td> <td>Piezo-electric acc. sensor (<strong>ACC</strong>)</td> <td>30</td> <td>15, 69, 97&nbsp;</td> <td>3 Bi-directional accelerometers at different levels. LAT 15 installed at 240 degree heading; LAT 69 and 97 at 60 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Resistive strain gauge (<strong>SG</strong>)</td> <td>30</td> <td>14</td> <td>6 SGs: equally spaced around the inner circumference of the can. Headings: 50, 110, 170, 230, 290, 350 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Fiber-Bragg Grating strain gauge (<strong>FBG</strong>)</td> <td>100</td> <td>-17, -19</td> <td>2 FBGs per level at 165 and 255 degree respectively.</td> </tr> </tbody> </table> <p><strong>Table 1. Description of sensor types.</strong></p> <p>The FBG strain time series have been synchronized with the SG time series using using a cross-correlation based approach. Therefore the SG data has been used to genereate refrence strain time series at the headings of the FBG sensors; the FBG data is subsequently synchronized with regard to this reference time series. No synchronization of the acceleration data was needed, since these are collected using the same data aquisition system as the SG data.&nbsp;</p> <p>The SG strain time series have been calibrated and temperature compensated, whereas this is not the case for the FBG strain time series. The latter have a yet to be determined calibration offset.&nbsp;&nbsp;</p> <p>In conjunction to the sensor channels presented in <strong>Tab. 1</strong>, 1 Hz SCADA data is provided. A summary of the provided SCADA parameters, all sampled at 1Hz, is presented in <strong>Tab 2.</strong></p> <table> <tbody> <tr> <td><strong>Parameter</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Wind speed</td> <td>m/s</td> <td>Wind speed as recorded in the turbine SCADA</td> </tr> <tr> <td>Wind direction</td> <td>&deg;</td> <td>Wind direction relative to North (0&deg;) as recorded in the turbine SCADA</td> </tr> <tr> <td>Yaw angle</td> <td>&deg;</td> <td>Yaw orientation of the nacelle relative to North (0&deg;) as recorded in the turbine SCADA</td> </tr> <tr> <td>Pitch angle</td> <td>&deg;</td> <td>Rotor blade pitch as recorded in the turbine SCADA</td> </tr> <tr> <td>Rotor speed</td> <td>rpm</td> <td>Rotor speed in rotations per minute as recorded in the turbine SCADA</td> </tr> <tr> <td>Power</td> <td>kW</td> <td>Active power of the turbine&nbsp;as recorded in the turbine SCADA</td> </tr> </tbody> </table> <p><strong>Table 2. </strong>List of provided SCADA parameters</p> <p>&nbsp;</p> <p>A summary of the selected intervals and relevant corresponding scada parameters is given in&nbsp;<strong>Tab 3</strong>.</p> <table> <tbody> <tr> <td><strong>Scenario&nbsp;</strong></td> <td><strong>T1 (UTC)</strong></td> <td><strong>T2 (UTC)&nbsp;</strong></td> <td><strong>Windspeed</strong></td> <td><strong>RPM&nbsp;</strong></td> <td><strong>Pitch&nbsp;</strong></td> </tr> <tr> <td>Parked</td> <td> <p>03/07&nbsp; 01:30</p> </td> <td> <p>03/07&nbsp;03:30</p> </td> <td>&lt; 4.5 m/s</td> <td>~1</td> <td>~18 &deg;</td> </tr> <tr> <td>Rated</td> <td> <p>05/07 22:30</p> </td> <td> <p>06/07 00:30&nbsp;</p> </td> <td>~15 m/s</td> <td>10.5</td> <td>8.1&deg;</td> </tr> </tbody> </table> <p><strong>Table 3. </strong>Selected data intervals and relevant scada parameters</p> <p>&nbsp;</p> <h1><em><strong>2. Included in this version&nbsp;</strong></em></h1> <h2><em><strong>2.1 Version - 0.1.0</strong></em></h2> <ul> <li>Relevant Design information can be found in: <ul> <li>Geometry data for NRT-WTG: "WILLOW-Geometry_v4.xlsx"</li> <li>Best estimate soil profile: "WILLOW-BE_soil_profile.xlsx"</li> </ul> </li> <li>Acceleration, strain and scada data can be found in the following parquet files: <ul> <li>Measurement data for the parked case: "NRT-WTG_Parked.parquet.gz"</li> <li>Measurement data for the rated case: "NRT-WTG_Rated.parquet.gz"</li> </ul> </li> </ul> <p>&nbsp;</p> <h1><em><strong>3. Importing parquet files&nbsp; &nbsp;</strong></em></h1> <p>To import the measurement data into Python it is recommended to use pandas:</p> <pre>import pandas as pd<br># Read Parquet file with Pandas: relative_file_path = '<a href="../api/records/11093262/draft/files/NRT-WTG_Parked.parquet.gz/content" target="_blank" rel="noopener noreferrer">NRT-WTG_Parked.parquet.gz</a>' data = pd.read_parquet(relative_file_path ) <br><br>Once the dataframe has been imported, the users can process/re-arrange the raw data according the their needs; it should be noted that the imported dataframe contains NAN values - these are caused by the different sampling rates of the provided signals. </pre>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Data associated with the following publication: Developing the Playground Play Value and Usability Audit Tool (PVUA): An Evaluation of Content Validity via an Expert Panel

<p>This data set contains the supporting data associated with the following publication:</p> <p>Morgenthaler, T., Loebach, J., Lynch, H., Pentland, D., Kottorp, A., &amp; Schulze, C. (in press). Developing the Playground Play Value and Usability Audit Tool (PVUA): An Evaluation of Content Validity via an Expert Panel. Children, Youth and Environments. [DOI was not yet available when the data set was published]</p> <p>The data set includes the following files:</p> <ul> <li>read me file [contains all relevant information to understand and reuse this data set]&nbsp;</li> <li>13 additional files [for description, see read me file]</li> </ul> <p>For more information, please contact the lead researcher, Thomas Morgenthaler (tom.morgenthaler@gmail.com or 121101888@umail.ucc.ie)</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo52/100

OpenFOAM cases of the paper "Development and validation of an open-source CFD model for the efficiency assessment of data centers"

<p>This dataset contains the<em>&nbsp;underling data</em>&nbsp;for the paper &quot;Development and validation of an open-source CFD model for the efficiency assessment of data centers&rdquo;, submitted&nbsp;for the consideration and open review in Open Research Europe (ORE).</p> <p><strong>Validation1.tar.xz:</strong> OpenFOAM files and scripts for the simulation of flow and thermal structures in an enclosed environment (Wang and Chen, 2009).</p> <p><em>Wang, Miao; Chen, Qingyan (2009). Assessment of Various Turbulence Models for Transitional Flows in an Enclosed Environment (RP-1271). HVAC&amp;R Research, 15(6), 1099&ndash;1119. doi:10.1080/10789669.2009.10390881</em></p> <p><strong>Validation2-kOmegaSSTModel.tar.xz:</strong>&nbsp;OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using k-omega SST turbulence model.&nbsp;</p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai &amp; Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2&mdash;Comparison with Experimental Data from Literature, HVAC&amp;R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation2-RNGkEpsilonModel.tar.xz:</strong>&nbsp;OpenFOAM files and scripts for the simulation of forced convection in a room (Zhang et al. 2007) using RNG k-epsilon turbulence model.&nbsp;</p> <p><em>Zhao Zhang, Wei Zhang, Zhiqiang John Zhai &amp; Qingyan Yan Chen (2007) Evaluation of Various Turbulence Models in Predicting Airflow and Turbulence in Enclosed Environments by CFD: Part 2&mdash;Comparison with Experimental Data from Literature, HVAC&amp;R Research, 13:6, 871-886, DOI: 10.1080/10789669.2007.10391460</em></p> <p><strong>Validation3.tar.xz:</strong>&nbsp;OpenFOAM files and scripts for the simulation of strong natural convection in a model fire room (Murakami et al. 1995).</p> <p><em>Murakami, S., S. Kato, and R. Yoshie. 1995. Measurement of turbulence statistics in a model fire room by LDV. ASHRAE Transactions 101(2):287&ndash;301.</em></p> <p><strong>Validation4.tar.xz:</strong> OpenFOAM files and scripts for the simulation of thermal distribution in an open-aisle data center (Abdelmaksoud et al. 2013).</p> <p><em>W.A. Abdelmaksoud, T.Q. Dang, H. Ezzat Khalifa, R.R. Schmidt Improved computational fluid dynamics model for open-aisle air-cooled data center simulations J. Electron. Packag., 135 (2013), pp. 030901-30913</em></p> <p><strong>Results_Validation1.tar.xz:</strong> Simulation results of the Validation case 1.</p> <p><strong>Results_Validation2.tar.xz:</strong> Simulation results of the Validation case 2.</p> <p><strong>Results_Validation3.tar.xz:</strong> Simulation results of the Validation case 3.</p> <p><strong>Results_Validation4.tar.xz:</strong> Simulation results of the Validation case 4.</p> <p><strong>layout.csv:</strong> Input file for the Validation case 4.</p>

opencc-by-4.0Feb 2022View details →
zenodo52/100

TESS-validated Gaia DR3 Pulsating Variables of δ Scuti and γ Doradus: II. 360+ Eclipsing Binaries with δ Scuti and γ Doradus Components

<div> <div> <div> <p>I present serendipitous discoveries of 380 eclipsing binaries with &delta; Scuti and &gamma; Doradus pulsators, 46 eclipsing binaries exhibiting rotational variability, and 8 new RR Lyrae stars, &nbsp;identified for the first time during a validation project of pulsating variables from Gaia Data Release 3. Gaia DR3 Part 4 Variability released 12.4 million variables, including 748,058 pulsating variable stars of `DSCT|GDOR|SXPHE' types among the variability classification results of all classifiers -- 9,976,881 objects (in the file vclassre.dat, https://cdsarc.cds.unistra.fr/viz-bin/cat/I/358}). Among 75,369 analyzed stars,&nbsp; I confirmed 12,145 &delta; Scuti stars (including 8,710 new) and 8,192 &gamma; Doradus stars (including 7,531 new). This work has significantly expanded the bona fide DSCT and GDOR catalogs to include 98,968 and 19,466 stars, respectively, providing a valuable resource for future studies. The discovery of the remarkable number of pulsating binaries underscores the significance of this project in validating Gaia&rsquo;s variable star catalog.&nbsp;</p> </div> </div> </div> <p>The attached CSV files report the current validation results. If you use any data from the catalogs in your research, I appreciate your citation to the paper:&nbsp;</p> <p><strong>Zhou, A.-Y., 2024, Research Notes of the AAS, Volume 8, Number 4, 110 (ADS bibcode: 2024RNAAS...8..110Z)&nbsp;</strong></p> <ul> <li>CSV file GaiaDR3_vari_DSCTgDorSXPhe_Validated_R2_NewEB_Pul.csv for the newly identified Eclipsing Binaries with Pulsating components;</li> <li>CSV file&nbsp;GaiaDR3_vari_DSCTgDorSXPhe_Validated_R2.csv for the entire validated and newly identified results from 75,369 analyzed samples.</li> </ul> <p>This is a developing story. Check back for updates.</p>

opencc-by-4.0Sep 2024View details →
zenodo52/100

Initial Sample of HYPERNETS Hyperspectral Surface Reflectance Measurements for Satellite Validation from the Bare soil at Marquardt, Germany

<p>The HYPERNETS&nbsp;project (www.hypernets.eu) aims to ensure that high-quality in situ measurements are available to support the (VNIR/SWIR) optical Copernicus products. Therefore, it established a new autonomous&nbsp;hyperspectral spectroradiometer (HYPSTAR&reg; - www.hypstar.eu) dedicated to land and water surface reflectance validation&nbsp;with instrument-pointing capabilities.&nbsp;In the prototype phase, the instrument is being deployed at 24 sites covering a range of water and land types and a range of climatic and logistic conditions. This dataset provides the first published data for the ATB HYPERNETS site in Marquardt, Germany [52&deg;27&#39;59.40&quot;N, 12&deg;57&#39;35.16&quot;E] (ATGE). It is a subset of the complete data record, consisting of the measurements which could be used&nbsp;for satellite validation.&nbsp;</p> <p>The provided&nbsp;NetCDF files are the L2A hypernets products with surface reflectances, their associated uncertainties and error-correlation information. The reflectance in the L2A products is&nbsp;the Hemispherical-directional Reflectance Factor (HDRF) defined as HDRF = &pi; L / E where L is the directional upwelling radiance (with the field o, view of 5&nbsp;dgrees), and E is the (hemispherical)&nbsp;downwelling irradiance (i.e. including both direct solar and diffuse sky irradiance). These reflectances have dimensions of wavelength and series, where each series is a set of measurements for a given geometry (combination of viewing zenith and azimuth angle). In addition to variables for&nbsp;wavelength and bandwidth, the files also contain variables that provide for each series the acquisition time, viewing and solar angles, number of valid scans used, and quality flags (typically, no flags are set in the data provided in this dataset).&nbsp;These NetCDF files also contain further relevant metadata as attributes. See&nbsp;https://hypernets-processor.readthedocs.io/ for further info.</p> <p>The HYPSTAR&reg;-XR sensor was installed on 11 Oct 2022 at the top of a 5m mast on an extended 5 m horizontal boom to minimise interruption of the field of view.&nbsp;The boom faces South at the right angle towards bare soil. The mast is located at 52.466778&deg;N, 12.959778&deg;E. Data are collected every 30 minutes between 9:00 and 17:00 (UTC) from different zenith and azimuth angle.</p> <p>The HYPSTAR&reg;-XR (eXtended Range) instruments deployed at each land HYPERNETS site consist of&nbsp;a VNIR and a SWIR sensor and autonomously collect data between 380-1700 nm at various viewing&nbsp;geometries and send it to a central server for quality control and processing. The VNIR sensor spans&nbsp;1330 channels between 380 and 1000 nm with a FWHM of 3 nm, and the SWIR sensor has 220 channels&nbsp;between 1000 and 1700 nm with a FWHM of 10 nm. The hypernets_processor (Goyens et al. 2021; De Vis et al.&nbsp;in prep.)&nbsp;automatically processes all this data into various products, including the&nbsp;L2A surface&nbsp;reflectance product provided here. All products have associated uncertainties (divided into random and systematic uncertainties, including error-correlation information) which were propagated using the CoMet toolkit (www.comet-toolkit.org).&nbsp;</p> <p>To obtain this dataset, we start&nbsp;from the full ATGE data record and omit&nbsp;all the data that do not pass all quality checks performed as part of the hypernets_processor. In addition, an additional screening procedure was also developed to remove outliers and supply the best quality data suitable for satellite validation. To remove the outliers, a sigma-clipping method is used. First, reflectances are extracted in separate 2-hour windows throughout the day (to account for BRDF differences due to different solar positions) for four different wavelengths (500, 900, 1100 and 1600 nm).&nbsp;Outliers in these reflectances are then identified by iteratively calculating the mean reflectance trend&nbsp;with time&nbsp;(by binning the data per maximum of 30 data points), calculating the standard deviation from this trend, and masking any data that is more than three standard deviations away from the trend. This process is repeated on the unmasked data until the standard deviation does not vary by more than 5% between two iterations. The masks for the four different wavelengths&nbsp;are then combined (keeping only measurements for which none of the four wavelengths is an outlier). The reflectances and associated uncertainties for any masked series (i.e. a geometry that is masked either by the sigma-clipping procedure or from the masks of the hypernets_processor) are replaced by NaNs. Any sequence that has more than half of its series masked is removed entirely.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

SALLO validation experiment

<p>dataset containing psychophysical raw data and psychometric curves&#39; points of subjective equality (PSE) obtained in the&nbsp;left-right discrimination&nbsp;and in the bisection tasks, repeatedly performed in the visual and in the acoustic domains, with the head turned at 45&deg; left (-45&deg;), center (0) and 45&deg; right (+45). The clean dataset also contains the values of guess rate and lapse rate used to fit each psychometric curve.</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Dataset and Scripts for: RefPlantNLR: a comprehensive collection of experimentally validated plant NLRs (v.20200528_415)

<p><strong>RefPlantNLR v.20200528_415</strong></p> <p><strong>See&nbsp;</strong>bioRxiv&nbsp;2020.07.08.193961;&nbsp;doi:&nbsp;<a href="https://doi.org/10.1101/2020.07.08.193961">https://doi.org/10.1101/2020.07.08.193961</a></p> <p>SUPPLEMENTAL DATA</p> <p>Table S1: Description of RefPlantNLR.</p> <p>Table S2: Plant orders represented in RefPlantNLR.</p> <p>Supplemental dataset 1: Amino acid sequences of RefPlantNLR entries (fasta format). This file contains 415 amino acid sequences.</p> <p>Supplemental dataset 2: CDS sequences of RefPlantNLR entries (fasta format). This file contains 400 CDS sequences. CDS sequences could not be retrieved for 15 RefPlantNLR entries.</p> <p>Supplemental dataset 3: Annotated genomic sequences of RefPlantNLR entries (GenBank flat file format). This file contains 329 genomic loci containing the gene models of 344 RefPlantNLR entries and 56 RefPlantNLR mRNA entries lacking genomic information.</p> <p>Supplemental dataset 4: InterProScan annotation of the RefPlantNLR amino acid sequences (GFF3 format). This file contains the InterProScan annotation of 415 amino acid sequences.</p> <p>Supplemental dataset 5: InterProScan annotation of the RefPlantNLR CDS sequences (GFF3 format). This file contains the InterProScan annotation of the 400 CDS sequences.</p> <p>Supplemental dataset 6: Amino acid sequences of the extracted RefPlantNLR NB-ARC domains (fasta format). This file contains 424 NB-ARC domain (SUPERFAMILY signature SSF52540) amino acid sequences belonging to 415 RefPlantNLR entries.</p> <p>Supplemental dataset 7: Amino acid sequences of the unique RefPlantNLR extracted NB-ARC domains (fasta format). This file contains 347 unique NB-ARC domain (SUPERFAMILY signature SSF52540) amino acid sequences.</p> <p>Supplemental dataset 8: Clustal Omega alignment of the unique RefPlantNLR extracted NB-ARC domains (PHYLIP format). This file contains the Clustal Omega alignment of 346 unique NB-ARC domains (SUPERFAMILY signature SSF52540) with all positions with less than 95% coverage removed. Pb1 was omitted from this alignment.</p> <p>Supplemental dataset 9: NB-ARC domain phylogeny of the RefPlantNLR entries using the Maximum likelihood method (Newick format). This file contains the phylogenetic analysis of the NB-ARC domain of the RefPlantNLR entries using the JTT method.</p> <p>Supplemental dataset 10: Amino acid sequences of the non-redundant RefPlantNLR entries (fasta format). This file contains 235 amino acid sequences representing the non-redundant RefPlantNLR entries at a 90% amino acid identity threshold per genus according to the NB-ARC domain.</p> <p>Supplemental dataset 11: Amino acid sequences of the NB-ARC domains of the non-redundant RefPlantNLR entries (fasta format). This file contains 241 amino acid sequences representing the extracted NB-ARC domains of the 235 non-redundant RefPlantNLR.</p> <p>Appendix S1: R script used to generate annotations and figures.</p> <p>Appendix S2: InterProScan descriptions used for generating annotations.</p>

opencc-by-4.0Jul 2020View details →
zenodo48/100

Defect Prediction Tool Validation Dataset 2

<p><strong>This dataset is used to address the Research Questions in the study at Transactions on Software Engineering</strong>: <strong>Within-Project</strong> <strong>Defect Prediction of Infrastructure-as-Code using Product and Process Metrics. </strong></p> <p><strong>See also: https://github.com/stefanodallapalma/TSE-2020-05-0217.</strong></p> <p>It provides</p> <p>* <strong>repositories.json</strong> - a list of repositories selected from open-source GitHub repositories based on the Ansible language.</p> <p>* <strong>fixing-commits.json</strong> - a list of defect-fixing commits extracted from those repositories.</p> <p>* <strong>fixed-files.json</strong> - a list of Ansible files fixed in those defect-fixing commits and respective bug-inducing commits.</p> <p>* <strong>failure-prone-files.json</strong> - a list of failure-prone files through the repository&#39;s commit history.</p> <p>* <strong>metrics.zip </strong>- csv files consisting of releases (set of files) and their IaC-oriented, delta and process metrics extracted from each analyzed repository</p> <p>* <strong>projects.zip </strong>- for each analyzed project, it contains the data (models, performance, and results of Recursive Feature Elimination) used to answer the Research Questions.</p> <p><strong>Context</strong></p> <p><em>Infrastructure-as-code&nbsp;(IaC)</em> is the DevOps strategy that allows management and provisioning of infrastructure through the definition of machine-readable files and automation around them, rather than physical hardware configuration or interactive configuration tools.</p> <p>On the one hand, although IaC represents an ever-increasing widely adopted practice nowadays, still little is known concerning how to best maintain, speedily evolve, and continuously improve the code behind the IaC strategy in a measurable fashion.&nbsp;<br> On the other hand, source code measurements are often computed and analyzed to evaluate the different quality aspects of the software developed.<br> In particular, Infrastructure-as-Code is simply &quot;code&quot;, as such it is prone to defects as any other programming languages.</p> <p>This dataset targets the YAML-based Ansible language to devise <strong>within-project defects prediction</strong> approaches for IaC based on Machine-learning.</p> <p><strong>Content</strong></p> <p>The dataset contains metrics extracted from 85 open-source GitHub repositories based on the Ansible language that satisfied the following criteria:</p> <p>* The repository has at least one push event to its master branch in the last six months;<br> * The repository has at least 2 releases;<br> * At least 10% of the files in the repository are IaC scripts;<br> * The repository has at least 2 core contributors;<br> * The repository has evidence of continuous integration&nbsp;practice, such as the presence of a &nbsp;.travis.yaml file;<br> * The repository has a comments ratio&nbsp;of at least 0.1%;<br> * The repository has commit frequency&nbsp;of at least 2 per month on average;<br> * The repository has an issue frequency of at least 0.01 events per month on average;<br> * The repository has evidence of a license, such as the presence of a LICENSE.md file<br> * The repository has at least 100 source lines of code.</p> <p>Metrics are grouped into three categories:</p> <p>* <strong>IaC-Oriented:</strong> metrics of structural properties derived from the source code of infrastructure scripts. Click [here](https://www.sciencedirect.com/science/article/pii/S0164121220301618) for more info.</p> <p>* <strong>Delta</strong>: metrics that capture the amount of change in a file between two successive releases, collected for each IaC-oriented metric.</p> <p>* <strong>Process</strong>: metrics that capture aspects of the development process rather than aspects about the code itself. Description of the process metrics in this dataset can be found [here](https://pydriller.readthedocs.io/en/latest/processmetrics.html).</p> <p>In addition to the metrics, the dataset contains the pre-trained models (*.joblib) in the folders rq1 and rq2 of projects.zip.</p> <p>You can load the model in Python as follows:</p> <p>```<br> from joblib import load<br> model = load(&#39;projects/owner/repository/rq1/random_forest.joblib&#39;), mmap_mode=&#39;r&#39;)</p> <p>best_estimator = model[&#39;estimator&#39;]&nbsp; # The estimator that maximized the AUC-PR</p> <p>cv_results = model[&#39;cv_results&#39;]&nbsp; # The results of each step of the validation procedure</p> <p>best_index = mode[&#39;best_index&#39;]&nbsp; # The index to access the best cv_results<br> ```</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>&nbsp;</p> <p>This work is supported by the European Commission grants no. 825040 (RADON H2020).</p> <p><br> <strong>Inspiration</strong></p> <p>What source code properties and properties about the development process are good predictors of defects in Infrastructure-as-Code scripts?</p>

opencc-by-4.0Jun 2020View details →
zenodo48/100

Wind measurement data from the publication: "Development of a load model validation framework applied to synthetic turbulent wind field evaluation"

<h3>Dataset description:</h3> <p>This datasat represents supplementary material used in the contribution "Development of a load model validation framework applied to<br>synthetic turbulent wind field evaluation" by Meyer, Huhn and Gottschall.</p> <p>Wind measurements from the Testfeld BHV are made available. For installation details, see the mentioned reference.</p> <p>&nbsp;</p> <h3>File description:</h3> <ul> <li>Lidar_HWS.nc - Horizontal wind speed measurements (10 min averages) from a WindCube V2 vertical profiler for one day with a low-level jet occurrence ( <div> <div>2021-04-20)</div> </div> </li> <li>Cups_HWS.nc - Horizontal wind speed measurements (10 min averages) from cup anemometer installed on a met mast for the same day</li> <li>Ensemble_averaged_Spectra.nc - Ensemble averaged spectra for neutral and near neutral situations from a Gill Windmaster at 110m above ground level, used to fit the Mann and KSEC model parameters</li> </ul> <h3>&nbsp;</h3> <h3>Referencing:</h3> <p>When used, please cite like the following:</p> <p>Meyer, Paul J., Matthias L. Huhn, and Julia Gottschall. 2024. "Development of a Load Model Validation Framework Applied to Synthetic Turbulent Wind Field Evaluation"&nbsp;<em>Energies</em> 17, no. 4: 797. https://doi.org/10.3390/en17040797</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Deflections of the s-FKLP model for validation purposes

<p>The test data included are for plate deflections obtained according to the s-FKLP model produced for model validation. The methodology and analysis of the obtained data is included in the Open Access article:</p> <p>Stempin, P.; Pawlak, T. P. &amp; Sumelka, W.<br>Formulation of non-local space-fractional plate model and validation for composite micro-plates <br><em>International Journal of Engineering Science, </em><em>Elsevier BV, </em><strong>2023</strong><em>, 192</em>, 103932.</p> <p>DOI: https://doi.org/10.1016/j.ijengsci.2023.103932</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

ERA5 based training, validation and evaluation data for retrievals combining 22-58 GHz with 175-340 GHz microwave radiometer measurements during MOSAiC

<p>This data set is used for the training, validation and evaluation of retrievals of temperature and specific humidity profiles, as well as integrated water vapour from simlulated or measured microwave brightness temperatures (TBs), which are described in <strong>[1]</strong>.</p> <p>The data set consists of yearly files (2001-2018, 6-hourly resolution) that include data from the European Centre for Medium-Range Weather Forecasts's ERA5 reanalysis <strong>[2]</strong> and simulated TBs in the microwave spectrum. TB simulations were performed with PAMTRA <strong>[3,4]</strong> on the native ERA5 model level resolution at frequencies of a low frequency Humidity and Temperature Profiler (HATPRO, 22-58 GHz) and of a Low Humidity Profiler (LHUMPRO-243-340, aka MiRAC-P, 175-340 GHz). Afterwards, the ERA5 model level data has been interpolated to a new height grid (dimension 'z'), of which the lowest 43 indices equal the height grid of the retrieval that is developed with this data set. The upper 11 indices are included for additional TB simulations needed for the information content estimation performed and are not used for the retrievals to avoid the tropopause.</p> <p>The trained retrieval is applied to observations from the HATPRO and MiRAC-P that were installed onboard the research vessel Polarstern during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition.</p> <p><strong>[1]:</strong> Walbr&ouml;l, A., Griesche, H. J., Mech, M., Crewell, S., and Ebell, K.: Combining low- and high-frequency microwave radiometer measurements from the MOSAiC expedition for enhanced water vapour products, Atmospheric Measurement Techniques, 17, 6223-6245, https://doi.org/10.5194/amt-17-6223-2024, 2024.</p> <p><strong>[2]:</strong> Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Hor&aacute;nyi, A., Mu&ntilde;oz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., H&oacute;lm, E., Janiskov&aacute;, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Th&eacute;paut, J.: The ERA5 global reanalysis, Quarterly Journal of the Royal Meteorological Society, 146, 1999&ndash;2049, https://doi.org/10.1002/qj.3803, 2020.</p> <p><strong>[3]:</strong> Mech, M., Maahn, M., Kneifel, S., Ori, D., Orlandi, E., Kollias, P., Schemann, V., and Crewell, S.: PAMTRA 1.0: the Passive and Active Microwave radiative TRAnsfer tool for simulating radiometer and radar measurements of the cloudy atmosphere, Geoscientific Model Development, 13, 4229&ndash;4251, https://doi.org/10.5194/gmd-13-4229-2020, 2020.</p> <p><strong>[4]:</strong> Mech, M., Maahn, M., Ori, D., Kneifel, S., and Orlandi, E.: PAMTRA Package &ndash; Passive and Active Microwave TRANsfer, available at: https://github.com/igmk/pamtra (last access: 6 September 2020), 2019c.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Galápagos Archipelago Refined Analysis Validation data

<p>This dataset includes (near) surface data variables from the&nbsp;<a href="https://data.klima.tu-berlin.de/GAR/">GAR</a> dataset for the model validation period from 2022-04-01 to 2023-03-31.</p> <p>As this data is part of the GAR dataset, please find additional data at <a href="https://data.klima.tu-berlin.de/GAR/">https://data.klima.tu-berlin.de/GAR/</a></p> <p>The <a href="https://www.unidata.ucar.edu/software/netcdf/">netCDF</a> format is self-describing, so that all needed metadata are included within the files.</p> <p>The file names are composed with the following structure:</p> <p>&lt;model-setup&gt;_&lt;horizontal-resolution&gt;_&lt;time-resolution&gt;_&lt;variable-name&gt;.nc</p> <p>The shorthands in the file names represent the following:</p> <p><strong>MM</strong> = Name of the model setup, described in Schmidt et al. (unpublished)</p> <p><strong>d02km</strong> = domain with a grid spacing of 2 km</p> <p><strong>2d</strong> = spatial dimensions (2d data, single level)</p> <p><strong>3d_press</strong> = spatial dimensions (3d data, pressure level)</p> <p><strong>d</strong> = time frequency of the data (daily)</p> <p><strong>m</strong> = time frequency of the data (monthly)</p> <p><strong>y</strong> = time frequency of the data (yearly)</p> <p><strong>psfc</strong> = surface (sfc) pressure</p> <p><strong>q2</strong> = water vapor mixing ratio (qv) st 2 m</p> <p><strong>q</strong> = mixing ratio</p> <p><strong>prcp</strong> = total precipitation (step-wise)</p> <p><strong>et</strong> = actual evapotranspiration (step-wise)</p> <p><strong>t2</strong> = temperature (temp) at 2 m</p> <p><strong>theta</strong> = potential temperature</p> <p><strong>sh2 </strong>= specific humidity at 2 m</p> <p><strong>rh2 </strong>= relative humidity at 2 m</p> <p><strong>u10</strong> = 10 m u-wind component</p> <p><strong>v10</strong> = 10 m v-wind component</p> <p><strong>ws10</strong> = 10 m wind speed</p> <p><strong>w</strong> = w-wind component</p> <p><strong>wd10</strong> = 10 m wind direction</p> <p><strong>hgt&nbsp;</strong>= surface height</p> <p><strong>landmask </strong>= landmask</p> <p>&nbsp;</p> <p>The data is in accordance with the <a href="https://cfconventions.org/">CF Conventions</a> CF-1.8</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Synthetic cryo electron subtomograms containing biomolecular complexes with continuous conformational variability, used for validating TomoFlow method

<p>Two datasets used for validating TomoFlow method, an optical-flow based approach for analyzing continuous conformational variability of biomolecular complexes in cryo electron subtomograms. The&nbsp;TomoFlow method and the methods used to synthesize the two test datasets have been fully described in the following article: &quot;M. Harastani, M. Eltsov, A. Leforestier, S. Jonic, TomoFlow: Analysis of continuous conformational variability of macromolecules in cryogenic subtomograms based on 3D dense optical flow, Journal of Molecular Biology (2021), doi: https://doi.org/10.1016/j.jmb.2021.167381&quot;. Additionally, this article describes a test of TomoFlow using one experimental cryo electron tomography dataset (available in EMPIAR and EMDB databases under the accession codes EMPIAR-10679 and EMD-12699).&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Dataset for the validation of a Computational Thinking test for upper primary school (grades 3-4)

<p>This dataset contains quantitative student&nbsp;data acquired during the administration of a new computational thinking assessment for upper primary school (grades 3 and 4). Over 1500 students (approximately half in grade 3 and half in grade 4) participated in the data collection which took place in&nbsp;January 2021 in the Canon Vaud in Switzerland. The data was used to validate the psychometric properties of the instrument in the referenced article.&nbsp;</p> <p>&nbsp;</p> <p>If you use any of the resources provided in this repository, please cite the following</p> <p>&bull; The Zenodo repository, DOI:&nbsp;10.5281/zenodo.5865573</p> <p>&bull; The corresponding journal article</p> <p>&bull; Licence : CC-BY-NC</p> <p>&nbsp;</p> <p>In case of inquiries, please contact laila.elhamamsy@epfl.ch</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Data files belonging to the paper "Dealing with clustered samples for assessing map accuracy by cross-validation"

<p>Mapping of environmental variables often relies on map accuracy assessment through cross-validation with the data used for calibrating the underlying mapping model. When the data points are spatially clustered, conventional cross-validation leads to optimistically biased estimates of map accuracy. Several papers have promoted spatial cross-validation as a means to tackle this over-optimism. Many of these papers blame spatial autocorrelation as the cause of the bias and propagate the widespread misconception that spatial proximity of calibration points to validation points invalidates classical statistical validation of maps. In the paper related to these data, we present and evaluate alternative cross-validation approaches for assessing map accuracy from clustered sample data.&nbsp;</p> <p>&nbsp;</p> <p>The study area is western Europe, constrained in the north at 52&deg; latitude&nbsp;and at -10&deg; and 24&deg; longitude The projection is IGNF:ETRS89LAEA (Lambert azimuthal equal area projection).</p> <p>&nbsp;</p> <p><strong>Files:</strong></p> <p>agb.tif&nbsp; = above ground biomass (AGB) map from&nbsp;version 3 of the 2017 CCI-Biomass product (<a href="https://catalogue.ceda.ac.uk/uuid/5f331c418e9f4935b8eb1b836f8a91b8">https://catalogue.ceda.ac.uk/uuid/5f331c418e9f4935b8eb1b836f8a91b8</a>)<br> AGBstack.tif&nbsp; = covariates used for predicting AGB<br> aggArea.tif&nbsp; = coarse&nbsp;grid used for simulation in the model-based methods<br> ocs.tif&nbsp; = soil organic carbon stock (OCS) map (0-30 cm) from&nbsp;Soilgrids (<a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a>)<br> OCSstack.tif&nbsp; = covariates used for predicting OCS<br> strata.xxx&nbsp;= 100 compact geo-strata (ESRI shape) created with the spcosa package; used for generating clustered samples<br> TOTmask.tif&nbsp; = mask of the area covered by the covariates</p> <p>&nbsp;</p> <p><strong>Details and data sources of the covariates in AGBstack.tif and OCSstack.tif:</strong></p> <table> <tbody> <tr> <td> <p><strong>Name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Note</strong></p> </td> </tr> <tr> <td> <p>ai</p> </td> <td> <p>Aridity Index</p> </td> <td> <p><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></p> </td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio1</p> </td> <td> <p>Mean annual air temperature [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio5</p> </td> <td> <p>Mean daily maximum air temperature of the warmest month [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio7</p> </td> <td> <p>Annual range of air temperature [&deg;C]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio12</p> </td> <td> <p>Annual precipitation [kg/m<sup>2</sup>]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>bio15</p> </td> <td> <p>Precipitation seasonality [kg/m<sup>2</sup>]</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>gdd10</p> </td> <td> <p>Growing degree days heat sum above 10&deg;C</p> </td> <td><a href="https://chelsa-climate.org/downloads/">https://chelsa-climate.org/downloads/</a></td> <td>Version 2.1</td> </tr> <tr> <td> <p>clay</p> </td> <td> <p>Clay content [g/kg] of the 0-5cm layer</p> </td> <td> <p><a href="https://soilgrids.org/">https://soilgrids.org/</a></p> <p>&nbsp;</p> </td> <td> <p>Only used for AGB</p> </td> </tr> <tr> <td> <p>sand</p> </td> <td> <p>Sand content [g/kg] of the 0-5cm layer</p> </td> <td><a href="https://soilgrids.org/">https://soilgrids.org/</a></td> <td>as above</td> </tr> <tr> <td> <p>pH</p> </td> <td> <p>Acidity (Ph(water)) of the 0-5cm layer</p> </td> <td><a href="https://soilgrids.org/">https://soilgrids.org/</a></td> <td>as above</td> </tr> <tr> <td> <p>glc2017</p> </td> <td> <p>Landcover 2017</p> </td> <td> <p><a href="https://land.copernicus.eu/global/products/lc">https://land.copernicus.eu/global/products/lc</a>, reclassified&nbsp; to: closed forest, open forest,&nbsp; natural non-forest veg., bare &amp; sparse veg. cropland, built-up, water</p> </td> <td> <p>Categorical variable</p> </td> </tr> <tr> <td> <p>dem</p> </td> <td> <p>Elevation</p> </td> <td> <p><a href="https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-eu-dem">https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-eu-dem</a></p> </td> <td> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>cosasp</p> </td> <td> <p>Cosine of slope aspect</p> </td> <td> <p>Computed with the terra package from elevation</p> </td> <td>Computed @25m resolution; next aggregated to 0.5km</td> </tr> <tr> <td> <p>sinasp</p> </td> <td> <p>Sine of slope aspect</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>slope</p> </td> <td> <p>Slope</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TPI</p> </td> <td> <p>Topographic position index</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TRI</p> </td> <td> <p>Terrain ruggedness index</p> </td> <td>Computed with the terra package from elevation</td> <td>as above</td> </tr> <tr> <td> <p>TWI</p> </td> <td> <p>Topographic wetness index</p> </td> <td> <p>Computed with SAGA from 500m resolution (aggregated) dem</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>gedi</p> </td> <td> <p>Forest height</p> </td> <td> <p><a href="https://glad.umd.edu/dataset/gedi">https://glad.umd.edu/dataset/gedi</a></p> </td> <td> <p>Zone: NAFR</p> </td> </tr> <tr> <td> <p>xcoord</p> </td> <td> <p>X coordinate</p> </td> <td> <p>Using a mask created from the other covariates</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>ycoord</p> </td> <td> <p>Y coordinate</p> </td> <td>Using a mask created from the other covariates</td> <td>&nbsp;</td> </tr> <tr> <td> <p>Dcoast</p> </td> <td> <p>Distance from coast</p> </td> <td> <p>Using a land mask created from the other covariates</p> </td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Characterization of Wind Turbine Wakes with Nacelle-Mounted Doppler LiDARs and Model Validation in the Presence of Wind Veer

<p>Dataset of the paper &quot;Characterization of Wind Turbine Wakes with Nacelle-Mounted Doppler LiDARs and Model Validation in the Presence of Wind Veer&quot; published in Remote Sensing [1].</p> <p>[1] Brugger P, Fuertes FC, Vahidzadeh M, Markfort CD, Port&eacute;-Agel F. Characterization of Wind Turbine Wakes with Nacelle-Mounted Doppler LiDARs and Model Validation in the Presence of Wind Veer. <em>Remote Sensing</em>. 2019; 11(19):2247. https://doi.org/10.3390/rs11192247.</p>

opencc-by-4.0Sep 2019View details →
zenodo48/100

Experimental validation data for in silico OA study

<p>The folder contains, ALP assay and&nbsp;PCR data, experimental protocols as well as scripts for plotting and&nbsp;analysis.</p> <p>This is related to the following git repository:&nbsp;https://github.com/Rapha-L/Experimental_validation_for_insilicoOA&nbsp;</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record