Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,155

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,155 results for “Data Base”

Learn how ShareScore rates datasets ↗
edi60/100

Wrack classification data based on UAV imagery from Dean Creek on Sapelo Island, GA

We used a DJI Matrice 210 UAV with a MicaSense Altum to collect a total of 20 images from January 2020 - December 2021 in a the Dean Creek marsh on Sapelo Island, GA. Wrack was classified using a principal component analysis. Wrack patches under 1 m2 were excluded from analyses. Wrack classifications were converted to polygon and point data where each point represents a 5 cm x 5 cm pixel. Those files were then used to analyze wrack characteristics, their relation to environmental drivers, and landscape based patterns. For both polygon and point data, we used the National Elevation Dataset (https://gdg.sc.egov.usda.gov/Catalog/ProductDescription/NED.html) to determine the elevation of each wrack patch. Creeks and shorelines were digitized and used to determine each wrack patches' distance to water. We calculated the frequency of wrack deposition at each point by adding together the number of images where that pixel was classified as wrack over the course of the study. Polygon data were related to tide height from a NOAA tidal station data product (Ft. Pulaski, Station 8670870; https://tidesandcurrents.noaa.gov) and wind speed and wind direction from the Marsh Landing weather station (downloaded data for the SAPMLMET met station from: https://cdmo.baruch.sc.edu/) to evaluate the relationship of wrack to environmental drivers.

openCC (other)Aug 2023View details →
zenodo56/100

Data from: Satellite-based Lagrangian model reveals how upwelling and oceanic circulation shape krill hotspots in the California Current System [updated]

<p><strong>Abstract</strong></p> <p>In the California Current System, wind-driven nutrient supply and primary production, computed from satellite data, provide a synoptic view of how phytoplankton production is coupled to upwelling. In contrast, linking upwelling to zooplankton populations is difficult due to relatively scarce observations and the inherent patchiness of zooplankton. While phytoplankton respond quickly to environmental forcing, zooplankton grow slower and tend to aggregate into mesoscale &ldquo;hotspot&rdquo; regions spatially decoupled from upwelling centers. To better understand mechanisms controlling the formation of zooplankton hotspots, we use a satellite-based Lagrangian method where variables from a plankton model, forced by wind-driven nutrient supply, are advected by near-surface currents following upwelling events. Modeled zooplankton distribution reproduces published accounts of euphausiid (krill) hotspots, including the location of major hotspots and their interannual variability. This satellite-based modeling tool is used to analyze the variability and drivers of krill hotspots in the California Current System, and to investigate how water masses of different origin and history converge to form predictable biological hotspots. The Lagrangian framework suggests that two conditions are necessary for a hotspot to form: a convergence of coastal water masses, and above average nutrient supply where these water masses originated from. The results highlight the role of upwelling, oceanic circulation, and plankton temporal dynamics in shaping krill mesoscale distribution, seasonal northward propagation, and interannual variability.</p> <p><strong>Data set description</strong></p> <p>This data set includes 2 files:</p> <ul> <li>a satellite-based 1993-2023 monthly retrospective of krill concentrations (Zbig) modeled using the growth-advection method in the California Current upwelling system. Inputs include the nitrate supply product described below and GlobCurrent 15 m oceanic currents. This dataset is updated monthly (using NRT data) at https://www.mbari.org/science/upper-ocean-systems/biological-oceanography/krill-hotspots-in-the-california-current/.</li> <li>a satellite-based 1993-2023 monthly retrospective of wind-driven nitrate supply estimated in a 150 km coastal band at 0.125&deg; latitudinal resolution. Nitrate supply was calculated based primarily on CCMP v3.1 winds, AVISO geostrophic currents, and a climatology of in situ nitrate at 60m. This dataset is updated monthly (using NRT data) at https://www.mbari.org/science/upper-ocean-systems/biological-oceanography/nitrate-supply-estimates-in-upwelling-systems/.</li> </ul> <p>See details regarding data sources and calculations in&nbsp;<a href="https://doi.org/10.3389/fmars.2022.835813">Messi&eacute; et al. (2022)</a>.</p> <p>[IMPORTANT NOTE:] There is an error in the Ekman pumping fields (trans_pump, Nsupply_pump, Nsupply_total) that will be corrected soon (those fields are not used in publications where only coastal transport was considered). Please contact me if you need Ekman pumping fields before this is fixed.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

WILLOW - Norther: data set for the full-scale validation of model-based virtual sensing methods for an operational offshore wind turbine

<h1><em><strong>1. General description&nbsp;</strong></em></h1> <p>This data set contains as-build design information, as well as full-scale vibration response measurements from an operational offshore wind-turbine. The turbine is part of the Norther wind farm which is located in the Belgian North Sea<em> </em>and includes a total of 44 Vestas V164 (8.4MW) wind turbines on monopile foundations, see <a href="../api/records/11093262/draft/files/Fig1_Norther_locaction.png/content" target="_blank" rel="noopener noreferrer">Fig1_Norther_locaction.png</a>. This data set is intended to verify and validate model-based virtual sensing algorithms, using data as well as modeling information from a real turbine.&nbsp;</p> <h2><em><strong>1.1 Summary of the shared structural information</strong></em></h2> <p>The included information entails a detailed description of the geometric properties of the monopile and transition piece, distributed and lumped structural masses&nbsp;. All information shared in this record is conform the as-designed documentation.&nbsp;An example of the lumped masses considered in the model input files is presented in "<a href="../api/records/11093262/draft/files/Fig2_Sensor_Network.png/content" target="_blank" rel="noopener">Fig2_Sensor_Network.png"</a></p> <h2><em><strong>1.2 Summary of the shared geotechnical information</strong></em></h2> <p>Monopiles are distinguished by the significant role of soil-structure interaction. Ground reaction is most typically included in the structural model as non-linear p-y curves. Different p-y curves are available for a certain number of soils in the standards applicable to offshore structures (API RP 2GEO, 2011, and ISO 19901-4:2016(E), 2016).</p> <p>The required soil properties to define p-y curves according to the API framework are given in the soil profile provided in a separate Excel. Rather than symbols, the name of the soil properties is generally used as column header (e.g.,&nbsp;<em>Undrained shear strength</em>). Therefore, it is straightforward to identify each soil parameter. The only soil parameter that might lead to confusion is:</p> <ul> <li><em>"epsilon50 [-]"&nbsp;</em>represents&nbsp;the vertical strain at half the maximum principal stress difference in a static undrained triaxial compression test on an undisturbed soil sample.</li> </ul> <p>It's worthy to note that estimates for the small shear strain stiffness, referred to as Gmax, are also included. Despite not being required as an input to define the API p-y curves, this parameter remains a key input for other soil reaction frameworks than the API (e.g., PISA).&nbsp;</p> <h2><em><strong>1.3 Summary of the shared measurement data</strong></em></h2> <p>Two sets of measurement data have been curated for validation purposes; the first interval has been collected during parked conditions, whereas the second interval has been collected during rated operational conditions. Both records have a length of 2 hours, and are subdivided into 10-minute data sets. Furthermore 1Hz SCADA data has been made available for the selected intervals. All different data sources are time synchronized and have been subjected to several internal quality checks.&nbsp;</p> <p>The sensor network on NRT-WTG is illustrated in in <strong>Fig. 2, </strong>whereas a description of the sensor types is presented in&nbsp;<strong>Tab.1.</strong> The acceleration sensors are installed in the horizontal plane, and measure tangential (Y) and orthogonal (X) to the wall, where the positive Y direction is pointing clockwise and the positive X direction is pointing inwards. All strain sensors are installed vertically and are located on the inside of the wall.</p> <table> <tbody> <tr> <td><strong>Data type&nbsp;</strong></td> <td><strong>Sensor type</strong></td> <td><strong>Fs (Hz)</strong></td> <td> <p><strong>Level mLAT (m)</strong></p> </td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>Acceleration (g)&nbsp;&nbsp;</td> <td>Piezo-electric acc. sensor (<strong>ACC</strong>)</td> <td>30</td> <td>15, 69, 97&nbsp;</td> <td>3 Bi-directional accelerometers at different levels. LAT 15 installed at 240 degree heading; LAT 69 and 97 at 60 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Resistive strain gauge (<strong>SG</strong>)</td> <td>30</td> <td>14</td> <td>6 SGs: equally spaced around the inner circumference of the can. Headings: 50, 110, 170, 230, 290, 350 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Fiber-Bragg Grating strain gauge (<strong>FBG</strong>)</td> <td>100</td> <td>-17, -19</td> <td>2 FBGs per level at 165 and 255 degree respectively.</td> </tr> </tbody> </table> <p><strong>Table 1. Description of sensor types.</strong></p> <p>The FBG strain time series have been synchronized with the SG time series using using a cross-correlation based approach. Therefore the SG data has been used to genereate refrence strain time series at the headings of the FBG sensors; the FBG data is subsequently synchronized with regard to this reference time series. No synchronization of the acceleration data was needed, since these are collected using the same data aquisition system as the SG data.&nbsp;</p> <p>The SG strain time series have been calibrated and temperature compensated, whereas this is not the case for the FBG strain time series. The latter have a yet to be determined calibration offset.&nbsp;&nbsp;</p> <p>In conjunction to the sensor channels presented in <strong>Tab. 1</strong>, 1 Hz SCADA data is provided. A summary of the provided SCADA parameters, all sampled at 1Hz, is presented in <strong>Tab 2.</strong></p> <table> <tbody> <tr> <td><strong>Parameter</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Wind speed</td> <td>m/s</td> <td>Wind speed as recorded in the turbine SCADA</td> </tr> <tr> <td>Wind direction</td> <td>&deg;</td> <td>Wind direction relative to North (0&deg;) as recorded in the turbine SCADA</td> </tr> <tr> <td>Yaw angle</td> <td>&deg;</td> <td>Yaw orientation of the nacelle relative to North (0&deg;) as recorded in the turbine SCADA</td> </tr> <tr> <td>Pitch angle</td> <td>&deg;</td> <td>Rotor blade pitch as recorded in the turbine SCADA</td> </tr> <tr> <td>Rotor speed</td> <td>rpm</td> <td>Rotor speed in rotations per minute as recorded in the turbine SCADA</td> </tr> <tr> <td>Power</td> <td>kW</td> <td>Active power of the turbine&nbsp;as recorded in the turbine SCADA</td> </tr> </tbody> </table> <p><strong>Table 2. </strong>List of provided SCADA parameters</p> <p>&nbsp;</p> <p>A summary of the selected intervals and relevant corresponding scada parameters is given in&nbsp;<strong>Tab 3</strong>.</p> <table> <tbody> <tr> <td><strong>Scenario&nbsp;</strong></td> <td><strong>T1 (UTC)</strong></td> <td><strong>T2 (UTC)&nbsp;</strong></td> <td><strong>Windspeed</strong></td> <td><strong>RPM&nbsp;</strong></td> <td><strong>Pitch&nbsp;</strong></td> </tr> <tr> <td>Parked</td> <td> <p>03/07&nbsp; 01:30</p> </td> <td> <p>03/07&nbsp;03:30</p> </td> <td>&lt; 4.5 m/s</td> <td>~1</td> <td>~18 &deg;</td> </tr> <tr> <td>Rated</td> <td> <p>05/07 22:30</p> </td> <td> <p>06/07 00:30&nbsp;</p> </td> <td>~15 m/s</td> <td>10.5</td> <td>8.1&deg;</td> </tr> </tbody> </table> <p><strong>Table 3. </strong>Selected data intervals and relevant scada parameters</p> <p>&nbsp;</p> <h1><em><strong>2. Included in this version&nbsp;</strong></em></h1> <h2><em><strong>2.1 Version - 0.1.0</strong></em></h2> <ul> <li>Relevant Design information can be found in: <ul> <li>Geometry data for NRT-WTG: "WILLOW-Geometry_v4.xlsx"</li> <li>Best estimate soil profile: "WILLOW-BE_soil_profile.xlsx"</li> </ul> </li> <li>Acceleration, strain and scada data can be found in the following parquet files: <ul> <li>Measurement data for the parked case: "NRT-WTG_Parked.parquet.gz"</li> <li>Measurement data for the rated case: "NRT-WTG_Rated.parquet.gz"</li> </ul> </li> </ul> <p>&nbsp;</p> <h1><em><strong>3. Importing parquet files&nbsp; &nbsp;</strong></em></h1> <p>To import the measurement data into Python it is recommended to use pandas:</p> <pre>import pandas as pd<br># Read Parquet file with Pandas: relative_file_path = '<a href="../api/records/11093262/draft/files/NRT-WTG_Parked.parquet.gz/content" target="_blank" rel="noopener noreferrer">NRT-WTG_Parked.parquet.gz</a>' data = pd.read_parquet(relative_file_path ) <br><br>Once the dataframe has been imported, the users can process/re-arrange the raw data according the their needs; it should be noted that the imported dataframe contains NAN values - these are caused by the different sampling rates of the provided signals. </pre>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Invasion Biology WikiProject Scientific Papers: Text Data Mining and LLM-based Information Extraction of Species, Locations, Habitats, and Ecosystems

<p>This dataset contains the abstract and full-text for publication DOIs from the Invasion Biology WikiProject (DOI:&nbsp;<a href="https://www.doi.org/10.5281/zenodo.12518036">10.5281/zenodo.12518036</a>). The data was retrieved using the <a href="https://ask.orkg.org/">ask.orkg.org</a> <a href="https://api.ask.orkg.org/docs#tag/Semantic-Neural-Search/operation/explore_documents_index_explore_get">API</a>. For the <a href="https://github.com/jd-coderepos/invasion-biology-IE/blob/main/scripts/ask-doi-list-fulltext-search.py">script</a> used to obtain the data, refer to the accompanying GitHub repository: <a href="https://github.com/jd-coderepos/invasion-biology-IE/" target="_blank" rel="noopener">https://github.com/jd-coderepos/invasion-biology-IE/</a>.</p> <p>The resulting CSV file includes the following fields: <code>"ASK ID"</code>, <code>"DOI"</code>, <code>"Title"</code>, <code>"Abstract"</code>, and <code>"Full-text"</code>.</p> <p>Of the 49,438 queried DOIs, the ASK database provided:</p> <ul> <li><strong>Total DOIs processed:</strong> 12,636</li> <li><strong>DOIs with neither abstract nor full-text:</strong> 36 (abstract token count was less than 10)</li> <li><strong>DOIs with abstracts but no full-text:</strong> 12,636</li> <li><strong>DOIs with both abstract and full-text:</strong> 2,834</li> </ul> <p>The second part of the dataset contains structured information extracted from the publications using the GPT-4o Large Language Model. This structured data is included in the zipped folder <code>structured-publications.zip</code>.</p> <p>The accompanying GitHub repository provides access to the code and scripts used at various stages of the information extraction (IE) process.</p> <p><strong>Theme of the Study:</strong><br>"Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models."</p>

opencc-by-4.0Oct 2024View details →
zenodo52/100

ELKI Multi-View Clustering Data Sets Based on the Amsterdam Library of Object Images (ALOI)

<p>These data sets were originally created for the following publications:</p> <p><em>M. E. Houle, H.-P. Kriegel, P. Kr&ouml;ger, E. Schubert, A. Zimek</em><br> <strong>Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?</strong><br> In Proceedings of the 22nd International Conference on Scientific and Statistical Database Management (SSDBM), Heidelberg, Germany, 2010.</p> <p><em>H.-P. Kriegel, E. Schubert, A. Zimek</em><br> <strong>Evaluation of Multiple Clustering Solutions</strong><br> In 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, 2011.</p> <p>The outlier data set versions were introduced in:</p> <p><em>E. Schubert, R. Wojdanowski, A. Zimek, H.-P. Kriegel</em><br> <strong>On Evaluation of Outlier Rankings and Outlier Scores</strong><br> In Proceedings of the 12th SIAM International Conference on Data Mining (SDM), Anaheim, CA, 2012.</p> <p>&nbsp;</p> <p>They are derived from the original image data available at <a href="https://aloi.science.uva.nl/">https://aloi.science.uva.nl/</a></p> <p>The image acquisition process is documented in the original ALOI work: <em>J. M. Geusebroek, G. J. Burghouts, and A. W. M. Smeulders</em>, <strong>The Amsterdam library of object images</strong>, Int. J. Comput. Vision, 61(1), 103-112, January, 2005</p> <p>Additional information is available at: <a href="https://elki-project.github.io/datasets/multi_view">https://elki-project.github.io/datasets/multi_view</a></p> <p>The following views are currently available:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>Object number</td> <td>Sparse 1000 dimensional vectors that give the <em>true</em> object assignment</td> <td><a href="6355684/files/objs.arff.gz">objs.arff.gz</a></td> </tr> <tr> <td>RGB color histograms</td> <td>Standard RGB color histograms (uniform binning)</td> <td><a href="6355684/files/aloi-8d.csv.gz">aloi-8d.csv.gz</a> <a href="6355684/files/aloi-27d.csv.gz">aloi-27d.csv.gz</a> <a href="6355684/files/aloi-64d.csv.gz">aloi-64d.csv.gz</a> <a href="6355684/files/aloi-125d.csv.gz">aloi-125d.csv.gz</a> <a href="6355684/files/aloi-216d.csv.gz">aloi-216d.csv.gz</a> <a href="6355684/files/aloi-343d.csv.gz">aloi-343d.csv.gz</a> <a href="6355684/files/aloi-512d.csv.gz">aloi-512d.csv.gz</a> <a href="6355684/files/aloi-729d.csv.gz">aloi-729d.csv.gz</a> <a href="6355684/files/aloi-1000d.csv.gz">aloi-1000d.csv.gz</a></td> </tr> <tr> <td>HSV color histograms</td> <td>Standard HSV/HSB color histograms in various binnings</td> <td><a href="6355684/files/aloi-hsb-2x2x2.csv.gz">aloi-hsb-2x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-3x3x3.csv.gz">aloi-hsb-3x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-4x4x4.csv.gz">aloi-hsb-4x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-5x5x5.csv.gz">aloi-hsb-5x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-6x6x6.csv.gz">aloi-hsb-6x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-7x7x7.csv.gz">aloi-hsb-7x7x7.csv.gz</a> <a href="6355684/files/aloi-hsb-7x2x2.csv.gz">aloi-hsb-7x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-7x3x3.csv.gz">aloi-hsb-7x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-14x3x3.csv.gz">aloi-hsb-14x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-8x4x4.csv.gz">aloi-hsb-8x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-9x5x5.csv.gz">aloi-hsb-9x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-13x4x4.csv.gz">aloi-hsb-13x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-14x5x5.csv.gz">aloi-hsb-14x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-10x6x6.csv.gz">aloi-hsb-10x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-14x6x6.csv.gz">aloi-hsb-14x6x6.csv.gz</a></td> </tr> <tr> <td>Color similiarity</td> <td>Average similarity to 77 reference colors (not histograms) 18 colors x 2 sat x 2 bri + 5 grey values (incl. white, black)</td> <td><a href="6355684/files/aloi-colorsim77.arff.gz">aloi-colorsim77.arff.gz</a> (feature subsets are meaningful here, as these features are computed independently of each other)</td> </tr> <tr> <td>Haralick features</td> <td>First 13 Haralick features (radius 1 pixel)</td> <td><a href="6355684/files/aloi-haralick-1.csv.gz">aloi-haralick-1.csv.gz</a></td> </tr> <tr> <td>Front to back</td> <td>Vectors representing front face vs. back faces of individual objects</td> <td><a href="6355684/files/front.arff.gz">front.arff.gz</a></td> </tr> <tr> <td>Basic light</td> <td>Vectors indicating basic light situations</td> <td><a href="6355684/files/light.arff.gz">light.arff.gz</a></td> </tr> <tr> <td>Manual annotations</td> <td>Manually annotated object groups of semantically related objects such as cups</td> <td><a href="6355684/files/manual1.arff.gz">manual1.arff.gz</a></td> </tr> </tbody></table> <p><strong>Outlier Detection Versions</strong></p> <p>Additionally, we generated a number of subsets for outlier detection:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>RGB Histograms</td> <td>Downsampled to 100000 objects (553 outliers)</td> <td><a href="6355684/files/aloi-27d-100000-max10-tot553.csv.gz">aloi-27d-100000-max10-tot553.csv.gz</a> <a href="6355684/files/aloi-64d-100000-max10-tot553.csv.gz">aloi-64d-100000-max10-tot553.csv.gz</a></td> </tr> <tr> <td>&nbsp;</td> <td>Downsampled to 75000 objects (717 outliers)</td> <td><a href="6355684/files/aloi-27d-75000-max4-tot717.csv.gz">aloi-27d-75000-max4-tot717.csv.gz</a> <a href="6355684/files/aloi-64d-75000-max4-tot717.csv.gz">aloi-64d-75000-max4-tot717.csv.gz</a></td> </tr> <tr> <td>&nbsp;</td> <td>Downsampled to 50000 objects (1508 outliers)</td> <td><a href="6355684/files/aloi-27d-50000-max5-tot1508.csv.gz">aloi-27d-50000-max5-tot1508.csv.gz</a> <a href="6355684/files/aloi-64d-50000-max5-tot1508.csv.gz">aloi-64d-50000-max5-tot1508.csv.gz</a></td> </tr> </tbody></table>

opencc-by-4.0Jun 2010View details →
zenodo52/100

Decadal BIOCLIM estimates based on ISIMIP3b climatic forcing data for the European continent

<p>This dataset contains BIOCLIM variables (plus huss, sfcwind, rsds) which have been prepared and calculated from the original ISIMIP3b bias-adjusted climate forcing data from 5 GCM models (obtained on 2023-08-07). <br><br>For more information on the original data and its properties, please see the ISIMIP3b modelling protocol and here specifically the climate forcing section <a href="https://protocol.isimip.org/#/ISIMIP3b/31-forcing-data" target="_blank" rel="noopener">https://protocol.isimip.org/#/ISIMIP3b/31-forcing-data</a> and <a href="https://doi.org/10.5194/gmd-17-1-2024">Frieler et al. (2024)</a>.</p> <p>The original climate forcing data (global extent, daily temporal grain) were cropped to the European extent and spatial-temporally aggregated. Here 10 year (decadal) steps were chosen as target climatology.<br><br>For each time slot (e.g. 10 years) and scenario (historical or ssps) the following 22 variables were calculated:</p> <p>bioclim01 = Annual Mean Temperature<br>bioclim02 = Mean Diurnal Range (Mean of monthly (max temp - min temp))<br>bioclim03 = Isothermality (BIO2/BIO7) (&times;100)<br>bioclim04 = Temperature Seasonality (standard deviation &times;100)<br>bioclim05 = Max Temperature of Warmest Month<br>bioclim06 = Min Temperature of Coldest Month<br>bioclim07 = Temperature Annual Range (BIO5-BIO6)<br>bioclim08 = Mean Temperature of Wettest Quarter<br>bioclim09 = Mean Temperature of Driest Quarter<br>bioclim10 = Mean Temperature of Warmest Quarter<br>bioclim11 = Mean Temperature of Coldest Quarter<br>bioclim12 = Annual Precipitation<br>bioclim13 = Precipitation of Wettest Month<br>bioclim14 = Precipitation of Driest Month<br>bioclim15 = Precipitation Seasonality (Coefficient of Variation)<br>bioclim16 = Precipitation of Wettest Quarter<br>bioclim17 = Precipitation of Driest Quarter<br>bioclim18 = Precipitation of Warmest Quarter<br>bioclim19 = Precipitation of Coldest Quarter<br>huss = Average (arithmetric mean) specific humidity<br>rsds = Average (arithmetric mean) Surface downwelling shortwave radiation<br>sfcwind = Average near-surface wind speed (arithmetric mean)<br><br>---<br><strong>Data properties:</strong></p> <table> <tbody> <tr> <td>Shared Socioeconomic Pathways (SSP)</td> <td>SSP1-2.6, SSP2-4.5, SSP3-7.0, SSP5-8.5</td> </tr> <tr> <td>General circulation models (GCMs)</td> <td>GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, UKESM1-0-LL</td> </tr> <tr> <td>Spatial grain</td> <td>0.5 degree (~50km&sup2;)</td> </tr> <tr> <td>Geographic projection</td> <td>WGS 84</td> </tr> <tr> <td>Temporal grain</td> <td>10 year steps</td> </tr> <tr> <td>Spatial extent</td> <td>Continental Europe including Turkey (see screenshot)</td> </tr> <tr> <td>Temporal extent</td> <td>1850 to 2010 (Historical), 2010 - 2100 (Future)</td> </tr> <tr> <td>Number of variables</td> <td>22</td> </tr> </tbody> </table> <p><br>All files are provided in netCDF (nc) format. The preprocessed datasets are provided as it and the author takes no responsibility for errors or misuse.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo52/100

Surface water and flooding dynamics data set based on seasonally continuous Landsat data (1986-2011) in a dryland river basin

<p>Animations of the data are available here:&nbsp;<a href="https://doi.org/10.5281/zenodo.2438110">https://doi.org/10.5281/zenodo.2438110</a></p> <p>If you are using this data set, please cite the following publication:</p> <p>Tulbure, M.G. and M. Broich (2018). Spatiotemporal patterns and effects of climate and land use on surface water extent dynamics in a dryland region with three decades of Landsat satellite data. Science of the Total Environment.&nbsp;https://www.sciencedirect.com/science/article/pii/S0048969718347466&nbsp;</p> <p>The data represent statistically validated surface water and flooding extent dynamics derived from seasonally continous Landsat TM/ETM+ data and random forest models, and summarised to the maximum extent of surface water per season between 1986-2011 over Australia&#39;s Murray-Darling Basin. The overall accuracy was over 99% and producer&#39;s accuracy for water 87% +/- 3%.&nbsp;</p> <p>The method is described in the following publication:&nbsp;<br> Tulbure, M.G., M. Broich, S.V. Stehman, A. Kommareddy. (2016). Surface water extent dynamics from three decades of seasonally continuous Landsat time series at subcontinental scale in a semi-arid region. Remote Sensing of Environment. 178: 142-157</p> <p>URL: https://www.sciencedirect.com/science/article/pii/S0034425716300621&nbsp;</p> <p>Data are provided in GeoTIFF format per season per year. File naming convention is as follows:<br> yy_inund_freq_season_SamplingMethod. For example, &quot;99_inund_freq_winter_max&quot; will represent inundation frequency for winter 1999 resampled using a maximum resampling method.&nbsp;</p> <p>Inundation frequency represents the number of times a pixel has been flagged as flooded out of the times that pixel had valid observations * 100. Valid observation exclude no data values and clouds.&nbsp;The valid range of inundation frequency is 0-100 [%], with 255 indicating no data values.&nbsp;Data type is&nbsp;eight bit unsigned integer (uint8).&nbsp;</p> <p>The data were resampled to 120m resolution to reduce file size. The resampling methods used include max (e.g. selects the max value of all non-NODATA contributing 30m pixels)&nbsp;and mean (median and min can be provided upon request). If you are unsure which resampling to use, you may want to start with the mean. &nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo52/100

DisVis-based filtering of contacts from co-evolution data (or other sources)

<p>Dataset described in the manuscript:&nbsp;<em>Improving the Quality of Co-evolution Intermolecular Contact Prediction with DisVis</em>Siri Camee van Keulen, Alexandre M.J.J. Bonvin</p> <p>Details about the data set can be found at: &nbsp;https://github.com/haddocking/contact-filtering</p> <p>This archive contains in addition all the models generated with HADDOCK.</p>

opencc-by-4.0Oct 2022View details →
edi52/100

Eight Mile Lake Research Watershed, Carbon in Permafrost Experimental Heating Research (CiPEHR): Half-hourly growing season, chamber-based, CO2 flux data, 2009-2021

The Carbon in Permafrost Experimental Heating Research (CiPEHR) project addresses the following questions: 1) Does ecosystem warming cause a net release of C from the ecosystem to the atmosphere?, 2) Does the decomposition of old C, that comprises the bulk of the soil C pool, influence ecosystem C loss?, and 3) How do winter and summer warming alone, and in combination, affect ecosystem C exchange? We are answering these questions using a combination of field and laboratory experiments to measure ecosystem carbon balance and radiocarbon isotope ratios at a warming experiment located in an upland tundra field site near Healy, Alaska in the foothills of the Alaska Range. This data contains CO2 fluxes measured using an automated chamber system that measures net ecosystem CO2 exchange (NEE). Measurements are made every ~1.5 hours and modeled half-hourly. Half hour ecosystem respiration is modeled using an exponential Q10 relationship when light conditions are low (PAR<5umol/m2/s) and using a hyperbolic light relationship when PAR>5umol/m2/s. GPP is calculated as the difference between NEE and Reco.

openOpenApr 2022View details →
edi52/100

Eight Mile Lake Research Watershed, Carbon in Permafrost Experimental Heating and Drying Research (DryPEHR): Growing season, chamber-based, CO2 flux data, 2009-2021

This drying and warming experiment addresses the following questions: 1) Does ecosystem drying, warming and permafrost thaw cause a net release or uptake of C from the ecosystem to the atmosphere?, 2) Does the decomposition of old C that comprises the bulk of the soil C pool influence ecosystem C loss? 3) How do drying and warming affect plant communities and ecosystem properties? We are answering these questions using a combined warming and drying experiment (DryPEHR), which is situated with the Carbon in Permafrost Experimental Heating Research (CiPEHR) project and located in an upland tundra field site near Healy, Alaska in the foothills of the Alaska Range. Warming treatment here refers to growing season air temperature warming (~1C) using open top chambers (OTC) combined with soil 'warming' using snow fences during the snow covered months. Drying is achieved using an automated pumping system that lowers the water table in the dry plots. Soil warming began in 2008; OTCs and drying in 2011. This data set includes measured values of CO2 fluxes during the growing season.

openOpenApr 2022View details →
zenodo48/100

Supplementary Data for MOCCASIN: A method for correcting known and unknown confounders in RNA-Seq-based splicing analysis

<p>Contents</p> <ol> <li><strong>moccasin_paper_env.yaml</strong>: conda environment file with R and Python packages and modules needed to reproduce &nbsp;analyses.</li> <li><strong>FigureReproduction.zip</strong>: data and code to reproduce main and supplemental figures.</li> <li><strong>MOCCASIN_ExampleDataset.zip</strong>: A small subset of the simulated data with example code to run MOCCASIN.</li> <li><strong>encode_corrected.zip</strong>: Folder with batch-corrected ENCODE differential splicing quantifications (dPSI).</li> </ol> <p>&nbsp;</p> <p>&nbsp;</p> <p>(1) <strong>moccasin_paper_env.yaml</strong></p> <p>Use the moccasin_paper_env.yaml file to create a conda environment from which all analyses for the paper can be reproduced.</p> <pre><code class="language-bash"># need to first install conda. See here: # https://docs.conda.io/en/latest/miniconda.html # Next, create a conda environment: conda env create --name moccasin_paper_env --file moccasin_paper_env.yaml --force # Activate the environment: conda activate moccasin_paper_env</code></pre> <p><br> The only Python packages not included in this environment are MAJIQ &amp; VOILA. Please see majiq.biocipers.org for installation instructions.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>(2) <strong>FigureReproduction.zip</strong></p> <p>Within FigureReproduction are folders with code and data to reproduce the main and supplemental figures of the publication. Each folder contains data, script(s) and a README.txt with instructions on how to reproduce figures.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>(3) <strong>MOCCASIN_ExampleDataset.zip</strong></p> <p>Within this folder is an example dataset to test MOCCASIN. The README.txt file contains detailed line-by-line instructions for how to run MOCCASIN and do post-MOCCASIN analyses. In this example, we show how to run MOCCASIN on a group of .majiq samples with one known confounding effect. Also demonstrated is how to run an &quot;explore unknown residuals&quot; analysis as described in the detailed methods in the supplemental of the paper.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>(4) <strong>encode_corrected.zip</strong></p> <p>Includes a file called ENCODE_BeforeAndAfterMOCCASIN.voila.tsv.zip which includes LSV quantifications before and after MOCCASIN. Each row in the file represents a junction from an LSV. Each column header starts with the prefix &quot;BeforeMOCCASIN&quot; or &quot;AfterMOCCASIN&quot; and headers ending in dPSI corresponds to the dPSI of an ENCODE knockdown vs control experiment.&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo48/100

TCOM-HF : Daily global gap-free stratospheric hydrogen fluoride (HF) profile data set based on TOMCAT CTM and Occultation Measurements

<p><strong>Methodology:&nbsp; TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated hydrogen fluoride (HF) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement&nbsp; differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the&nbsp; differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids.&nbsp; TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time hydrogen fluoride profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values.&nbsp; For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used.</strong></p> <p><strong>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride&nbsp; profiles on height (10-50 km) and pressure (300-0.1 hPa) levels:</strong></p> <p><strong>zmhf_TCOM_hlev_T2Dz_2000_2024.nc &ndash; height level data (10 to 50 km)</strong></p> <p><strong>zmhf_TCOM_plev_T2Dz_2000_2024.nc &ndash; pressure level data (300 to 0.1 hPa)</strong></p> <p><strong>Daily 3D profiles on height and pressure levels would be made available on request.</strong></p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

mzrtsim: Raw Data Simulation for Reproducible Gas/Liquid Chromatography–Mass Spectrometry Based Non-targeted Metabolomics Data Analysis

<p>All the data for 'mzrtsim: Raw Data Simulation for Reproducible Gas/Liquid Chromatography&ndash;Mass Spectrometry Based Non-targeted Metabolomics Data Analysis'</p> <p>sim.zip is stimulated data for intensity cutoff 0.05. simxcms.csv is peak intensity profiles for their simulated peaks.</p> <p>sim3.zip are simulated data for normal/leading/tailing peaks with tailing factor of 1, 0.8, and 1.5, respectively.</p> <p>All the csv files begin with sim3 are extracted peaks list from the sim3.zip with corresponding data analysis software.</p> <p>csv.zip recorded the m/z, retention time, intensity, and compounds name for simulated compound for each condition (sim.zip and sim3.zip).</p> <p>sep1.mzML: simulation for 8 isomers with similar m/z while different retention times. 7 peaks are non baseline separation peaks. Peaks profile is saved in spe1.csv file.</p> <p>xcms.csv, mzmine.csv, openms.csv: peaks found in sep1.mzML by xcms, mzmine 4.5 and openms, respectively.</p> <p>R code:&nbsp;<a href="https://github.com/yufree/democode/blob/master/meta/simfin.R">https://github.com/yufree/democode/blob/master/meta/simfin.R</a></p> <p>Website of mzrtsim package: https://yufree.github.io/mzrtsim/</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

BASE-9 binarity and stellar masses from Gaia DR3, 2MASS, and Pan-STARRS data for six open clusters: NGC 2168, NGC 7789, NGC 6819, NGC 2682, NGC 188, NGC 6791

<h2>Data sets as described in "Goodbye to Chi-by-Eye: A Bayesian Analysis of Photometric Binaries in Six Open Clusters", Childs et al. 2023 <a href="https://ui.adsabs.harvard.edu/abs/2023arXiv230816282C/abstract">https://ui.adsabs.harvard.edu/abs/2023arXiv230816282C/abstract</a></h2>

openmit-licenseNov 2023View details →
zenodo48/100

Experimental data of dissipative embedded column base connections tested under cyclic lateral loading

<p>This experimental dataset is comprised of the following items:</p> <p>(a) the deduced experimental data of conventional/dissipative embedded column base connection specimens, which contains base moment, column drift ratio, and axial shortening responses (TestData.xlsx);</p> <p>(b) photos of&nbsp;each specimen taken during cyclic loading (C-N-0_Test_Photos.7z, D-M1-1_Test_Photos.7z, D-M1-3_Test_Photos.7z, D-M1-5_Test_Photos.7z, D-M2-2_Test_Photos.7z);</p> <p>(c) characteristic videos for each specimen that demonstrate the cyclic behavior (Test_Video.7z);</p> <p>(d) Digital image correlation (DIC) images taken during&nbsp;cyclic loading to obtain strain fields near the steel column/reinforced concrete foundation interface (C-N-0_DIC_Photos.7z, D-M1-1_DIC_Photos.7z, D-M1-3_DIC_Photos.7z, D-M1-5_DIC_Photos.7z, D-M2-2_DIC_Photos.7z);&nbsp;</p> <p>(e) Videos that demonstrate strain fields of column flanges of both conventional and dissipative embedded column base connection specimens (DIC_Video.7z)&nbsp;</p> <p>Please read the &quot;README&quot; file contained in each folder for more detailed information regarding each data.</p> <p>&nbsp;</p>

opencc-by-2.0Jul 2021View details →
zenodo48/100

Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics

<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from&nbsp;Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness,&nbsp;<br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al.&nbsp;</p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule &ldquo;create_morphological_analysis&rdquo;. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from&nbsp;Burress et al. 2017.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Rating curves based on satellite altimetry and in-situ discharge data

<h1>Context:&nbsp;</h1> <p>The ESA river discharge Climate Change Initiative (CCI) project is a precursor study. It aims to derive long term climate data records (at least over 20-years) of river discharge for some selected river basins (and some locations in the river network) using satellite remote sensing observations (altimetry and multispectral images) and ancillary data. It aims to provide a proof-of-concept for the feasibility for a potential River Discharge ECV product to meet the requirements for the&nbsp;<a href="https://gcos.wmo.int/en/essential-climate-variables/rivers/" target="_blank" rel="noopener">Global Climate Observing System</a>. This project covers precursor activities towards the production of data products that address the GCOS-defined requirements for the River Discharge ECV.</p> <h1>Data description :</h1> <p>Just as in-situ stage measurements can be used to gauge river discharge, altimetry-derived water surface elevation (WSE) can serve as an alternative means of estimating river discharge when discharge time series data is available. Several methodologies have been documented for deriving discharge time series from multimission altimetry observations and supplementary data (Biancamaria et al., 2024). At least two approaches will be used, depending on the available in situ discharge and altimetry water surface elevation (WSE) time series:</p> <p>&sdot; <strong><em>Method 1</em>: </strong>The preferred approach relies on the altimetry water surface elevation time series and in situ discharge time series to create a rating curve (RC) characterized by a power relationship between these two variables following a Bayesian approach (Rantz et al., 1982). However, this method necessitates a significant overlap period between discharge data and radar altimetry measurements (e.g., Biancamaria et al., 2011; Papa et al., 2012), or it requires the assumption that the rating curve remains valid and consistent when discharge data is only available prior to the altimetry observation period.</p> <p>&sdot; <em><strong>Method 2:</strong></em> The final option, in cases where there is no temporal overlap between in-situ or simulated discharge and water surface elevation data, assumes that the validity and stability of the rating curve persist across the various time periods covered by the two datasets. Both of these time periods should be sufficiently long to encompass a wide range of events. With this assumption, Tourian et al. (2013, 2017) introduced a method for calculating the rating curve, not based on the time series of discharge and water surface elevation, but on the distribution of their quantiles. This method has been adopted by a limited number of recent studies (e.g., Belloni et al., 2021). However, it&rsquo;s important to note that this methodology naturally introduces higher errors when compared to the preferred approach. For this reason, this methodology will be validated over some stations with various hydrological dynamics and satisfying previous methods (overlap period exists between WSE and Q).</p> <h1>Approaches to derive Rating Curve (RC) :</h1> <h2>Bayesian Approach :</h2> <p>The Bayesian method is a robust statistical approach used for constructing a rating curve, frequently applied in the field of hydrology when the goal is to estimate unknown parameters from observed data, while taking into consideration the associated uncertainty in these estimates.&nbsp;</p> <p>According to this, the estimation of the rating curve using the Bayesian method involves several steps:</p> <ul> <li>The initial step entails defining a probabilistic model that describes the relationship between observed data and the parameters we aim to estimate. In many hydrological applications, the relationship between discharge data (Q) and water surface elevation data (WSE) is often expressed as a power function:</li> </ul> <p><em>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Q = a&sdot;(WSE-z</em><em>0</em><em>)</em><sup><em>b</em></sup></p> <p>Here,&nbsp;<em>a, z0</em> and <em>b</em> are the parameters of the rating curve. <em>a,</em> is a scaling coefficient governing the magnitude of the Q-WSE relationship, <em>b,</em> characterizes the nature of this relationship, and <em>z0</em>, represents the height of the free surface above the reference point, corresponding to the river bottom's altitude.&nbsp;The power relationship is especially pertinent due to its consistency with numerous hydrodynamic phenomena. The exponent b within the equation allows for the representation of distinctive flow characteristics, including factors like roughness and channel geometry. Moreover, it offers adaptability in modelling to accommodate variations in flow characteristics, whether they are turbulent or laminar. This relationship, despite its mathematical simplicity, facilitates the fine-tuning of model adjustments in accordance with observed data (Chow, 1959).</p> <ul> <li>The second step involves the use of prior normal distributions, reflecting our prior knowledge about these parameters. These distributions can either be informative or uninformative, depending on our level of knowledge.&nbsp;The limits and ranges for a, z0 and b can vary depending on the specific context of the study, the dataset used, and the characteristics of the river or channel being analysed.</li> </ul> <p><u>- Coefficient &ldquo;a&rdquo;</u>:&nbsp; adjustment parameter for the rating curve representing the scaling factor for discharge. Its value can significantly fluctuate based on various factors such as the characteristics of the river or channel, hydraulic conditions, and other influencing factors. Consequently, "a" must be non-negative and constrained within a sensible range specific to the system under study. Following the Manning equation, &ldquo;a&rdquo; must be equal to W/n*S<sup>1/2</sup> (Chow et al., 1988) where W is the river&rsquo;s width (m), n the Manning&rsquo;s roughness coefficient and S the slope (m/m). Given the considerable variability in river width and slope across different stations, a feasible range for this coefficient can be considered as:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; a &isin; [0; 3000]</p> <p><u>- Coefficient &ldquo;b&rdquo;</u>: adjustment parameter representing the exponent of the rating curve and indicating the hydraulic condition of the study site. Like "a," this value must comply with physical constraints and cannot be negative. Following the Manning equation, &ldquo;b&rdquo; must be equal to 5/3 for reference hydraulic condition (Rantz et al., 1982). To accommodate the variability in system characteristics across sites, the following range values can be considered for this coefficient:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; b &isin; [0; 5]</p> <p><u>- Coefficient &ldquo;z0&rdquo;</u>:&nbsp;offset or the elevation at which discharge begins. It should be within the range of elevations relevant to your study. For this <em>reason, the value</em> cannot exceed the minimum value of water surface elevation (WSE) and the range value need to consider of the variability in term of water depth over the sites. A feasible range for this coefficient can be considered as:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; z0 &isin; [min(WSE)-30; min(WSE)]</p> <ul> <li>The final step involves parameter estimation. The posterior distribution of the parameters yields probabilistic estimates of the rating curve parameters in the form of mean values (optimal values) and credibility intervals (95th percentiles). This accounts for the uncertainty associated with these parameters and is achieved through Markov Chain Monte Carlo (MCMC) sampling from the posterior distribution. Two commonly employed MCMC algorithms are "NUTS" (No-U-Turn Sampler) and "Metropolis-Hastings." The Metropolis-Hasting sampler "MH" algorithm, which is relatively simple and efficient where a balance between exploration and exploitation is desired. This algorithm can be adapted to sample from discrete state spaces.</li> </ul> <h2>Quantile approach :&nbsp;</h2> <p>The Quantile approach employs statistical modelling using quantile functions to create a rating curve, eliminating the necessity for overlapping measurements. This algorithmic method enables the estimation of river discharge using satellite altimetry, even in instances where there are no in situ measurements within the altimeter's timeframe. This approach has undergone application and validation in diverse river basins spanning different climatic zones, such as the Amazon, Brahmaputra, Danube, Niger, and Ob (Tourian et al., 2013).</p> <p>Assuming a stationary flow behaviour and no modification in the river bathymetry both at the altimetry virtual station and at the in-situ gage, this approach ensures the utilization of historical in situ data in current applications. This method computes the quantile functions of the altimetry water surface elevation on one hand and of the discharge time series on the other hand. Then a scatter plot of these in-situ discharge quantiles versus altimetry water surface elevation quantiles is computed to establish the rating curve using the bayesian approach described previously.</p> <h1>File description :</h1> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>basin-station</td> <td>Basin name in capital letters and Station name in capital letters separated by "_" and where spaces have been replaced by "-".</td> </tr> <tr> <td>lon</td> <td>Longitude in decimal degrees [-180,180] with 4 decimals - corresponding to the insitu discharge station.</td> </tr> <tr> <td>lat</td> <td>Latitude in decimal degrees [-90,90] with 4 decimals &ndash; corresponding to the insitu discharge station.</td> </tr> <tr> <td>a</td> <td>Adjustment parameter for the rating curve representing the scaling factor for discharge. Number with 3 decimals.</td> </tr> <tr> <td>b</td> <td>Adjustment parameter representing the exponent of the RC and indicating the hydraulic condition of the study site. Number with 3 decimals.</td> </tr> <tr> <td>z0</td> <td>Offset of the elevation at which discharge begins. Number with 3 decimals.</td> </tr> <tr> <td>a_sd</td> <td>Standard deviation of the coefficient "a". Number with 3 decimals.</td> </tr> <tr> <td>b_sd</td> <td>Standard deviation of the coefficient "b". Number with 3 decimals.</td> </tr> <tr> <td>z0_sd</td> <td>Standard deviation of the coefficient "z0". Number with 3 decimals.</td> </tr> <tr> <td>period</td> <td>Period used to compute the rating curve under the format %Y-%m-%d where the start and the end dates are separated by ":"</td> </tr> <tr> <td>nb</td> <td>Number of overlap dates to compute the rating curve.</td> </tr> <tr> <td>Methodology</td> <td>Methodology used to compute the rating curve. The first part describes the approach used to compute the RC and the second part, separated by &ldquo;_&rdquo;, describes the algorithm used. To avoid any issue for the reader the spaces have been replaced by &ldquo;-&rdquo;. At the end 2 approaches has been used: &ldquo;Overlap-approach&rdquo; or &ldquo;Quantile-approach&rdquo; and 2 algorithms: &ldquo;Bayesian-algorithm&rdquo; or &ldquo;Multiple-algorithms&rdquo; designed for Arctic rivers experiencing frozen periods.&nbsp;</td> </tr> <tr> <td>Source</td> <td>In-situ data sources to compute the rating curve. If multiple sources has been used, the sources are separate by "/"</td> </tr> </tbody> </table> <p>---------</p> <p><em>THE DATASET IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR&nbsp;</em><em>IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,</em><br><em>FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE&nbsp;</em><em>AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER&nbsp;</em><em>LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,&nbsp;</em><em>OUT OF OR IN CONNECTION WITH THE DATASET OR THE USE OR OTHER DEALINGS IN THE&nbsp;</em><em>DATASET.</em></p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

ERA5 based training, validation and evaluation data for retrievals combining 22-58 GHz with 175-340 GHz microwave radiometer measurements during MOSAiC

<p>This data set is used for the training, validation and evaluation of retrievals of temperature and specific humidity profiles, as well as integrated water vapour from simlulated or measured microwave brightness temperatures (TBs), which are described in <strong>[1]</strong>.</p> <p>The data set consists of yearly files (2001-2018, 6-hourly resolution) that include data from the European Centre for Medium-Range Weather Forecasts's ERA5 reanalysis <strong>[2]</strong> and simulated TBs in the microwave spectrum. TB simulations were performed with PAMTRA <strong>[3,4]</strong> on the native ERA5 model level resolution at frequencies of a low frequency Humidity and Temperature Profiler (HATPRO, 22-58 GHz) and of a Low Humidity Profiler (LHUMPRO-243-340, aka MiRAC-P, 175-340 GHz). Afterwards, the ERA5 model level data has been interpolated to a new height grid (dimension 'z'), of which the lowest 43 indices equal the height grid of the retrieval that is developed with this data set. The upper 11 indices are included for additional TB simulations needed for the information content estimation performed and are not used for the retrievals to avoid the tropopause.</p> <p>The trained retrieval is applied to observations from the HATPRO and MiRAC-P that were installed onboard the research vessel Polarstern during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition.</p> <p><strong>[1]:</strong> Walbr&ouml;l, A., Griesche, H. J., Mech, M., Crewell, S., and Ebell, K.: Combining low- and high-frequency microwave radiometer measurements from the MOSAiC expedition for enhanced water vapour products, Atmospheric Measurement Techniques, 17, 6223-6245, https://doi.org/10.5194/amt-17-6223-2024, 2024.</p> <p><strong>[2]:</strong> Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Hor&aacute;nyi, A., Mu&ntilde;oz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., H&oacute;lm, E., Janiskov&aacute;, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Th&eacute;paut, J.: The ERA5 global reanalysis, Quarterly Journal of the Royal Meteorological Society, 146, 1999&ndash;2049, https://doi.org/10.1002/qj.3803, 2020.</p> <p><strong>[3]:</strong> Mech, M., Maahn, M., Kneifel, S., Ori, D., Orlandi, E., Kollias, P., Schemann, V., and Crewell, S.: PAMTRA 1.0: the Passive and Active Microwave radiative TRAnsfer tool for simulating radiometer and radar measurements of the cloudy atmosphere, Geoscientific Model Development, 13, 4229&ndash;4251, https://doi.org/10.5194/gmd-13-4229-2020, 2020.</p> <p><strong>[4]:</strong> Mech, M., Maahn, M., Ori, D., Kneifel, S., and Orlandi, E.: PAMTRA Package &ndash; Passive and Active Microwave TRANsfer, available at: https://github.com/igmk/pamtra (last access: 6 September 2020), 2019c.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Data Files for Climate-based Maize Loss Rate Simulations

<p>This archive contains data files from an <a href="../records/13356711">open source pipeline</a> looking at how crop insurance rates may change in the future within the US Corn Belt using <a href="https://www.sciencedirect.com/science/article/pii/S0034425715001637">SCYM</a> and <a href="https://www.chc.ucsb.edu/data/chc-cmip6">CHC-CMIP6</a>. These are available under a Creative Commons license. Unless otherwise specified, these report on SSP245.</p> <p>See README for more details including column-level description of each resource. Funded by the <a href="https://dse.berkeley.edu/">Eric and Wendy Schmidt Center for Data Science and Environment</a> at the University of California, Berkeley.</p>

opencc-by-nc-4.0Aug 2024View details →
zenodo48/100

Global Ocean Heat Content Anomalies and Ocean Heat Uptake based on mapping Argo data using local Gaussian processes

<p>Monthly Ocean Heat Content Anomalies (OHCA) in the top 2000 dbar of the ocean are calculated (during 2004-2024, equatorward of 65 degree latitude) subtracting the mean over the period 2004-2024 from the monthly time series of OHC. Yearly OHCA time series are then calculated that include 1. one point per year, i.e., from averaging Jan to Dec (see files ending in &ldquo;yearly.nc&rdquo;), and 2. two points per year, i.e., from averaging Jan to Dec and Jul to Jun, respectively&nbsp; (see files ending in &ldquo;yearly2.nc&rdquo;). OHC fields are mapped using locally stationary Gaussian processes (defined over space and time) with data-driven decorrelation scales (Kuusela and Stein, 2018). A linear time trend was included in the estimate of the mean field (along with spatial terms and harmonics for the annual cycle). Mapping is done separately for different vertical sections: 15-20 dbar, 15-300 dbar, 300-700 dbar, 700-1850 dbar, 1800-1850 dbar. The 15-20 dbar (1800-1850 dbar) section is used to estimate OHCA for 0-15 dbar (1850-2000 dbar), where observations are sparser. Different vertical sections are combined to estimate global OHCA time series for 0-2000 dbar, 0-700 dbar, 700-2000 dbar (as indicated in the file names). The attribute "area" is included in the netcdf files and it tells the corresponding surface area for the estimates. Regions of the ocean that are shallower than 300 m or are not sufficiently well sampled by the Argo array are not included. Maps of the ocean masks used for the different vertical sections can be found in the .png files (blue shading indicates the area used for the horizontal integral); the bathymetry mask by Roemmich and Gilson (included in the file RG_ArgoClim_Temperature_2019.nc at https://sio-argo.ucsd.edu/RG_Climatology.html) is also used to define the ocean mask. Ocean Heat Uptake is calculated from the monthly OHCA and then averaged as described above to produce yearly time series included in the files for the different layers.</p> <p>For the uncertainty at each time point, the standard deviation of each OHCA/OHU value in the time series is included. When plotting a time series, the user may consider, e.g., shading plus/minus 1* or 1.96*standard deviation (corresponding to a&nbsp; confidence level of 68% or 95% respectively). These standard deviations in the files are estimated using spatially and temporally dependent conditional simulations of monthly gridded anomalies. When combining different layers, the standard deviation of the sum is conservatively estimated as the sum of the standard deviations.&nbsp;</p> <p>Finally, OHCA/OHU trends are estimated via a least-squares fit and reported in the variable metadata with uncertainties (confidence level of 68%). Trend uncertainties are estimated by repeating the fit for each member of the conditional simulation ensemble described above.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record