Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,195
datasets available to search
ShareScore release 0.9.0
Dataset results
7,195 results for “SET”
PIE LTER surface elevation table (SET) pin height data from twelve marsh sites in northeast Massachusetts.
Surface elevation table (SET) measurements from 26 SETs at 9 marsh sites in the Plum Island Sound Long-Term Ecological Research Site in the Great Marsh, Massachusetts. SET measurements are useful for determining the relative elevation change of marsh sediments. Precise measurements of sediment elevation in marshes is useful for determining rates of elevation change in response to changes in sea level.
Biogeochemistry data set for Imnavait Creek Weir on the North Slope of Alaska 2002-2024.
Data file containing biogeochemical data of water samples collected in Imnavait Creek, North Slope of Alaska. Sample site descriptors include a unique assigned number (sortchem), site, date, time, depth, distance (downstream), and elevation. Values of variables measured in the field include temperature, conductivity, pH. Chemical analysis for samples include alkalinity, dissolved organic carbon (DOC), inorganic and total dissolved nutrients particulate carbon, nitrogen, and phosphorus, cations and anions.
Biogeochemistry data set for soil waters, streams, and lakes near Toolik Lake on the North Slope of Alaska, 2012 through 2020
Data file of the biogeochemistry of samples collected at various sites near Toolik Lake, North Slope of Alaska. Sample site descriptors include a unique assigned number (sortchem), site, date, time, depth, distance (downstream from a reference location), elevation, treatment, date-time, category, and water type (lake, surface, soil). Physical measures collected in the field include temperature (water, soil, well water), conductivity, pH, and average thaw depth in soil. Chemical analyses for the sample include alkalinity; dissolved inorganic and organic carbon (DIC and DOC); dissolved gases CO2 and CH4; inorganic and total dissolved nutrients (NH4, PO4, NO3, TDN, TDP); particulate carbon, nitrogen, and phosphorus (PC, PN, and PP); cations (Ca, Mg, Na, K, and Si); and anions (SO4 and Cl).
North Temperate Lakes LTER: Manure Managment in Urbanizing Settings 2003 - 2004
The management of manure in urbanizing settings is a critical issue in the Lake Mendota watershed. The primary focus of this project was to examine the difficulties faced by livestock operations when managing manure on field systems that are fragmented by development. A survey regarding manure management was sent to Dane County, WI farms within the Lake Mendota watershed. The survey was conducted in two phases; March to May 2003 and March to May 2004. This dataset and accompanying survey entitled "Manure Management on the Urban Fringe" is available for users wishing to ascertain animal feeding operation size and management patterns in the Lake Mendota watershed. The data also include the distance from animal feeding operations to nearest urban centers via Euclidean (crow flies) and Road Network distances. Results suggest that exurban developments exert a strong influence on manure management routines of livestock producers. This influence is very local. Farmers in an urbanizing setting were more likely to encounter problems during manure hauling when the fields they were accessing were in close proximity to urban developments, regardless of their proximity to the urban core. The distances and times required to haul manure between the farm and the most distant field increased in the last five years. Land rental rates steadily increased at the same time that lease lengths shortened. Cash grain land tends to be sparse as livestock producers compete with developers for tracts on which to distribute manure. Manure brokering is a possible strategy to monitor land availability and coordinate manure placement between farms Cabot, P. E., S. K. Bowen, and P. J. Nowak. 2004. Manure management in urbanizing settings. Journal of Soil and Water Conservation 59:235-243. The survey "Manure Management on the Urban Fringe" was developed with assistance from Roger Schmidt and Charmaine Tryon-Petith with the Integrated Crop and Pest Management Program, University of Wisconsin-Madison.
North Temperate Lakes LTER General Lake Model Parameter Set for Lake Mendota, Summer 2016 Calibration
The General Lake Model (GLM), an open source, one-dimensional hydrodynamic model, was used to simulate various physical, chemical, and biological variables on Lake Mendota between 15 April 2016 and 11 November 2016. GLM (v.2.1.8) was coupled to the Aquatic EcoDynamics (AED) module library via the Framework for Aquatic Biogeochemical Modeling (FABM). GLM-AED requires four major “scripts†to run the model. First, the glm2.nml file configures lake metadata, meteorological driver data, stream inflow and outflow driver data, and physical response variables. Second, the aed2.nml file configures various biogeochemical modules for the simulation of oxygen, carbon, phosphorus, and nitrogen, among others. Third, aed2_phyto_pars.nml configures all parameters pertaining to phytoplankton dynamics. And fourth, aed2_zoop_pars.nml configures all parameters pertaining to zooplankton dynamics. This dataset contains parameter descriptions and values as they were used to simulate organic carbon and greenhouse gas production on Lake Mendota in summer 2016. Meteorological data and stream files used in this calibration are also included in this dataset. Additional methods and model descriptions can be found in J.A. hart’s Masters Thesis, University of Wisconsin-Madison Center for Limnology, May 2017. Readers are referred to the GLM (Hipsey et al. 2014) and AED (Hipsey et al. 2013) science manuals for further details on model configuration.
Pollinator visitation, flower count, and seed set in Black Sand plots, 2020.
Anthropogenic climate change is altering interactions among numerous species, including plants and pollinators. Plant-pollinator interactions, crucial for the persistence of most plant and many insect species, are threatened by climate change-driven phenological shifts. Phenological mismatches between plants and their pollinators may affect pollination services, and simulations indicated that these mismatches may reduce floral resources available to up to 50% of insect pollinator species. Although alpine plants rely heavily on vegetative reproduction, seedling recruitment and seed dispersal are likely to be important drivers of alpine community structure. Similarly, advanced flowering may expose plants to increased risk of frost damage and shifted soil moisture regimes; phenologically advanced plants will experience these environmental factors differently, which may alter their floral resource production. These effects may be dependent upon topography. Some species of alpine plants on the Niwot Ridge have displayed advanced phenology under treatments of advanced snowmelt (Forrester, 2021). However, little is understood about how these differences in distribution and phenology affect pollinator community composition and plant fecundity. Here we strive to examine how experimentally-induced changes in the timing of flowering and number of flowers produced by plants impact plant-pollinator interactions and seed set. We also ask how topography and the number of flowers interact with early snowmelt to affect pollination rates and the diversity of pollinating insects. Finally, we ask how seed set of Geum rossii is affected by pollinator visitation at different times of the season, under experimentally advanced snowmelt versus unmanipulated snowmelt, and with visitation by different insect taxa. In summer 2020, we found that plots with advanced phenology experienced peaks in pollinator visitation rates and pollinator diversity earlier than plots with unmanipulated snowmelt.
The NANOGrav 12.5-year Wideband Data Set (version 12yv4)
<p>The NANOGrav 12.5-year wideband data set (public release "12yv4") is the supplemental data set accompanying Alam et al. 2021, "The NANOGrav 12.5 yr Data Set: Wideband Timing of 47 Millisecond Pulsars," The Astrophysical Journal Supplement Series, 252, 5, DOI 10.3847/1538-4365/abc6a1. It contains wideband pulse times of arrival, models describing frequency-dependent template profiles, pulsar timing models, timing residuals, and clock files.</p> <p>Details about the contents of these files are contained in NANOGrav_12yv4_wideband/README, as well as in NANOGrav_12yv4_wideband/wideband/README.wideband. The narrowband version of this dataset (published in Alam et al. 2021, ApJS, 252, 4, DOI: 10.3847/1538-4365/abc6a0) can be found at Zenodo DOI: 10.5281/zenodo.4312297. Both the narrowband and wideband data sets are also available at <a href="http://data.nanograv.org">data.nanograv.org</a>.</p>
An updated mass-radius analysis of the 2017-2018 NICER data set of PSR J0030+0451
<p>Summarised posterior sample files associated with the preprint "An updated mass-radius analysis of the 2017-2018 NICER data set of PSR J0030+0451" by Vinciguerra et al. (2023; <a href="https://doi.org/10.48550/arXiv.2308.09469">arXiv</a>; accepted for publication in ApJ).</p><p>Also included are examples of model modules in the Python language using the X-PSI framework; and Jupyter analysis notebooks.</p><p>Please refer to the READme for detailed information.</p>
Data set for the journal article ''Nanoscale chemical reaction exploration with a quantum magnifying glass''
<div>This data set includes the raw data of the esterification and hydrogenation discussed in the journal article alongside with the Scine Puffin Singularity container, steering protocol files, Swoose parameters, (pre-)releases of the software, and Python scripts for individual steps without the graphical user interface to reproduce the data.</div>
WILLOW - Norther: data set for the full-scale validation of model-based virtual sensing methods for an operational offshore wind turbine
<h1><em><strong>1. General description </strong></em></h1> <p>This data set contains as-build design information, as well as full-scale vibration response measurements from an operational offshore wind-turbine. The turbine is part of the Norther wind farm which is located in the Belgian North Sea<em> </em>and includes a total of 44 Vestas V164 (8.4MW) wind turbines on monopile foundations, see <a href="../api/records/11093262/draft/files/Fig1_Norther_locaction.png/content" target="_blank" rel="noopener noreferrer">Fig1_Norther_locaction.png</a>. This data set is intended to verify and validate model-based virtual sensing algorithms, using data as well as modeling information from a real turbine. </p> <h2><em><strong>1.1 Summary of the shared structural information</strong></em></h2> <p>The included information entails a detailed description of the geometric properties of the monopile and transition piece, distributed and lumped structural masses . All information shared in this record is conform the as-designed documentation. An example of the lumped masses considered in the model input files is presented in "<a href="../api/records/11093262/draft/files/Fig2_Sensor_Network.png/content" target="_blank" rel="noopener">Fig2_Sensor_Network.png"</a></p> <h2><em><strong>1.2 Summary of the shared geotechnical information</strong></em></h2> <p>Monopiles are distinguished by the significant role of soil-structure interaction. Ground reaction is most typically included in the structural model as non-linear p-y curves. Different p-y curves are available for a certain number of soils in the standards applicable to offshore structures (API RP 2GEO, 2011, and ISO 19901-4:2016(E), 2016).</p> <p>The required soil properties to define p-y curves according to the API framework are given in the soil profile provided in a separate Excel. Rather than symbols, the name of the soil properties is generally used as column header (e.g., <em>Undrained shear strength</em>). Therefore, it is straightforward to identify each soil parameter. The only soil parameter that might lead to confusion is:</p> <ul> <li><em>"epsilon50 [-]" </em>represents the vertical strain at half the maximum principal stress difference in a static undrained triaxial compression test on an undisturbed soil sample.</li> </ul> <p>It's worthy to note that estimates for the small shear strain stiffness, referred to as Gmax, are also included. Despite not being required as an input to define the API p-y curves, this parameter remains a key input for other soil reaction frameworks than the API (e.g., PISA). </p> <h2><em><strong>1.3 Summary of the shared measurement data</strong></em></h2> <p>Two sets of measurement data have been curated for validation purposes; the first interval has been collected during parked conditions, whereas the second interval has been collected during rated operational conditions. Both records have a length of 2 hours, and are subdivided into 10-minute data sets. Furthermore 1Hz SCADA data has been made available for the selected intervals. All different data sources are time synchronized and have been subjected to several internal quality checks. </p> <p>The sensor network on NRT-WTG is illustrated in in <strong>Fig. 2, </strong>whereas a description of the sensor types is presented in <strong>Tab.1.</strong> The acceleration sensors are installed in the horizontal plane, and measure tangential (Y) and orthogonal (X) to the wall, where the positive Y direction is pointing clockwise and the positive X direction is pointing inwards. All strain sensors are installed vertically and are located on the inside of the wall.</p> <table> <tbody> <tr> <td><strong>Data type </strong></td> <td><strong>Sensor type</strong></td> <td><strong>Fs (Hz)</strong></td> <td> <p><strong>Level mLAT (m)</strong></p> </td> <td><strong>Description </strong></td> </tr> <tr> <td>Acceleration (g) </td> <td>Piezo-electric acc. sensor (<strong>ACC</strong>)</td> <td>30</td> <td>15, 69, 97 </td> <td>3 Bi-directional accelerometers at different levels. LAT 15 installed at 240 degree heading; LAT 69 and 97 at 60 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Resistive strain gauge (<strong>SG</strong>)</td> <td>30</td> <td>14</td> <td>6 SGs: equally spaced around the inner circumference of the can. Headings: 50, 110, 170, 230, 290, 350 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Fiber-Bragg Grating strain gauge (<strong>FBG</strong>)</td> <td>100</td> <td>-17, -19</td> <td>2 FBGs per level at 165 and 255 degree respectively.</td> </tr> </tbody> </table> <p><strong>Table 1. Description of sensor types.</strong></p> <p>The FBG strain time series have been synchronized with the SG time series using using a cross-correlation based approach. Therefore the SG data has been used to genereate refrence strain time series at the headings of the FBG sensors; the FBG data is subsequently synchronized with regard to this reference time series. No synchronization of the acceleration data was needed, since these are collected using the same data aquisition system as the SG data. </p> <p>The SG strain time series have been calibrated and temperature compensated, whereas this is not the case for the FBG strain time series. The latter have a yet to be determined calibration offset. </p> <p>In conjunction to the sensor channels presented in <strong>Tab. 1</strong>, 1 Hz SCADA data is provided. A summary of the provided SCADA parameters, all sampled at 1Hz, is presented in <strong>Tab 2.</strong></p> <table> <tbody> <tr> <td><strong>Parameter</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Wind speed</td> <td>m/s</td> <td>Wind speed as recorded in the turbine SCADA</td> </tr> <tr> <td>Wind direction</td> <td>°</td> <td>Wind direction relative to North (0°) as recorded in the turbine SCADA</td> </tr> <tr> <td>Yaw angle</td> <td>°</td> <td>Yaw orientation of the nacelle relative to North (0°) as recorded in the turbine SCADA</td> </tr> <tr> <td>Pitch angle</td> <td>°</td> <td>Rotor blade pitch as recorded in the turbine SCADA</td> </tr> <tr> <td>Rotor speed</td> <td>rpm</td> <td>Rotor speed in rotations per minute as recorded in the turbine SCADA</td> </tr> <tr> <td>Power</td> <td>kW</td> <td>Active power of the turbine as recorded in the turbine SCADA</td> </tr> </tbody> </table> <p><strong>Table 2. </strong>List of provided SCADA parameters</p> <p> </p> <p>A summary of the selected intervals and relevant corresponding scada parameters is given in <strong>Tab 3</strong>.</p> <table> <tbody> <tr> <td><strong>Scenario </strong></td> <td><strong>T1 (UTC)</strong></td> <td><strong>T2 (UTC) </strong></td> <td><strong>Windspeed</strong></td> <td><strong>RPM </strong></td> <td><strong>Pitch </strong></td> </tr> <tr> <td>Parked</td> <td> <p>03/07 01:30</p> </td> <td> <p>03/07 03:30</p> </td> <td>< 4.5 m/s</td> <td>~1</td> <td>~18 °</td> </tr> <tr> <td>Rated</td> <td> <p>05/07 22:30</p> </td> <td> <p>06/07 00:30 </p> </td> <td>~15 m/s</td> <td>10.5</td> <td>8.1°</td> </tr> </tbody> </table> <p><strong>Table 3. </strong>Selected data intervals and relevant scada parameters</p> <p> </p> <h1><em><strong>2. Included in this version </strong></em></h1> <h2><em><strong>2.1 Version - 0.1.0</strong></em></h2> <ul> <li>Relevant Design information can be found in: <ul> <li>Geometry data for NRT-WTG: "WILLOW-Geometry_v4.xlsx"</li> <li>Best estimate soil profile: "WILLOW-BE_soil_profile.xlsx"</li> </ul> </li> <li>Acceleration, strain and scada data can be found in the following parquet files: <ul> <li>Measurement data for the parked case: "NRT-WTG_Parked.parquet.gz"</li> <li>Measurement data for the rated case: "NRT-WTG_Rated.parquet.gz"</li> </ul> </li> </ul> <p> </p> <h1><em><strong>3. Importing parquet files </strong></em></h1> <p>To import the measurement data into Python it is recommended to use pandas:</p> <pre>import pandas as pd<br># Read Parquet file with Pandas: relative_file_path = '<a href="../api/records/11093262/draft/files/NRT-WTG_Parked.parquet.gz/content" target="_blank" rel="noopener noreferrer">NRT-WTG_Parked.parquet.gz</a>' data = pd.read_parquet(relative_file_path ) <br><br>Once the dataframe has been imported, the users can process/re-arrange the raw data according the their needs; it should be noted that the imported dataframe contains NAN values - these are caused by the different sampling rates of the provided signals. </pre>
ELKI Multi-View Clustering Data Sets Based on the Amsterdam Library of Object Images (ALOI)
<p>These data sets were originally created for the following publications:</p> <p><em>M. E. Houle, H.-P. Kriegel, P. Kröger, E. Schubert, A. Zimek</em><br> <strong>Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?</strong><br> In Proceedings of the 22nd International Conference on Scientific and Statistical Database Management (SSDBM), Heidelberg, Germany, 2010.</p> <p><em>H.-P. Kriegel, E. Schubert, A. Zimek</em><br> <strong>Evaluation of Multiple Clustering Solutions</strong><br> In 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, 2011.</p> <p>The outlier data set versions were introduced in:</p> <p><em>E. Schubert, R. Wojdanowski, A. Zimek, H.-P. Kriegel</em><br> <strong>On Evaluation of Outlier Rankings and Outlier Scores</strong><br> In Proceedings of the 12th SIAM International Conference on Data Mining (SDM), Anaheim, CA, 2012.</p> <p> </p> <p>They are derived from the original image data available at <a href="https://aloi.science.uva.nl/">https://aloi.science.uva.nl/</a></p> <p>The image acquisition process is documented in the original ALOI work: <em>J. M. Geusebroek, G. J. Burghouts, and A. W. M. Smeulders</em>, <strong>The Amsterdam library of object images</strong>, Int. J. Comput. Vision, 61(1), 103-112, January, 2005</p> <p>Additional information is available at: <a href="https://elki-project.github.io/datasets/multi_view">https://elki-project.github.io/datasets/multi_view</a></p> <p>The following views are currently available:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>Object number</td> <td>Sparse 1000 dimensional vectors that give the <em>true</em> object assignment</td> <td><a href="6355684/files/objs.arff.gz">objs.arff.gz</a></td> </tr> <tr> <td>RGB color histograms</td> <td>Standard RGB color histograms (uniform binning)</td> <td><a href="6355684/files/aloi-8d.csv.gz">aloi-8d.csv.gz</a> <a href="6355684/files/aloi-27d.csv.gz">aloi-27d.csv.gz</a> <a href="6355684/files/aloi-64d.csv.gz">aloi-64d.csv.gz</a> <a href="6355684/files/aloi-125d.csv.gz">aloi-125d.csv.gz</a> <a href="6355684/files/aloi-216d.csv.gz">aloi-216d.csv.gz</a> <a href="6355684/files/aloi-343d.csv.gz">aloi-343d.csv.gz</a> <a href="6355684/files/aloi-512d.csv.gz">aloi-512d.csv.gz</a> <a href="6355684/files/aloi-729d.csv.gz">aloi-729d.csv.gz</a> <a href="6355684/files/aloi-1000d.csv.gz">aloi-1000d.csv.gz</a></td> </tr> <tr> <td>HSV color histograms</td> <td>Standard HSV/HSB color histograms in various binnings</td> <td><a href="6355684/files/aloi-hsb-2x2x2.csv.gz">aloi-hsb-2x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-3x3x3.csv.gz">aloi-hsb-3x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-4x4x4.csv.gz">aloi-hsb-4x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-5x5x5.csv.gz">aloi-hsb-5x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-6x6x6.csv.gz">aloi-hsb-6x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-7x7x7.csv.gz">aloi-hsb-7x7x7.csv.gz</a> <a href="6355684/files/aloi-hsb-7x2x2.csv.gz">aloi-hsb-7x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-7x3x3.csv.gz">aloi-hsb-7x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-14x3x3.csv.gz">aloi-hsb-14x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-8x4x4.csv.gz">aloi-hsb-8x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-9x5x5.csv.gz">aloi-hsb-9x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-13x4x4.csv.gz">aloi-hsb-13x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-14x5x5.csv.gz">aloi-hsb-14x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-10x6x6.csv.gz">aloi-hsb-10x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-14x6x6.csv.gz">aloi-hsb-14x6x6.csv.gz</a></td> </tr> <tr> <td>Color similiarity</td> <td>Average similarity to 77 reference colors (not histograms) 18 colors x 2 sat x 2 bri + 5 grey values (incl. white, black)</td> <td><a href="6355684/files/aloi-colorsim77.arff.gz">aloi-colorsim77.arff.gz</a> (feature subsets are meaningful here, as these features are computed independently of each other)</td> </tr> <tr> <td>Haralick features</td> <td>First 13 Haralick features (radius 1 pixel)</td> <td><a href="6355684/files/aloi-haralick-1.csv.gz">aloi-haralick-1.csv.gz</a></td> </tr> <tr> <td>Front to back</td> <td>Vectors representing front face vs. back faces of individual objects</td> <td><a href="6355684/files/front.arff.gz">front.arff.gz</a></td> </tr> <tr> <td>Basic light</td> <td>Vectors indicating basic light situations</td> <td><a href="6355684/files/light.arff.gz">light.arff.gz</a></td> </tr> <tr> <td>Manual annotations</td> <td>Manually annotated object groups of semantically related objects such as cups</td> <td><a href="6355684/files/manual1.arff.gz">manual1.arff.gz</a></td> </tr> </tbody></table> <p><strong>Outlier Detection Versions</strong></p> <p>Additionally, we generated a number of subsets for outlier detection:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>RGB Histograms</td> <td>Downsampled to 100000 objects (553 outliers)</td> <td><a href="6355684/files/aloi-27d-100000-max10-tot553.csv.gz">aloi-27d-100000-max10-tot553.csv.gz</a> <a href="6355684/files/aloi-64d-100000-max10-tot553.csv.gz">aloi-64d-100000-max10-tot553.csv.gz</a></td> </tr> <tr> <td> </td> <td>Downsampled to 75000 objects (717 outliers)</td> <td><a href="6355684/files/aloi-27d-75000-max4-tot717.csv.gz">aloi-27d-75000-max4-tot717.csv.gz</a> <a href="6355684/files/aloi-64d-75000-max4-tot717.csv.gz">aloi-64d-75000-max4-tot717.csv.gz</a></td> </tr> <tr> <td> </td> <td>Downsampled to 50000 objects (1508 outliers)</td> <td><a href="6355684/files/aloi-27d-50000-max5-tot1508.csv.gz">aloi-27d-50000-max5-tot1508.csv.gz</a> <a href="6355684/files/aloi-64d-50000-max5-tot1508.csv.gz">aloi-64d-50000-max5-tot1508.csv.gz</a></td> </tr> </tbody></table>
Code and data set for data analysis published as manuscript "Bacttle: a microbiology educational board game for lay public and schools"
<p>Code that processed raw data and plots the figures of the manuscript "Bacttle: a microbiology educational board game for lay public and schools"</p> <p>Below is a table with the original survey questions. The ID corresponds to the column displayed on the data set. When letters are followed by a number (1 or 2), it means that the question was answered before playing the game (1) and after playing the game (2).</p> <table> <tbody> <tr> <td> <p><em>ID<sup>1</sup></em></p> </td> <td> <p><em>Question text</em></p> </td> <td> <p><em>Possible answers<sup>2</sup></em></p> </td> </tr> <tr> <td> <p><em>A</em></p> </td> <td> <p>How old are you?</p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p><em>B</em></p> </td> <td> <p>Do you know what a bacterium is?</p> </td> <td> <p>y/n</p> </td> </tr> <tr> <td> <p><em>C</em></p> </td> <td> <p>Do you know what a bacterial capsule is?</p> </td> <td> <p>y/n</p> </td> </tr> <tr> <td> <p><em>D</em></p> </td> <td> <p>Do bacteria have tools to harm each other?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>E</em></p> </td> <td> <p>Do bacteria reproduce at the same pace?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>F</em></p> </td> <td> <p>What is sporulation?</p> </td> <td> <p>A resistant state that some bacteria can achieve under unfavorable conditions.</p> </td> </tr> <tr> <td> <p>The release of toxins by bacteria.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>G</em></p> </td> <td> <p>What are flagella used for?</p> </td> <td> <p>Sticking to surfaces.</p> </td> </tr> <tr> <td> <p>Motility in liquid environments.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>H</em></p> </td> <td> <p>What does it mean to be lithotrophic?</p> </td> <td> <p>A bacterium can get energy from minerals.</p> </td> </tr> <tr> <td> <p>A bacterium can get energy from the sunlight.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>I</em></p> </td> <td> <p>Can bacteria be infected by viruses?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>J</em></p> </td> <td> <p>Are all bacteria harmful for humans?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>K</em></p> </td> <td> <p>How many bacteria are in a coffee spoon of yoghurt?</p> </td> <td> <p>Millions</p> </td> </tr> <tr> <td> <p>Hundreds</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>L</em></p> </td> <td> <p>How easy did you find the gameplay?</p> </td> <td> <p>VE/E/A/D/VD</p> </td> </tr> <tr> <td> <p><em>M</em></p> </td> <td> <p>Did you find the card content easy to understand?</p> </td> <td> <p>VE/E/A/D/VD</p> </td> </tr> <tr> <td> <p><em>N</em></p> </td> <td> <p>Did you like the setup of the game?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>O</em></p> </td> <td> <p>Would you like to play this game again?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>P</em></p> </td> <td> <p>What can we improve?</p> </td> <td> <p> </p> </td> </tr> </tbody> </table> <p>1) Question A categorizes the player’s age; B and C assess the initial level of knowledge in microbiology (none -both questions are answered negatively-, basic -player knows what a bacterium is but not a bacterial capsule-, or advanced -both answers are positive-); questions D-I score knowledge acquisition; J and K are control questions; L-O evaluate the appreciation of the game; and P is an optional free text-entry answer for additional feedback. <br>2) y= yes, n=no, idk=I don’t know, VE=very easy, E=easy, A=adequate, D=difficult, VD=very difficult.</p>
Iceland as stepping stone for intercontinental spread of highly pathogenic avian influenza H5N1 virus between Europe and North America: data set on phylogeographic analysis
<p>Highly pathogenic avian influenza viruses (HPAIV) subtype H5 clade 2.3.4.4b have widely spread within the northern hemisphere since 2020 and threaten wild bird populations as well as poultry production. For the very first time, HPAIV were detected in wild birds and, subsequently, in poultry holdings in Iceland.</p> <p>Here, we present phylogeographic evidence that Iceland has been used as a stepping stone for HPAIV translocation from Northern Europe to North America in 2021 and describe two independent incursions of HPAI H5N1 clade 2.3.4.4b viruses of two different genotypes to Iceland in 2021 and 2022.</p>
Quantifying the Sensitivity of Sea Level Change in Coastal Localities to the Geometry of Polar Ice Mass Flux -- Supplemental Data Set: Sea Level Sensitivity Kernels
<p><strong>Quantifying the Sensitivity of Sea Level Change in Coastal Localities to the Geometry of Polar Ice Mass Flux<br> SUPPLEMENTAL DATA SET: SEA LEVEL SENSITIVITY KERNELS</strong></p> <p>To accompany</p> <p> Jerry X. Mitrovica, Carling C. Hay, Robert E. Kopp, Christopher Harig, and<br> Konstantin Laytchev (2018). Quantifying the Sensitivity of Sea Level Change<br> in Coastal Localities to the Geometry of Polar Ice Mass Flux. Journal of<br> Climate. doi: 10.1175/JCLI-D-17-0465.1.</p> <p>We provide sea level kernels for ~740 tide gauge sites in the Permanent Service for Mean Sea Level (PSMSL) database (Holgate et al., 2013). Kernels associated with sensitivities to Greenland and Alaskan glacier melt are given on a spatial grid covering the globe, with 512 latitude rows (i=1,512) and 1024 longitude (j=1,1024) columns.</p> <p>Longitude values are evenly spaced moving eastward from Greenwich (the jth grid point has an east longitude value of (j-1)×360°/1024). Latitude values are Gauss-Legendre points beginning close to the North Pole and ending near the South Pole. Kernels associated with sensitivities to Antarctic melt are given on a spatial grid covering the globe, with 256 (Gauss-Legendre) latitude rows (i=1,256) and 512 longitude (j=1,512) columns. Longitude values are evenly spaced moving eastward from Greenwich.</p> <p>The format of the files is: </p> <p> grid_sitenumber_region.txt</p> <p>where “region” is either “green” (Greenland), “ant” (Antarctic) or “Alaska” (Alaska). The list of sites (and site numbers) is provided in the sites.txt file. The first 8 sites in this list were test sites and can be ignored.</p>
Unsteady Aerodynamics Open Data Set
<p>A selection of four different unsteady aerodynamic experiments have been done to prepare a database which will serve for the analysis, investigation and tool validation of airfoil unsteady behavior of wind turbine blades.<br> The four experiments and selected data are:</p> <ul> <li>University of Glasgow dynamic stall experiments: NACA0015 and NACA0030 airfoils tested at sinusoidal type motion of the pitch.</li> <li>NREL OSU experiments: LS(1)0417MOD, NACA4415 and S809 airfoils tested at sinusoidal type motion of the pitch.</li> <li>CENER unsteady airfoil pitching and flapping tests at DTU: NACA643-418 airfoil tested at sinusoidal type motion of the pitch, the flap and combined pitch and flap.</li> <li>ForWind airfoil tests under tailored inflow turbulence: DU00W212 airfoil with laminar flow, open grid condition and one sinusoidal dynamic grid condition.</li> </ul>
S42 | HDXNOEX | Hydrogen Deuterium Exchange (HDX) Standard Set
<p>This is the collection associated with list S42 HDXNOEX on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S42</p> <p>HDXNOEX</p> <p><strong>Hydrogen Deuterium Exchange (HDX) Standard Set</strong></p> <p>HDXNOEX <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/120219Update/HDXNOEX_14022019.xlsx">XLSX</a>, <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/120219Update/HDXNOEX_14022019.csv">CSV</a> (14/02/2019)<br> CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/hdxnoex">HDXNOEX List</a><br> CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/hdxexch">HDXEXCH List</a></p> <p>HDXNOEX <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/120219Update/HDXNOEX_InChIKeys_14022019.txt">InChIKeys</a> (14/02/2019)</p> <p>Environmental standard set used to investigate hydrogen deuterium exchange in small molecule HRMS (Ruttkies et al. accepted). <a href="https://comptox.epa.gov/dashboard/chemical_lists/hdxexch">HDXEXCH</a> list also contains observed deuterated species. </p>
Data set of the manuscript titled: Follicular Immune Landscaping Reveals a distinct profile of FOXP3hi CD4+ T cells in Treated compared to Untreated HIV
<p>Multiplex imaging data were collected using a scanning confocal system (STELARIS, Leica) and proccessed with the Imaris and Fiji imaging programs. csv files incuding the position identifiers and intensities for each fluorochrome used were generated and data were further analysed using the FlowJo10 program. Neighboring analysis was performed using the G function and mean of minimum distances of relevant cell type pairs. </p>
PE-HRI-temporal: A Multimodal Temporal Dataset in a robot mediated Collaborative Educational Setting
<p><em><strong>Please note that this dataset corresponds to the training data used in "Social robots as skilled ignorant peers for supporting learning "[7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint). </strong></em></p> <p> </p> <p>This data set consists of <strong>multi-modal temporal team behaviors as well as learning outcomes </strong>collected in the context of a robot mediated collaborative and constructivist learning activity called JUSThink [1,2]. The data set can be useful for those looking to explore evolution of log actions, speech behavior, affective states, and gaze patterns for students to model constructs such as engagement, motivation, collaboration, etc. in educational settings. </p> <p>In this data set, team level data is collected from 34 teams of two (68 children) where the children are aged between 9 and 12. There are two files: </p> <p><strong>PE-HRI_learning_and_performance.csv:</strong> This file consists of the <strong>team level performance and learning metrics</strong> which are defined below: </p> <ul> <li> <p><em>last_error:</em> This is the error of the last submitted solution. Note that if a team has found an optimal solution (error = 0) the game stops, therefore making last error = 0. This is a metric for performance in the task. </p> </li> <li> <p><em>T_LG_absolute:</em> It is a team-level learning outcome that we calculate by taking the average of the two individual absolute learning gains of the team members. The individual absolute gain is the difference between a participant’s post-test and pre-test score, divided by the maximum score that can be achieved (10), which grasps how much the participant learned of all the knowledge available.</p> </li> <li> <p><em>T_LG_relative:</em> It is a team-level learning outcome that we calculate by taking the average of the two individual relative learning gains of the team members. The individual relative gain is the difference between a participant’s post-test and pre-test score, divided by the difference between the maximum score that can be achieved and the pre-test score. This grasps how much the participant learned of the knowledge that he/she didn’t possess before the activity. </p> </li> <li> <p><em>T_LG_joint_abs: </em>It is a team-level learning outcome defined as the difference between the number of questions that both of the team members answer correctly in the post-test and in the pre-test, which grasps the amount of knowledge acquired together by the team members during the activity</p> </li> </ul> <p><strong>PE-HRI_behavioral_timeseries_w_labels.csv:</strong> In this file, for each team, the interaction of around 20-25 minutes is organized in windows of 10 seconds; hence, we have a total of 5048 windows of 10 seconds each. We report team level log actions, speech behavior, affective states, and gaze patterns for each window. More specifically, within each window, 26 features are generated in two ways: </p> <ol> <li>non-incremental</li> <li>incremental</li> </ol> <p>A non-incremental type would mean the value of a feature <em>in</em> that particular time window while an incremental type would mean the value of a feature <em>until</em> that particular time window. The incremental type is indicated by an "_inc" at the end of the feature name. Hence, in the end, within each window, we have 52 values: </p> <ul> <li> <p><em>T_add/(_inc): </em>The number of times a team added an edge on the map in that window/(until that window).</p> </li> <li> <p><em>T_remove/(_inc): </em>The number of times a team removed an edge from the map in that window/(until that window).</p> </li> <li> <p><em>T_ratio_add_rem/(_inc): </em>The ratio of addition of edges over deletion of edges by a team in that window/(until that window).</p> </li> <li> <p><em>T_action/(_inc):</em> The total number of actions taken by a team (add, delete, submit, presses on the screen) in that window/(until that window).</p> </li> <li> <p><em>T_hist/(_inc): </em>The number of times a team opened the sub-window with history of their previous solutions in that window/(until that window).</p> </li> <li> <p><em>T_help/(_inc): </em>The number of times a team opened the instructions manual in that window/(until that window). Please note that the robot initially gives all the instructions before the game-play while a video is played for demonstration of the functionality of the game. </p> </li> <li> <p><em>T1_T1_rem/(_inc): </em>The number of times either of the two members in the team followed the pattern consecutively: I add an edge, I then delete it in that window/(until that window).</p> </li> <li> <p><em>T1_T1_add/(_inc): </em>The number of times either of the two members in the team followed the pattern consecutively: I delete an edge, I add it back in that window/(until that window).</p> </li> <li> <p><em>T1_T2_rem/(_inc): </em>The number of times the members of the team followed the pattern consecutively: I add an edge, you then delete it in that window/(until that window).</p> </li> <li> <p><em>T1_T2_add/(_inc): </em>The number of times the members of the team followed the pattern consecutively: I delete an edge, you add it back in that window/(until that window).</p> </li> <li> <p><em>redundant_exist/(_inc): </em>The number of times the team had redundant edges in their map in that window/(until that window).</p> </li> <li> <p><em>positive_valence/(_inc): </em>The average value of positive valence for the team in that window/(until that window).</p> </li> <li> <p><em>negative_valence/(_inc): </em>The average value of negative valence for the team in that window/(until that window).</p> </li> <li> <p><em>difference_in_valence/(_inc): </em>The difference of the average value of positive and negative valence for the team in that window/(until that window).</p> </li> <li> <p><em>arousal/(_inc): </em>The average value of arousal for the team in that window/(until that window).</p> </li> <li> <p><em>gaze_at_partner/(_inc): </em>The average of the the two team member's gaze when looking at their partner in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_robot/(_inc): </em>The average of the the two team member's gaze when looking at the robot in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_other/(_inc): </em>The average of the the two team member's gaze when looking in the direction opposite to the robot in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_screen_left/(_inc): </em>The average of the the two team member's gaze when looking at the left side of the screen in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>gaze_at_screen_right/(_inc):</em> The average of the the two team member's gaze when looking at the right side of the screen in that window/(until that window). Each individual member's gaze is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_speech_activity/(_inc): </em>The average of the two team member's speech activity in that window/(until that window). Each individual member's speech activity is calculated as a percentage of time that they are speaking in that window/(until that window). </p> </li> <li> <p><em>T_silence/(_inc): </em>The average of the two team member's silence in that window/(until that window). Each individual member's silence is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_short_pauses/(_inc): </em>The average of the two team member's short pauses over their speech activity in that window/(until that window). Each individual member's short pause refers to a brief pause of 0.15 seconds and is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_long_pauses/(_inc): </em>The average of the two team members long pauses over their speech activity in that window/(until that window). Each individual member's long pause refers to a pause of 1.5 seconds and is calculated as a percentage of time in that window/(until that window). </p> </li> <li> <p><em>T_overlap/(_inc): </em>The average percentage of time the speech of the team members overlaps in that window/(until that window).</p> </li> <li> <p><em>T_overlap_to_speech_ratio/(_inc): </em>The ratio of the speech overlap over the speech activity of the team in that window/(until that window).</p> </li> </ul> <p>Apart from these 52 values, within each window, we also indicate: </p> <ul> <li><em>team: </em>The team to which the window belongs to.</li> <li><em>time_in_secs:</em> Time in seconds until that window.</li> <li><em>window: </em>The window number.</li> <li><em>normalized_time: </em>The time when this window occurred with respect to the total duration of the task for a particular team. </li> <li>cluster_labels: The cluster number associated with each time window in reference to the productive and non-productive clusters found in [3]</li> <li>PE_score: The Productive Engagement score in each window</li> </ul> <p>Lastly, we briefly elaborate on how the features are operationalised. We extract log behaviors from the recorded rosbags while the behaviors related to both gaze and affective states are computed through the open source library OpenFace [6] that returns both facial actions units (AUs) as well as gaze angles. For voice activity detection (VAD), that classifies if a piece of audio is voiced or unvoiced, we made use of the python wrapper for the open source Google WebRTC VAD. The literature that inspired our log, audio and video features as well as the tools used to extract them are described in more detail in [3,4]. However, in those papers, we make use of only the aggregate version of this data [5].</p> <p><em><strong>Please note that this dataset corresponds to the training data used in [7]. This (second) version of the dataset additionally includes labels (PE score and cluster labels for each datapoint). </strong></em></p>
Exploring AdaBoost and Random Forests machine learning approaches for infrared pathology on unbalanced data sets
<p>The use of infrared spectroscopy to augment decision-making in histopathology is a promising direction for the diagnosis of many disease types. Hyperspectral images of healthy and diseased tissue, generated by infrared spectroscopy, are used to build chemometric models that can provide objective metrics of disease state. It is important to build robust and stable models to provide confidence to the end user. The data used to develop such models can have a variety of characteristics which can pose problems to many model-building approaches. Here we have compared the performance of two machine learning algorithms – AdaBoost and Random Forests – on a variety of non-uniform data sets. Using samples of breast cancer tissue, we devised a range of training data capable of describing the problem space. Models were constructed from these training sets and their characteristics compared. In terms of separating infrared spectra of cancerous epithelium tissue from normal-associated tissue on the tissue microarray, both AdaBoost and Random Forests algorithms were shown to give excellent classification performance (over 95% accuracy) in this study. AdaBoost models were more robust when datasets with large imbalance were provided. The outcomes of this work are a measure of classification accuracy as a function of training data available, and a clear recommendation for choice of machine learning approach.</p>
Surface water and flooding dynamics data set based on seasonally continuous Landsat data (1986-2011) in a dryland river basin
<p>Animations of the data are available here: <a href="https://doi.org/10.5281/zenodo.2438110">https://doi.org/10.5281/zenodo.2438110</a></p> <p>If you are using this data set, please cite the following publication:</p> <p>Tulbure, M.G. and M. Broich (2018). Spatiotemporal patterns and effects of climate and land use on surface water extent dynamics in a dryland region with three decades of Landsat satellite data. Science of the Total Environment. https://www.sciencedirect.com/science/article/pii/S0048969718347466 </p> <p>The data represent statistically validated surface water and flooding extent dynamics derived from seasonally continous Landsat TM/ETM+ data and random forest models, and summarised to the maximum extent of surface water per season between 1986-2011 over Australia's Murray-Darling Basin. The overall accuracy was over 99% and producer's accuracy for water 87% +/- 3%. </p> <p>The method is described in the following publication: <br> Tulbure, M.G., M. Broich, S.V. Stehman, A. Kommareddy. (2016). Surface water extent dynamics from three decades of seasonally continuous Landsat time series at subcontinental scale in a semi-arid region. Remote Sensing of Environment. 178: 142-157</p> <p>URL: https://www.sciencedirect.com/science/article/pii/S0034425716300621 </p> <p>Data are provided in GeoTIFF format per season per year. File naming convention is as follows:<br> yy_inund_freq_season_SamplingMethod. For example, "99_inund_freq_winter_max" will represent inundation frequency for winter 1999 resampled using a maximum resampling method. </p> <p>Inundation frequency represents the number of times a pixel has been flagged as flooded out of the times that pixel had valid observations * 100. Valid observation exclude no data values and clouds. The valid range of inundation frequency is 0-100 [%], with 255 indicating no data values. Data type is eight bit unsigned integer (uint8). </p> <p>The data were resampled to 120m resolution to reduce file size. The resampling methods used include max (e.g. selects the max value of all non-NODATA contributing 30m pixels) and mean (median and min can be provided upon request). If you are unsure which resampling to use, you may want to start with the mean. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.