Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
PIE LTER surface elevation table (SET) pin height data from twelve marsh sites in northeast Massachusetts.
Surface elevation table (SET) measurements from 26 SETs at 9 marsh sites in the Plum Island Sound Long-Term Ecological Research Site in the Great Marsh, Massachusetts. SET measurements are useful for determining the relative elevation change of marsh sediments. Precise measurements of sediment elevation in marshes is useful for determining rates of elevation change in response to changes in sea level.
Biogeochemistry data set for Imnavait Creek Weir on the North Slope of Alaska 2002-2024.
Data file containing biogeochemical data of water samples collected in Imnavait Creek, North Slope of Alaska. Sample site descriptors include a unique assigned number (sortchem), site, date, time, depth, distance (downstream), and elevation. Values of variables measured in the field include temperature, conductivity, pH. Chemical analysis for samples include alkalinity, dissolved organic carbon (DOC), inorganic and total dissolved nutrients particulate carbon, nitrogen, and phosphorus, cations and anions.
Biogeochemistry data set for soil waters, streams, and lakes near Toolik Lake on the North Slope of Alaska, 2012 through 2020
Data file of the biogeochemistry of samples collected at various sites near Toolik Lake, North Slope of Alaska. Sample site descriptors include a unique assigned number (sortchem), site, date, time, depth, distance (downstream from a reference location), elevation, treatment, date-time, category, and water type (lake, surface, soil). Physical measures collected in the field include temperature (water, soil, well water), conductivity, pH, and average thaw depth in soil. Chemical analyses for the sample include alkalinity; dissolved inorganic and organic carbon (DIC and DOC); dissolved gases CO2 and CH4; inorganic and total dissolved nutrients (NH4, PO4, NO3, TDN, TDP); particulate carbon, nitrogen, and phosphorus (PC, PN, and PP); cations (Ca, Mg, Na, K, and Si); and anions (SO4 and Cl).
The NANOGrav 12.5-year Wideband Data Set (version 12yv4)
<p>The NANOGrav 12.5-year wideband data set (public release "12yv4") is the supplemental data set accompanying Alam et al. 2021, "The NANOGrav 12.5 yr Data Set: Wideband Timing of 47 Millisecond Pulsars," The Astrophysical Journal Supplement Series, 252, 5, DOI 10.3847/1538-4365/abc6a1. It contains wideband pulse times of arrival, models describing frequency-dependent template profiles, pulsar timing models, timing residuals, and clock files.</p> <p>Details about the contents of these files are contained in NANOGrav_12yv4_wideband/README, as well as in NANOGrav_12yv4_wideband/wideband/README.wideband. The narrowband version of this dataset (published in Alam et al. 2021, ApJS, 252, 4, DOI: 10.3847/1538-4365/abc6a0) can be found at Zenodo DOI: 10.5281/zenodo.4312297. Both the narrowband and wideband data sets are also available at <a href="http://data.nanograv.org">data.nanograv.org</a>.</p>
An updated mass-radius analysis of the 2017-2018 NICER data set of PSR J0030+0451
<p>Summarised posterior sample files associated with the preprint "An updated mass-radius analysis of the 2017-2018 NICER data set of PSR J0030+0451" by Vinciguerra et al. (2023; <a href="https://doi.org/10.48550/arXiv.2308.09469">arXiv</a>; accepted for publication in ApJ).</p><p>Also included are examples of model modules in the Python language using the X-PSI framework; and Jupyter analysis notebooks.</p><p>Please refer to the READme for detailed information.</p>
Data set for the journal article ''Nanoscale chemical reaction exploration with a quantum magnifying glass''
<div>This data set includes the raw data of the esterification and hydrogenation discussed in the journal article alongside with the Scine Puffin Singularity container, steering protocol files, Swoose parameters, (pre-)releases of the software, and Python scripts for individual steps without the graphical user interface to reproduce the data.</div>
WILLOW - Norther: data set for the full-scale validation of model-based virtual sensing methods for an operational offshore wind turbine
<h1><em><strong>1. General description </strong></em></h1> <p>This data set contains as-build design information, as well as full-scale vibration response measurements from an operational offshore wind-turbine. The turbine is part of the Norther wind farm which is located in the Belgian North Sea<em> </em>and includes a total of 44 Vestas V164 (8.4MW) wind turbines on monopile foundations, see <a href="../api/records/11093262/draft/files/Fig1_Norther_locaction.png/content" target="_blank" rel="noopener noreferrer">Fig1_Norther_locaction.png</a>. This data set is intended to verify and validate model-based virtual sensing algorithms, using data as well as modeling information from a real turbine. </p> <h2><em><strong>1.1 Summary of the shared structural information</strong></em></h2> <p>The included information entails a detailed description of the geometric properties of the monopile and transition piece, distributed and lumped structural masses . All information shared in this record is conform the as-designed documentation. An example of the lumped masses considered in the model input files is presented in "<a href="../api/records/11093262/draft/files/Fig2_Sensor_Network.png/content" target="_blank" rel="noopener">Fig2_Sensor_Network.png"</a></p> <h2><em><strong>1.2 Summary of the shared geotechnical information</strong></em></h2> <p>Monopiles are distinguished by the significant role of soil-structure interaction. Ground reaction is most typically included in the structural model as non-linear p-y curves. Different p-y curves are available for a certain number of soils in the standards applicable to offshore structures (API RP 2GEO, 2011, and ISO 19901-4:2016(E), 2016).</p> <p>The required soil properties to define p-y curves according to the API framework are given in the soil profile provided in a separate Excel. Rather than symbols, the name of the soil properties is generally used as column header (e.g., <em>Undrained shear strength</em>). Therefore, it is straightforward to identify each soil parameter. The only soil parameter that might lead to confusion is:</p> <ul> <li><em>"epsilon50 [-]" </em>represents the vertical strain at half the maximum principal stress difference in a static undrained triaxial compression test on an undisturbed soil sample.</li> </ul> <p>It's worthy to note that estimates for the small shear strain stiffness, referred to as Gmax, are also included. Despite not being required as an input to define the API p-y curves, this parameter remains a key input for other soil reaction frameworks than the API (e.g., PISA). </p> <h2><em><strong>1.3 Summary of the shared measurement data</strong></em></h2> <p>Two sets of measurement data have been curated for validation purposes; the first interval has been collected during parked conditions, whereas the second interval has been collected during rated operational conditions. Both records have a length of 2 hours, and are subdivided into 10-minute data sets. Furthermore 1Hz SCADA data has been made available for the selected intervals. All different data sources are time synchronized and have been subjected to several internal quality checks. </p> <p>The sensor network on NRT-WTG is illustrated in in <strong>Fig. 2, </strong>whereas a description of the sensor types is presented in <strong>Tab.1.</strong> The acceleration sensors are installed in the horizontal plane, and measure tangential (Y) and orthogonal (X) to the wall, where the positive Y direction is pointing clockwise and the positive X direction is pointing inwards. All strain sensors are installed vertically and are located on the inside of the wall.</p> <table> <tbody> <tr> <td><strong>Data type </strong></td> <td><strong>Sensor type</strong></td> <td><strong>Fs (Hz)</strong></td> <td> <p><strong>Level mLAT (m)</strong></p> </td> <td><strong>Description </strong></td> </tr> <tr> <td>Acceleration (g) </td> <td>Piezo-electric acc. sensor (<strong>ACC</strong>)</td> <td>30</td> <td>15, 69, 97 </td> <td>3 Bi-directional accelerometers at different levels. LAT 15 installed at 240 degree heading; LAT 69 and 97 at 60 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Resistive strain gauge (<strong>SG</strong>)</td> <td>30</td> <td>14</td> <td>6 SGs: equally spaced around the inner circumference of the can. Headings: 50, 110, 170, 230, 290, 350 degree.</td> </tr> <tr> <td>Strain (micro strain)</td> <td>Fiber-Bragg Grating strain gauge (<strong>FBG</strong>)</td> <td>100</td> <td>-17, -19</td> <td>2 FBGs per level at 165 and 255 degree respectively.</td> </tr> </tbody> </table> <p><strong>Table 1. Description of sensor types.</strong></p> <p>The FBG strain time series have been synchronized with the SG time series using using a cross-correlation based approach. Therefore the SG data has been used to genereate refrence strain time series at the headings of the FBG sensors; the FBG data is subsequently synchronized with regard to this reference time series. No synchronization of the acceleration data was needed, since these are collected using the same data aquisition system as the SG data. </p> <p>The SG strain time series have been calibrated and temperature compensated, whereas this is not the case for the FBG strain time series. The latter have a yet to be determined calibration offset. </p> <p>In conjunction to the sensor channels presented in <strong>Tab. 1</strong>, 1 Hz SCADA data is provided. A summary of the provided SCADA parameters, all sampled at 1Hz, is presented in <strong>Tab 2.</strong></p> <table> <tbody> <tr> <td><strong>Parameter</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>Wind speed</td> <td>m/s</td> <td>Wind speed as recorded in the turbine SCADA</td> </tr> <tr> <td>Wind direction</td> <td>°</td> <td>Wind direction relative to North (0°) as recorded in the turbine SCADA</td> </tr> <tr> <td>Yaw angle</td> <td>°</td> <td>Yaw orientation of the nacelle relative to North (0°) as recorded in the turbine SCADA</td> </tr> <tr> <td>Pitch angle</td> <td>°</td> <td>Rotor blade pitch as recorded in the turbine SCADA</td> </tr> <tr> <td>Rotor speed</td> <td>rpm</td> <td>Rotor speed in rotations per minute as recorded in the turbine SCADA</td> </tr> <tr> <td>Power</td> <td>kW</td> <td>Active power of the turbine as recorded in the turbine SCADA</td> </tr> </tbody> </table> <p><strong>Table 2. </strong>List of provided SCADA parameters</p> <p> </p> <p>A summary of the selected intervals and relevant corresponding scada parameters is given in <strong>Tab 3</strong>.</p> <table> <tbody> <tr> <td><strong>Scenario </strong></td> <td><strong>T1 (UTC)</strong></td> <td><strong>T2 (UTC) </strong></td> <td><strong>Windspeed</strong></td> <td><strong>RPM </strong></td> <td><strong>Pitch </strong></td> </tr> <tr> <td>Parked</td> <td> <p>03/07 01:30</p> </td> <td> <p>03/07 03:30</p> </td> <td>< 4.5 m/s</td> <td>~1</td> <td>~18 °</td> </tr> <tr> <td>Rated</td> <td> <p>05/07 22:30</p> </td> <td> <p>06/07 00:30 </p> </td> <td>~15 m/s</td> <td>10.5</td> <td>8.1°</td> </tr> </tbody> </table> <p><strong>Table 3. </strong>Selected data intervals and relevant scada parameters</p> <p> </p> <h1><em><strong>2. Included in this version </strong></em></h1> <h2><em><strong>2.1 Version - 0.1.0</strong></em></h2> <ul> <li>Relevant Design information can be found in: <ul> <li>Geometry data for NRT-WTG: "WILLOW-Geometry_v4.xlsx"</li> <li>Best estimate soil profile: "WILLOW-BE_soil_profile.xlsx"</li> </ul> </li> <li>Acceleration, strain and scada data can be found in the following parquet files: <ul> <li>Measurement data for the parked case: "NRT-WTG_Parked.parquet.gz"</li> <li>Measurement data for the rated case: "NRT-WTG_Rated.parquet.gz"</li> </ul> </li> </ul> <p> </p> <h1><em><strong>3. Importing parquet files </strong></em></h1> <p>To import the measurement data into Python it is recommended to use pandas:</p> <pre>import pandas as pd<br># Read Parquet file with Pandas: relative_file_path = '<a href="../api/records/11093262/draft/files/NRT-WTG_Parked.parquet.gz/content" target="_blank" rel="noopener noreferrer">NRT-WTG_Parked.parquet.gz</a>' data = pd.read_parquet(relative_file_path ) <br><br>Once the dataframe has been imported, the users can process/re-arrange the raw data according the their needs; it should be noted that the imported dataframe contains NAN values - these are caused by the different sampling rates of the provided signals. </pre>
ELKI Multi-View Clustering Data Sets Based on the Amsterdam Library of Object Images (ALOI)
<p>These data sets were originally created for the following publications:</p> <p><em>M. E. Houle, H.-P. Kriegel, P. Kröger, E. Schubert, A. Zimek</em><br> <strong>Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?</strong><br> In Proceedings of the 22nd International Conference on Scientific and Statistical Database Management (SSDBM), Heidelberg, Germany, 2010.</p> <p><em>H.-P. Kriegel, E. Schubert, A. Zimek</em><br> <strong>Evaluation of Multiple Clustering Solutions</strong><br> In 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, 2011.</p> <p>The outlier data set versions were introduced in:</p> <p><em>E. Schubert, R. Wojdanowski, A. Zimek, H.-P. Kriegel</em><br> <strong>On Evaluation of Outlier Rankings and Outlier Scores</strong><br> In Proceedings of the 12th SIAM International Conference on Data Mining (SDM), Anaheim, CA, 2012.</p> <p> </p> <p>They are derived from the original image data available at <a href="https://aloi.science.uva.nl/">https://aloi.science.uva.nl/</a></p> <p>The image acquisition process is documented in the original ALOI work: <em>J. M. Geusebroek, G. J. Burghouts, and A. W. M. Smeulders</em>, <strong>The Amsterdam library of object images</strong>, Int. J. Comput. Vision, 61(1), 103-112, January, 2005</p> <p>Additional information is available at: <a href="https://elki-project.github.io/datasets/multi_view">https://elki-project.github.io/datasets/multi_view</a></p> <p>The following views are currently available:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>Object number</td> <td>Sparse 1000 dimensional vectors that give the <em>true</em> object assignment</td> <td><a href="6355684/files/objs.arff.gz">objs.arff.gz</a></td> </tr> <tr> <td>RGB color histograms</td> <td>Standard RGB color histograms (uniform binning)</td> <td><a href="6355684/files/aloi-8d.csv.gz">aloi-8d.csv.gz</a> <a href="6355684/files/aloi-27d.csv.gz">aloi-27d.csv.gz</a> <a href="6355684/files/aloi-64d.csv.gz">aloi-64d.csv.gz</a> <a href="6355684/files/aloi-125d.csv.gz">aloi-125d.csv.gz</a> <a href="6355684/files/aloi-216d.csv.gz">aloi-216d.csv.gz</a> <a href="6355684/files/aloi-343d.csv.gz">aloi-343d.csv.gz</a> <a href="6355684/files/aloi-512d.csv.gz">aloi-512d.csv.gz</a> <a href="6355684/files/aloi-729d.csv.gz">aloi-729d.csv.gz</a> <a href="6355684/files/aloi-1000d.csv.gz">aloi-1000d.csv.gz</a></td> </tr> <tr> <td>HSV color histograms</td> <td>Standard HSV/HSB color histograms in various binnings</td> <td><a href="6355684/files/aloi-hsb-2x2x2.csv.gz">aloi-hsb-2x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-3x3x3.csv.gz">aloi-hsb-3x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-4x4x4.csv.gz">aloi-hsb-4x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-5x5x5.csv.gz">aloi-hsb-5x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-6x6x6.csv.gz">aloi-hsb-6x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-7x7x7.csv.gz">aloi-hsb-7x7x7.csv.gz</a> <a href="6355684/files/aloi-hsb-7x2x2.csv.gz">aloi-hsb-7x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-7x3x3.csv.gz">aloi-hsb-7x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-14x3x3.csv.gz">aloi-hsb-14x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-8x4x4.csv.gz">aloi-hsb-8x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-9x5x5.csv.gz">aloi-hsb-9x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-13x4x4.csv.gz">aloi-hsb-13x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-14x5x5.csv.gz">aloi-hsb-14x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-10x6x6.csv.gz">aloi-hsb-10x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-14x6x6.csv.gz">aloi-hsb-14x6x6.csv.gz</a></td> </tr> <tr> <td>Color similiarity</td> <td>Average similarity to 77 reference colors (not histograms) 18 colors x 2 sat x 2 bri + 5 grey values (incl. white, black)</td> <td><a href="6355684/files/aloi-colorsim77.arff.gz">aloi-colorsim77.arff.gz</a> (feature subsets are meaningful here, as these features are computed independently of each other)</td> </tr> <tr> <td>Haralick features</td> <td>First 13 Haralick features (radius 1 pixel)</td> <td><a href="6355684/files/aloi-haralick-1.csv.gz">aloi-haralick-1.csv.gz</a></td> </tr> <tr> <td>Front to back</td> <td>Vectors representing front face vs. back faces of individual objects</td> <td><a href="6355684/files/front.arff.gz">front.arff.gz</a></td> </tr> <tr> <td>Basic light</td> <td>Vectors indicating basic light situations</td> <td><a href="6355684/files/light.arff.gz">light.arff.gz</a></td> </tr> <tr> <td>Manual annotations</td> <td>Manually annotated object groups of semantically related objects such as cups</td> <td><a href="6355684/files/manual1.arff.gz">manual1.arff.gz</a></td> </tr> </tbody></table> <p><strong>Outlier Detection Versions</strong></p> <p>Additionally, we generated a number of subsets for outlier detection:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>RGB Histograms</td> <td>Downsampled to 100000 objects (553 outliers)</td> <td><a href="6355684/files/aloi-27d-100000-max10-tot553.csv.gz">aloi-27d-100000-max10-tot553.csv.gz</a> <a href="6355684/files/aloi-64d-100000-max10-tot553.csv.gz">aloi-64d-100000-max10-tot553.csv.gz</a></td> </tr> <tr> <td> </td> <td>Downsampled to 75000 objects (717 outliers)</td> <td><a href="6355684/files/aloi-27d-75000-max4-tot717.csv.gz">aloi-27d-75000-max4-tot717.csv.gz</a> <a href="6355684/files/aloi-64d-75000-max4-tot717.csv.gz">aloi-64d-75000-max4-tot717.csv.gz</a></td> </tr> <tr> <td> </td> <td>Downsampled to 50000 objects (1508 outliers)</td> <td><a href="6355684/files/aloi-27d-50000-max5-tot1508.csv.gz">aloi-27d-50000-max5-tot1508.csv.gz</a> <a href="6355684/files/aloi-64d-50000-max5-tot1508.csv.gz">aloi-64d-50000-max5-tot1508.csv.gz</a></td> </tr> </tbody></table>
Code and data set for data analysis published as manuscript "Bacttle: a microbiology educational board game for lay public and schools"
<p>Code that processed raw data and plots the figures of the manuscript "Bacttle: a microbiology educational board game for lay public and schools"</p> <p>Below is a table with the original survey questions. The ID corresponds to the column displayed on the data set. When letters are followed by a number (1 or 2), it means that the question was answered before playing the game (1) and after playing the game (2).</p> <table> <tbody> <tr> <td> <p><em>ID<sup>1</sup></em></p> </td> <td> <p><em>Question text</em></p> </td> <td> <p><em>Possible answers<sup>2</sup></em></p> </td> </tr> <tr> <td> <p><em>A</em></p> </td> <td> <p>How old are you?</p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p><em>B</em></p> </td> <td> <p>Do you know what a bacterium is?</p> </td> <td> <p>y/n</p> </td> </tr> <tr> <td> <p><em>C</em></p> </td> <td> <p>Do you know what a bacterial capsule is?</p> </td> <td> <p>y/n</p> </td> </tr> <tr> <td> <p><em>D</em></p> </td> <td> <p>Do bacteria have tools to harm each other?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>E</em></p> </td> <td> <p>Do bacteria reproduce at the same pace?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>F</em></p> </td> <td> <p>What is sporulation?</p> </td> <td> <p>A resistant state that some bacteria can achieve under unfavorable conditions.</p> </td> </tr> <tr> <td> <p>The release of toxins by bacteria.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>G</em></p> </td> <td> <p>What are flagella used for?</p> </td> <td> <p>Sticking to surfaces.</p> </td> </tr> <tr> <td> <p>Motility in liquid environments.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>H</em></p> </td> <td> <p>What does it mean to be lithotrophic?</p> </td> <td> <p>A bacterium can get energy from minerals.</p> </td> </tr> <tr> <td> <p>A bacterium can get energy from the sunlight.</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>I</em></p> </td> <td> <p>Can bacteria be infected by viruses?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>J</em></p> </td> <td> <p>Are all bacteria harmful for humans?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>K</em></p> </td> <td> <p>How many bacteria are in a coffee spoon of yoghurt?</p> </td> <td> <p>Millions</p> </td> </tr> <tr> <td> <p>Hundreds</p> </td> </tr> <tr> <td> <p>idk</p> </td> </tr> <tr> <td> <p><em>L</em></p> </td> <td> <p>How easy did you find the gameplay?</p> </td> <td> <p>VE/E/A/D/VD</p> </td> </tr> <tr> <td> <p><em>M</em></p> </td> <td> <p>Did you find the card content easy to understand?</p> </td> <td> <p>VE/E/A/D/VD</p> </td> </tr> <tr> <td> <p><em>N</em></p> </td> <td> <p>Did you like the setup of the game?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>O</em></p> </td> <td> <p>Would you like to play this game again?</p> </td> <td> <p>y/n/idk</p> </td> </tr> <tr> <td> <p><em>P</em></p> </td> <td> <p>What can we improve?</p> </td> <td> <p> </p> </td> </tr> </tbody> </table> <p>1) Question A categorizes the player’s age; B and C assess the initial level of knowledge in microbiology (none -both questions are answered negatively-, basic -player knows what a bacterium is but not a bacterial capsule-, or advanced -both answers are positive-); questions D-I score knowledge acquisition; J and K are control questions; L-O evaluate the appreciation of the game; and P is an optional free text-entry answer for additional feedback. <br>2) y= yes, n=no, idk=I don’t know, VE=very easy, E=easy, A=adequate, D=difficult, VD=very difficult.</p>
Iceland as stepping stone for intercontinental spread of highly pathogenic avian influenza H5N1 virus between Europe and North America: data set on phylogeographic analysis
<p>Highly pathogenic avian influenza viruses (HPAIV) subtype H5 clade 2.3.4.4b have widely spread within the northern hemisphere since 2020 and threaten wild bird populations as well as poultry production. For the very first time, HPAIV were detected in wild birds and, subsequently, in poultry holdings in Iceland.</p> <p>Here, we present phylogeographic evidence that Iceland has been used as a stepping stone for HPAIV translocation from Northern Europe to North America in 2021 and describe two independent incursions of HPAI H5N1 clade 2.3.4.4b viruses of two different genotypes to Iceland in 2021 and 2022.</p>
Quantifying the Sensitivity of Sea Level Change in Coastal Localities to the Geometry of Polar Ice Mass Flux -- Supplemental Data Set: Sea Level Sensitivity Kernels
<p><strong>Quantifying the Sensitivity of Sea Level Change in Coastal Localities to the Geometry of Polar Ice Mass Flux<br> SUPPLEMENTAL DATA SET: SEA LEVEL SENSITIVITY KERNELS</strong></p> <p>To accompany</p> <p> Jerry X. Mitrovica, Carling C. Hay, Robert E. Kopp, Christopher Harig, and<br> Konstantin Laytchev (2018). Quantifying the Sensitivity of Sea Level Change<br> in Coastal Localities to the Geometry of Polar Ice Mass Flux. Journal of<br> Climate. doi: 10.1175/JCLI-D-17-0465.1.</p> <p>We provide sea level kernels for ~740 tide gauge sites in the Permanent Service for Mean Sea Level (PSMSL) database (Holgate et al., 2013). Kernels associated with sensitivities to Greenland and Alaskan glacier melt are given on a spatial grid covering the globe, with 512 latitude rows (i=1,512) and 1024 longitude (j=1,1024) columns.</p> <p>Longitude values are evenly spaced moving eastward from Greenwich (the jth grid point has an east longitude value of (j-1)×360°/1024). Latitude values are Gauss-Legendre points beginning close to the North Pole and ending near the South Pole. Kernels associated with sensitivities to Antarctic melt are given on a spatial grid covering the globe, with 256 (Gauss-Legendre) latitude rows (i=1,256) and 512 longitude (j=1,512) columns. Longitude values are evenly spaced moving eastward from Greenwich.</p> <p>The format of the files is: </p> <p> grid_sitenumber_region.txt</p> <p>where “region” is either “green” (Greenland), “ant” (Antarctic) or “Alaska” (Alaska). The list of sites (and site numbers) is provided in the sites.txt file. The first 8 sites in this list were test sites and can be ignored.</p>
Unsteady Aerodynamics Open Data Set
<p>A selection of four different unsteady aerodynamic experiments have been done to prepare a database which will serve for the analysis, investigation and tool validation of airfoil unsteady behavior of wind turbine blades.<br> The four experiments and selected data are:</p> <ul> <li>University of Glasgow dynamic stall experiments: NACA0015 and NACA0030 airfoils tested at sinusoidal type motion of the pitch.</li> <li>NREL OSU experiments: LS(1)0417MOD, NACA4415 and S809 airfoils tested at sinusoidal type motion of the pitch.</li> <li>CENER unsteady airfoil pitching and flapping tests at DTU: NACA643-418 airfoil tested at sinusoidal type motion of the pitch, the flap and combined pitch and flap.</li> <li>ForWind airfoil tests under tailored inflow turbulence: DU00W212 airfoil with laminar flow, open grid condition and one sinusoidal dynamic grid condition.</li> </ul>
Data set of the manuscript titled: Follicular Immune Landscaping Reveals a distinct profile of FOXP3hi CD4+ T cells in Treated compared to Untreated HIV
<p>Multiplex imaging data were collected using a scanning confocal system (STELARIS, Leica) and proccessed with the Imaris and Fiji imaging programs. csv files incuding the position identifiers and intensities for each fluorochrome used were generated and data were further analysed using the FlowJo10 program. Neighboring analysis was performed using the G function and mean of minimum distances of relevant cell type pairs. </p>
Exploring AdaBoost and Random Forests machine learning approaches for infrared pathology on unbalanced data sets
<p>The use of infrared spectroscopy to augment decision-making in histopathology is a promising direction for the diagnosis of many disease types. Hyperspectral images of healthy and diseased tissue, generated by infrared spectroscopy, are used to build chemometric models that can provide objective metrics of disease state. It is important to build robust and stable models to provide confidence to the end user. The data used to develop such models can have a variety of characteristics which can pose problems to many model-building approaches. Here we have compared the performance of two machine learning algorithms – AdaBoost and Random Forests – on a variety of non-uniform data sets. Using samples of breast cancer tissue, we devised a range of training data capable of describing the problem space. Models were constructed from these training sets and their characteristics compared. In terms of separating infrared spectra of cancerous epithelium tissue from normal-associated tissue on the tissue microarray, both AdaBoost and Random Forests algorithms were shown to give excellent classification performance (over 95% accuracy) in this study. AdaBoost models were more robust when datasets with large imbalance were provided. The outcomes of this work are a measure of classification accuracy as a function of training data available, and a clear recommendation for choice of machine learning approach.</p>
Surface water and flooding dynamics data set based on seasonally continuous Landsat data (1986-2011) in a dryland river basin
<p>Animations of the data are available here: <a href="https://doi.org/10.5281/zenodo.2438110">https://doi.org/10.5281/zenodo.2438110</a></p> <p>If you are using this data set, please cite the following publication:</p> <p>Tulbure, M.G. and M. Broich (2018). Spatiotemporal patterns and effects of climate and land use on surface water extent dynamics in a dryland region with three decades of Landsat satellite data. Science of the Total Environment. https://www.sciencedirect.com/science/article/pii/S0048969718347466 </p> <p>The data represent statistically validated surface water and flooding extent dynamics derived from seasonally continous Landsat TM/ETM+ data and random forest models, and summarised to the maximum extent of surface water per season between 1986-2011 over Australia's Murray-Darling Basin. The overall accuracy was over 99% and producer's accuracy for water 87% +/- 3%. </p> <p>The method is described in the following publication: <br> Tulbure, M.G., M. Broich, S.V. Stehman, A. Kommareddy. (2016). Surface water extent dynamics from three decades of seasonally continuous Landsat time series at subcontinental scale in a semi-arid region. Remote Sensing of Environment. 178: 142-157</p> <p>URL: https://www.sciencedirect.com/science/article/pii/S0034425716300621 </p> <p>Data are provided in GeoTIFF format per season per year. File naming convention is as follows:<br> yy_inund_freq_season_SamplingMethod. For example, "99_inund_freq_winter_max" will represent inundation frequency for winter 1999 resampled using a maximum resampling method. </p> <p>Inundation frequency represents the number of times a pixel has been flagged as flooded out of the times that pixel had valid observations * 100. Valid observation exclude no data values and clouds. The valid range of inundation frequency is 0-100 [%], with 255 indicating no data values. Data type is eight bit unsigned integer (uint8). </p> <p>The data were resampled to 120m resolution to reduce file size. The resampling methods used include max (e.g. selects the max value of all non-NODATA contributing 30m pixels) and mean (median and min can be provided upon request). If you are unsure which resampling to use, you may want to start with the mean. </p>
Selecting for infectivity across metapopulations can increase virulence in the social microbe Bacillus thuringiensis:data set.
<p>Passage experiments that sequentially infect hosts with parasites have long been used to manipulate virulence. However, for many invertebrate pathogens passage has been applied naively without a full theoretical understanding of how best to select for increased virulence and this has led to very mixed results. Understanding the evolution of virulence is complex because selection on parasites occurs across multiple spatial scales with potentially different conflicts operating on parasites with different life-histories. For example, in social microbes, strong selection on replication rate within hosts can lead to cheating and loss of virulence, because investment in public goods virulence reduces replication rate. </p> <p>In this study<em> </em>we tested how varying mutation supply and selection for infectivity or pathogen yield (population size in hosts) affected evolution of virulence against resistant hosts in the specialist insect pathogen <em>Bacillus thuringiensis</em>, aiming to optimize methods for strain improvement against a difficult to kill insect target. We show that selection for infectivity using competition between sub-populations in a metapopulation prevents social cheating, acts to retain key virulence plasmids and facilitates increased virulence. Increased virulence was associated with reduced efficiency of sporulation, and possible loss of function in putative regulatory genes but not with altered expression of the primary virulence factors. Selection in a metapopulation provides a broadly applicable tool for improving the efficacy of biocontrol agents. Moreover, a structured host population can facilitate artificial selection on infectivity, while selection on life history traits such as faster replication or larger population sizes can reduce virulence in social microbes.</p>
Data set discussed in "Beyond Fortune 500: Women in a Global Network of Directors"
<p>Bipartite graph of directors and companies. Generated from information on the Financial Times website (<a href="https://markets.ft.com/data/equities/results">https://markets.ft.com/data/equities/results</a>), retrieved on 17 September 2016.</p> <p>Blank fields are used for missing data.</p> <p><strong>comp_nodes.csv:</strong></p> <ul> <li>id: unique identifier</li> <li>ft_country: name of the country</li> <li>ft_sector: segment of the economy in which a company operates</li> <li>ft_industry: specific business (i.e., subset of sector) in which a company operates</li> <li>ft_employees_num: number of company's employees. "NA" if the vertex represents a person or if the company's number of employees is unknown.</li> </ul> <p><strong>comp_people_edges.csv:</strong></p> <ul> <li>person_id:</li> <li>comp_id: company identifier. It matches the identifier in comp_nodes.csv</li> </ul> <p><strong>people_one_mode_edges.csv:</strong></p> <p>Edges in the one-mode projection, in which two directors are connected if and only if they sit together on at least one board. Numbers correspond to the identifiers in unique_people_nodes.csv.</p> <p><strong>unique_people_nodes.csv:</strong></p> <ul> <li>ID: unique identifier</li> <li>age: years of age</li> <li>gender_base: "Male" or "Female"</li> </ul>
Water temperature in the hidden, subglacial lake at Uruguay Island, Antarctic Peninsula region, 2020-2021, and additional data sets.
The dataset contains temperature measurements in a small subglicer (hidden) lake of Antarctic Peninsula region at several levels of depth. The measurements cover almost a full year and provide an understanding of the temperature and hydrological regime of the water body. Weather measurement data and statistics is provided additionally.
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - BISC1 SET Surface Water level data from in Biscayne National Park, Florida, USA (2016-2025)
Surface water level data (m) was collected in Biscayne National Park (BISC) by the South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2016 to 2025 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This dataset belongs to Site 1, known as BISC-SET-1 or BISC1. This data-package is complete.
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - BISC2 SET Surface Water level data from in Biscayne National Park, Florida, USA (2017-2025)
Water level data (m) was collected in Biscayne National Park (BISC) by the National Park Service - South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2017-2025 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This dataset belongs to Site 2, known as BISC-SET-2 or BISC2. This data-package is complete.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.