Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “SCADA data”
Aventa AV-7 ETH Zurich Research Wind Turbine SCADA and high frequency Structural Health Monitoring (SHM) data
<p><strong>General description of wind turbine: </strong>The ETH owned wind turbine is Aventa AV-7, manufactured by Aventa AG in Switzerland and was commissioned in December 2002. The turbine is operated via a belt-driven generator and a frequency converter with a variable speed drive. The rated power of the Aventa AV-7 is 7 kW, beginning production at a wind speed of 2 m/s and having a cut-off speed of 14 m/s. The rotor diameter is 12.8 m with 3 rotor blades, and a hub height is 18m. The maximum rotational speed of the turbine is 63 rpm. The tower is a tubular steel-reinforced concrete structure, supported on concrete foundation, while the blades are made of glassfiber with a tubular steel main-spar. The turbine is regulated via a variable-speed and variable pitch control system.</p> <p><strong>Location of site: </strong>The wind turbine is located in Taggenberg, about 5 km from the city centre of Winterthur, Switzerland. This site is easily accessible by public transport and on foot with direct road access right next to the turbine. This prime location reduces the cost of site visits and allows for frequent personal monitoring of the site when test equipment is installed. The coordinates of the site are: 47°31'12.2"N 8°40'55.7"E.</p> <p><strong>Control and measurement systems and signals: </strong>The turbine is regulated via a variable-speed and collective variable pitch control system.</p> <p><strong>SHM Motivation: </strong>Designed and commissioned in 2002, the Aventa wind turbine in Winterthur is soon reaching its end of design lifetime. In order to assess the various techniques of predicting the remaining useful lifetime, a Structural Health Monitoring (SHM) campaign was implemented by ETH Zurich. The monitoring campaign started in 2020, and is still ongoing. In addition, the setup is used as a research platform on topics such as system identification, operational modal analysis, faults/damage detection and classification. We analyze the influence of operational and environmental conditions on the modal parameters and to further infer Performance Indicators (PIs) for assessing structural behavior in terms of deterioration processes.</p> <p><strong>Data Description: </strong>The tower and nacelle have been instrumented with 11 accelerometers distributed along the length of the tower, nacelle main frame, main bearing and generator. Two full bridge strain gauges are installed on the concrete tower based measuring fore-aft and side-side strain (and can be converted to bending moments) – all acceleration and strain signals sampled at 200Hz. Temperature and humidity are measured at the tower base – 1Hz data. In additional we are collecting operational performance data (SCADA), namely: wind speed, nacelle yaw orientation, rotor RPM, power output and turbine status – SCADA signals are sampled at 10Hz. See appendix for further details of the sensors layout.</p> <p>The measurements/instrumentation setup, type and layout is provided in the pdf files.</p> <p><strong>The data:</strong> the data is provided in zip files corresponding to four use-cases as follows:</p> <ul> <li>Normal operation data for system identification</li> <li>Aerodynamic imbalance on one blade</li> <li>Rotor icing event</li> <li>Failure of the flexible coupling of the linear drive of the collective pitch system</li> </ul> <p>The data for each of the four uses-cases is organized in zip files. The content of each zip file is as follows:</p> <ul> <li>Time-series data in HDF5 format</li> <li>Metadata: <ul> <li>Turbine specification (Aventa-AV-7.json and Aventa-AV-7.yaml)</li> <li>Sensor specification (Aventa_sensors.json )</li> <li>Unstructured description of the Aventa Turbine and the installed sensors (Aventa_Sensors_Specs.xlsx)</li> </ul> </li> <li>Semantic artifacts: <ul> <li>WindIO Wind Turbine YAML schema describing turbine specifications (IEAontology_schema.yaml)</li> <li>Sensor specification JSON schema (sensors_schema.json)</li> </ul> </li> <li>Media: Pictures of leading edge roughness and a clip of wind turbine operation</li> <li>Code: Jupyter notebook containing example code to load metadata from JSON and data from HDF5 files (example.ipynb)</li> </ul> <p>Additional data is available upon request, please contact:</p> <ul> <li>Prof. Dr. Eleni Chatzi (chatzi@ibk.baug.ethz.ch)</li> <li>Dr. Imad Abdallah (ai@rtdt.ai , abdallah@ibk.baug.ethz.ch)</li> </ul> <p>For further details or questions, please contact:</p> <p>Prof. Dr. Eleni Chatzi<br> Chair of Structural Mechanics & Monitoring</p> <p>ETH Zürich<br> <a href="http://www.chatzi.ibk.ethz.ch/">http://www.chatzi.ibk.ethz.ch/</a></p>
Data format figures-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY
<p>The original data set included noisy, missing and inconsistent data. Data<br> preprocessing improved the quality of the data and facilitated e±cient data<br> mining tasks.<br> Before the experiment, we prepared data suitable to next operation as<br> following steps:<br> ² Delete or replace missing values;<br> ² Delete redundant properties (columns);<br> ² Data Transformation;<br> ² Data Discretization;<br> ² Export data to a required .ar® or .csv format ¯le [11].<br> The original and modi¯ed formats of data set are shown in Figure 1 and<br> Figure 2.<br> Data visualization is also a very useful technique because it helps to deter-<br> mine the di±culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The ¯gure 3 shows the variation<br> of the temperature in time.</p>
Figure 3. Data visualization-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY
<p>Data visualization is also a very useful technique because it helps to deter-<br> mine the di±culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The ¯gure 3 shows the variation<br> of the temperature in time.</p>
Industrial Big Data Innovation Platform SCADA Dataset
<p>This is the SCADA dataset from the first Industrial Big Data Innovation Competition, for more information, visit: <a href="https://www.industrial-bigdata.com/Challenge/title?competitionId=LEIREZMM8TT5VBU0TLJ61FPAI6WWJOJY&type=">数境创新大赛平台(industrial-bigdata.com)。</a></p>
Wind Turbine SCADA Data For Early Fault Detection
<p>This dataset is published together with the <a href="https://doi.org/10.3390/data9120138">paper</a> "CARE to Compare: A real-world dataset for anomaly detection in wind turbine data" which explains the dataset in detail and defines the CARE score that can be used to evaluate anomaly detection algorithms on this dataset. When referring to this dataset, please cite the paper mentioned in the related work section. </p> <p>The data consists of 95 datasets, containing 89 years of SCADA time series distributed across 36 different wind turbines<br>from the three wind farms A, B and C. The number of features depends on the wind farm; Wind farm A has 86 features, wind farm B has 257 features and wind farm C has 957 features. </p> <p>The overall dataset is balanced, as 45 out the 95 datasets contain a labeled anomaly event that leads up to a turbine fault and the other 50 datasets represent normal behavior. Additionally, the quality of training data is ensured by turbine-status-based labels for each data point and further information about some of the given turbine faults are included.</p> <p>The data for Wind farm A is based on data from the EDP open data platform (https://www.edp.com/en/innovation/open-data/data), <br>and consists of 5 wind turbines of an onshore wind farm in Portugal. <br>It contains SCADA data and information derived by a given fault logbook which defines start timestamps for specified faults. <br>From this data 22 datasets were selected to be included in this data collection. <br>The other two wind farms are offshore wind farms located in Germany. All three datasets were anonymized due to confidentiality reasons for the wind farms B and C.<br>Each dataset is provided in form of a csv-file with columns defining the features and rows representing the data points of the time series. Files</p> <p>More detailed information can be found in the included README-file.</p> <p><strong>Notes</strong></p> <p>In wind farm A status_type_id labels can be ignored while evaluating prediction time frames of error events with metrics like the CARE-score since the status_type_id is of wind farm A is based on the EDP failure logbook and it is intended to be used for filtering of the training data.</p> <p><strong>Version Changes:</strong></p> <p><em>Version 5 -> 6:</em></p> <ul> <li>Changed unit of sensor_40 and sensor_61 for wind farm C to hPa instead of bar. This unit error became obvious when looking at the data and comparing it to the standard air pressure.</li> <li>Edited event_description of events 34, 7 and 19 to high temperature in transformer cell.</li> <li>Changed date in event description of event 44 since it was not affected by the change in the date anonymization procedure from version 2.</li> <li>Changed date in event description of event 47 since it was not affected by the change in the date anonymization procedure from version 2 and edited the description text</li> <li> Changed date format in event_info files to match the date format in the dataset files.</li> <li>Fixed typo in Readme</li> <li>Re-added Readme files</li> </ul> <p><em>Version</em> 4->5:</p> <p>Corrections to labels were made:</p> <ul> <li>Previously missing status_type_id 4 labels were added to datasets in Wind Farm A. </li> <li>Event 51 from Wind Farm A was wrongly labeled as a normal event. With the newly added status_type_id 4 occurences, it is to be considered an anomaly event due to a gearbox bearing damage within the prediction data.</li> <li>Wind Farm A no longer contains status_type_id 5. All occurences of status_type_id 5 have been changed to 0 and are considered normal time stamps. This change is done, because status_type_id 5 was set as a result of a wind speed and power analysis, flagging potential anomalous data. This is not based on a fixed ground truth, so status_type_id 5 was removed. For Wind Farms B and C status_type_id 5 is still valid since it is based on real SCADA-status codes.</li> <li>The event_info.csv files now contain an additional column 'asset_id'.</li> </ul> <p><em>Version 3->4:<br></em></p> <ul> <li>The change of the timestamp anonymization lead to duplicate timestamps when transitioning from a leap year to 2022. This is now fixed in Version 4.</li> </ul> <p><em>Version 2->3:</em></p> <ul> <li>In version 2 timestamp changes were not consistent with the timestamps in the event-info-files. Version 3 fixes this.</li> </ul> <p><em>Version 1->2:<br></em></p> <ul> <li>Version 2 contains one deviation from version 1 regarding the anonymization procedure. Instead of shifting the timestamps of each sub-dataset by a random number of years, the size of the time shift is now determined to be the number of years so that each sub-dataset starts in 2022. This change is made to make the timestamp anonymization more consistent and to avoid future timestamps being present within the data.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.