Skip to main content
zenodoopen

GoiEner smart meters data

<ul> <li><strong>Name</strong>: GoiEner smart meters data</li> <li><strong>Summary</strong>: The dataset contains hourly time series of electricity consumption (kWh) provided by the Spanish electricity retailer GoiEner. The time series are arranged in four compressed files: <ul> <li><strong>raw.tzst</strong>, contains raw time series of all GoiEner clients (any date, any length, may have missing samples).</li> <li><strong>imp-pre.tzst</strong>, contains processed time series (imputation of missing samples), longer than one year, collected before March 1, 2020.</li> <li><strong>imp-in.tzst</strong>, contains processed time series (imputation of missing samples), longer than one year, collected between March 1, 2020 and May 30, 2021.</li> <li><strong>imp-post.tzst</strong>, contains processed time series (imputation of missing samples), longer than one year, collected after May 30, 2020.</li> <li><strong>metadata.csv</strong>, contains relevant information for each time series.</li> </ul> </li> <li><strong>License</strong>: CC-BY-SA</li> <li><strong>Acknowledge</strong>: These data have been collected in the framework of the WHY project. This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 891943.</li> <li><strong>Disclaimer</strong>: The sole responsibility for the content of this publication lies with the authors. It does not necessarily reflect the opinion of the Executive Agency for Small and Medium-sized Enterprises (EASME) or the European Commission (EC). EASME or the EC are not responsible for any use that may be made of the information contained therein.</li> <li><strong>Collection Date</strong>: From November 2, 2014 to June 8, 2022.</li> <li><strong>Publication Date</strong>: December 1, 2022.</li> <li><strong>DOI</strong>: 10.5281/zenodo.7362094</li> <li><strong>Other repositories</strong>: None.</li> <li><strong>Author</strong>: GoiEner, University of Deusto.</li> <li><strong>Objective of collection</strong>: This dataset was originally used to establish a methodology for clustering households according to their electricity consumption.</li> <li><strong>Description</strong>: The meaning of each column is described next for each file. <ul> <li><strong>raw.tzst</strong>: (no column names provided) <ul> <li>timestamp;</li> <li>electricity consumption in kWh.</li> </ul> </li> <li><strong>imp-pre.tzst</strong>, <strong>imp-in.tzst</strong>, <strong>imp-post.tzst</strong>: <ul> <li>&ldquo;<em>timestamp</em>&rdquo;: timestamp;</li> <li>&ldquo;<em>kWh</em>&rdquo;: electricity consumption in kWh;</li> <li>&ldquo;<em>imputed</em>&rdquo;: binary value indicating whether the row has been obtained by imputation.</li> </ul> </li> <li><strong>metadata.csv</strong>: <ul> <li>&ldquo;<em>user</em>&rdquo;: 64-character identifying a user;</li> <li>&ldquo;<em>start_date</em>&rdquo;: initial timestamp of the time series;</li> <li>&ldquo;<em>end_date</em>&rdquo;: final timestamp of the time series;</li> <li>&ldquo;<em>length_days</em>&rdquo;: number of days elapsed between the initial and the final timestamps;</li> <li>&ldquo;<em>length_years</em>&rdquo;: number of years elapsed between the initial and the final timestamps;</li> <li>&ldquo;<em>potential_samples</em>&rdquo;: number of samples that should be between the initial and the final timestamps of the time series if there were no missing values;</li> <li>&ldquo;<em>actual_samples</em>&rdquo;: number of actual samples of the time series;</li> <li>&ldquo;<em>missing_samples_abs</em>&rdquo;: number of potential samples minus actual samples;</li> <li>&ldquo;<em>missing_samples_pct</em>&rdquo;: potential samples minus actual samples as a percentage;</li> <li>&ldquo;<em>contract_start_date</em>&rdquo;: contract start date; &ldquo;<em>contract_end_date</em>&rdquo;: contract end date;</li> <li>&ldquo;<em>contracted_tariff</em>&rdquo;: type of tariff contracted (2.X: households and SMEs, 3.X: SMEs with high consumption, 6.X: industries, large commercial areas, and farms);</li> <li>&ldquo;<em>self_consumption_type</em>&rdquo;: the type of self-consumption to which the users are subscribed;</li> <li>&ldquo;<em>p1</em>&rdquo;, &ldquo;<em>p2</em>&rdquo;, &ldquo;<em>p3</em>&rdquo;, &ldquo;<em>p4</em>&rdquo;, &ldquo;<em>p5</em>&rdquo;, &ldquo;<em>p6</em>&rdquo;: contracted power (in kW) for each of the six time slots;</li> <li>&ldquo;<em>province</em>&rdquo;: province where the user is located;</li> <li>&ldquo;<em>municipality</em>&rdquo;: municipality where the user is located (municipalities below 50.000 inhabitants have been removed);</li> <li>&ldquo;<em>zip_code</em>&rdquo;: post code (post codes of municipalities below 50.000 inhabitants have been removed);</li> <li>&ldquo;<em>cnae</em>&rdquo;: CNAE (<em>Clasificaci&oacute;n Nacional de Actividades Econ&oacute;micas</em>) code for economic activity classification.</li> </ul> </li> </ul> </li> <li><strong>5 star</strong>: ⭐⭐⭐</li> <li><strong>Preprocessing steps</strong>: Data cleaning (imputation of missing values using the Last Observation Carried Forward algorithm using weekly seasons); data integration (combination of multiple SIMEL files, i.e. the data sources); data transformation (anonymization, unit conversion, metadata generation).</li> <li><strong>Reuse:</strong> This dataset is related to datasets: <ul> <li>&quot;A database of features extracted from different electricity load profiles datasets&quot; (DOI 10.5281/zenodo.7382818), where time series feature extraction has been performed.&nbsp;</li> <li>&quot;Measuring the flexibility achieved by a change of tariff&quot; (DOI 10.5281/zenodo.7382924), where the metadata has been extended to include the results of a socio-economic characterization and the answers to a survey about barriers to adapt to a change of tariff.</li> </ul> </li> <li><strong>Update policy:</strong> There might be a single update in mid-2023.</li> <li><strong>Ethics and legal aspects:</strong> The data provided by GoiEner contained values of the CUPS (Meter Point Administration Number), which are personal data. A pre-processing step has been carried out to replace the CUPS by random 64-character hashes.</li> <li><strong>Technical aspects:</strong> <ul> <li><strong>raw.tzst</strong> contains a 15.1 GB folder with 25,559 CSV files;</li> <li><strong>imp-pre.tzst </strong>contains a 6.28 GB folder with 12,149 CSV files;</li> <li><strong>imp-in.tzst</strong> contains a 4.36 GB folder with 15.562 CSV files; and</li> <li><strong>imp-post.tzst</strong> contains a 4.01 GB folder with 17.519 CSV files.</li> </ul> </li> <li><strong>Other:</strong> None.</li> </ul>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
8

Topics