Skip to main content
zenodoopen

Public BI benchmark - part 1

<p>Originally published on:&nbsp;https://github.com/cwida/public_bi_benchmark&nbsp;</p> <p>Originally compressed with bzip2. Compressed here with gzip.</p> <p>&quot;User generated benchmark derived from the DBTest&#39;18 paper [1] by Tableau. It contains real data and queries from 46 public workbooks in Tableau Public [2].</p> <p>We downloaded 46 of the biggest workbooks and converted the data to&nbsp;<em>.csv</em>&nbsp;files and collected the SQL queries that appear in the Tableau log when the workbooks are visualized. We processed the&nbsp;<em>.csv</em>&nbsp;files and queries with the purpose of making them load and run on different database systems.</p> <p>Each directory is associated with a workbook and contains:</p> <pre><code>samples: a sample of each .csv file (first 20 rows) tables: .sql files containing the schema of each .csv file queries: .sql files containing the queries data-urls.txt: links for downloading the full .csv.bz2 compressed files </code></pre> <p>There are 46 workbooks containing 206 tables (.csv files) with the total size of 41 GB compressed and 386 GB uncompressed.</p> <p>Multiple .csv files may overlap but are not identical. This is because Tableau extracts the same workbook in multiple different ways for different queries.&quot;</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0

Topics