Public BI benchmark - part 2 & results
<p>Originally published on: https://github.com/cwida/public_bi_benchmark </p> <p>Originally compressed with bzip2. Compressed here with gzip.</p> <p>"User generated benchmark derived from the DBTest'18 paper [1] by Tableau. It contains real data and queries from 46 public workbooks in Tableau Public [2].</p> <p>We downloaded 46 of the biggest workbooks and converted the data to <em>.csv</em> files and collected the SQL queries that appear in the Tableau log when the workbooks are visualized. We processed the <em>.csv</em> files and queries with the purpose of making them load and run on different database systems.</p> <p>Each directory is associated with a workbook and contains:</p> <pre><code>samples: a sample of each .csv file (first 20 rows) tables: .sql files containing the schema of each .csv file queries: .sql files containing the queries data-urls.txt: links for downloading the full .csv.bz2 compressed files </code></pre> <p>There are 46 workbooks containing 206 tables (.csv files) with the total size of 41 GB compressed and 386 GB uncompressed.</p> <p>Multiple .csv files may overlap but are not identical. This is because Tableau extracts the same workbook in multiple different ways for different queries."</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0