Skip to main content
zenodoopen

Spark application traces of the run of 9 differents instance of BigDataBench applications

<p>Spark application logs of the run of 9 different instances of BigDataBench applications. Each run was done with an input size in GB of 32, 64, 128 and with 4, 8, 16 executor respectively. There were 2 executors per node; so, it runs on 2, 4, and 8 nodes + a master node for&nbsp; Hadoop services.</p> <p>The input data were generated with the Data generator provided by BigDataBench using this procedure:</p> <p>https://gitlab.inria.fr/mmercier/bebida/blob/master/experiments/generate_dataset/journal.md</p> <p>List of the applications and their parameters:</p> <ul> <li>Grep: <ul> <li>&nbsp;parameter: &quot;word&quot;</li> </ul> </li> <li>WordCount <ul> <li>&nbsp; no parameters</li> </ul> </li> <li>Kmean: <ul> <li>&nbsp; parameter: &quot;4 3&quot; input size in GB: &quot;32, 64, 128&quot;</li> </ul> </li> </ul> <p>BigDataBench implementation can be downloaded here: http://prof.ict.ac.cn/download.html</p> <p>It was run on Debian 8, with Spark 2.1.0, on top of Hadoop 2.7.1 with Yarn and HDFS, using openjdk-7-jre-headless.</p> <p>All the details of the environment can be found here:</p> <p>https://gitlab.inria.fr/mmercier/bebida/blob/master/environments/bebida-slave.yaml</p> <p>The experiment in itself is described here:</p> <p>https://gitlab.inria.fr/mmercier/bebida/tree/master/experiments/run_big_data_workload</p> <p>All nodes Hardware description of the nodes:</p> <p>https://public-api.grid5000.fr/stable/sites/nancy/clusters/graphene/nodes.json?pretty=1</p> <p>They can be visualized using the Spark History server.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics