Published February 1, 2019 | Version 1.0

Spark application traces of the run of 9 differents instance of BigDataBench applications

Authors/Creators

  • 1. Inria

Description

Spark application logs of the run of 9 different instances of BigDataBench applications. Each run was done with an input size in GB of 32, 64, 128 and with 4, 8, 16 executor respectively. There were 2 executors per node; so, it runs on 2, 4, and 8 nodes + a master node for  Hadoop services.

The input data were generated with the Data generator provided by BigDataBench using this procedure:

https://gitlab.inria.fr/mmercier/bebida/blob/master/experiments/generate_dataset/journal.md

List of the applications and their parameters:

  • Grep:
    •  parameter: "word"
  • WordCount
    •   no parameters
  • Kmean:
    •   parameter: "4 3" input size in GB: "32, 64, 128"

BigDataBench implementation can be downloaded here: http://prof.ict.ac.cn/download.html

It was run on Debian 8, with Spark 2.1.0, on top of Hadoop 2.7.1 with Yarn and HDFS, using openjdk-7-jre-headless.

All the details of the environment can be found here:

https://gitlab.inria.fr/mmercier/bebida/blob/master/environments/bebida-slave.yaml

The experiment in itself is described here:

https://gitlab.inria.fr/mmercier/bebida/tree/master/experiments/run_big_data_workload

All nodes Hardware description of the nodes:

https://public-api.grid5000.fr/stable/sites/nancy/clusters/graphene/nodes.json?pretty=1

They can be visualized using the Spark History server.

Files

Files (24.0 MB)

Name Size Download all
md5:f0039ec36bab56b856491d859d730708
24.0 MB Download