Published August 10, 2022 | Version v1
Conference paper Open

SIM-PIPE DryRunner: An approach for testing container-based big data pipelines and generating simulation data

  • 1. SINTEF AS
  • 2. Oslo Metropolitan University

Description

Big data pipelines are becoming increasingly vital in a wide range of data intensive application domains such as digital healthcare, telecommunication, and manufacturing for efficiently processing data. Data pipelines in such domains are complex and dynamic and involve a number of data processing steps that are deployed on heterogeneous computing resources under the realm of the Edge-Cloud paradigm. The processes of testing and simulating big data pipelines on heterogeneous resources need to be able to accurately represent this complexity. However, since big data processing is heavily resource-intensive, it makes testing and simulation based on historical execution data impractical. In this paper, we introduce the SIM - PIPE Dry Runner approach - a dry run approach that deploys a big data pipeline step by step in an isolated environment and executes it with sample data; this approach could be used for testing big data pipelines and realising practical simulations using existing simulators.

Files

SIM-PIPE_DryRunner_An_approach_for_testing_container-based_big_data_pipelines_and_generating_simulation_data.pdf

Additional details

Funding

European Commission
DataCloud - ENABLING THE BIG DATA PIPELINE LIFECYCLE ON THE COMPUTING CONTINUUM 101016835