Published May 29, 2020 | Version v1
Thesis Open

Design and implementation of a distributed synopsis data engine on Apache Flink

  • 1. Athena Research Center, Technical University of Crete

Description

This work, it details the design and structure of a Synopses Data Engine (SDE) which combines the virtues of parallel processing and stream summarization towards delivering interactive analytics at extreme scale. The SDE is built on top of Apache Flink and implements a synopsis-as-a-service paradigm. In that it achieves (a) concurrently maintaining thousands of synopses of various types for thousands of streams on demand, (b) reusing maintained synopses among various concurrent workflows, (c) providing data summarization facilities even for cross-(Big Data) platform workflows, (d) pluggability of new synopses on-the-fly, (e) increased potential for workflow execution optimization. The proposed SDE is useful for interactive analytics at extreme scales because it enables (i) enhanced horizontal scalability, i.e., not only scaling out the computation to a number of processing units available in a computer cluster, but also harnessing the processing load assigned to each by operating on carefully-crafted data summaries, (ii) vertical scalability, i.e., scaling the computation to very high numbers of processed streams and (iii) federated scalability i.e., scaling the computation beyond single clusters and clouds by controlling the communication required to answer global queries posed over a number of potentially geo-dispersed clusters.

Files

Kontaxakis_Antonios_MSc_2020.pdf

Files (1.8 MB)

Name Size Download all
md5:0db7a3a1f87749458f04bb7deedbf1a2
1.8 MB Preview Download

Additional details

Funding

European Commission
INFORE - Interactive Extreme-Scale Analytics and Forecasting 825070