YADL Data Lake files and artifacts
Authors/Creators
- 1. INRIA
- 2. Eurecom
- 3. Dataiku
Description
Files composing the YADL data lake, for the paper "Benchmarking Data Lake for Join Discovery and
Learning with Relational Data"
Tabular representation learning is gaining traction as machine learning techniques are increasingly applied to database problems, including data integration across multiple tables. A challenge lies in the split focus among research questions: merging a set of different tables to assemble a larger feature matrix without considering the downstream task (common in database research), or utilizing the assembled feature matrix for model training (prevalent in machine learning). Despite significant work conducted separately, there has been limited emphasis on the end-to-end process – as evidenced by the lack of suitable benchmarks for this issue. This paper introduces the first benchmark data lake on the subject, aiming to encourage reproducible research on learning from data lakes. Using a proof-of-principle complete analytic pipeline, it demonstrates the benefit of studying assembling tables for a supervised-learning goal.