Published June 7, 2023 | Version 0.0.1

YADL Data Lake files and artifacts

  • 1. INRIA
  • 2. Eurecom
  • 3. Dataiku

Description

Files composing the YADL data lake, for the paper "Benchmarking Data Lake for Join Discovery and
Learning with Relational Data"

Tabular representation learning is gaining traction as machine learning techniques are increasingly applied to database problems, including data integration across multiple tables. A challenge lies in the split focus among research questions: merging a set of different tables to assemble a larger feature matrix without considering the downstream task (common in database research), or utilizing the assembled feature matrix for model training (prevalent in machine learning). Despite significant work conducted separately, there has been limited emphasis on the end-to-end process – as evidenced by the lack of suitable benchmarks for this issue. This  paper introduces the first benchmark data lake on the subject, aiming to encourage reproducible research on learning from data lakes. Using a proof-of-principle complete analytic pipeline, it demonstrates the benefit of studying assembling tables for a supervised-learning goal.
 

Files

Files (781.0 MB)

Name Size
md5:ef3ea674a79f17f247523f9bd1a684d5
288.7 MB Download
md5:931110f4c4083467f2f3266bc0adf3fd
492.3 MB Download