From Bronze to Broken: A Grounded Theory Study of Anti-Patterns and Accruing Data Debt in Medallion Lakehouse Deployments
Authors/Creators
Description
The medallion lakehouse architecture has rapidly gained prominence as a hybrid data management paradigm, promising to unify the flexibility of data lakes with the governance and structure of data warehouses. Its prescribed, multi-layered flow from raw Bronze to cleansed Silver to refined Gold tables is championed for enabling scalable data refinement and self-service analytics. However, this prescriptive design, when enacted in complex organizational environments, may inadvertently foster systematic deviations that compromise its foundational value proposition. While foundational literature, such as Armbrust et al.'s "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics" (2021), establishes the architectural blueprint and its theoretical benefits, and Heller et al. in "Data Engineering Zoomcamp: An Open-Source Curriculum for Modern Data Platforms" (2022) detail practical implementation patterns, a significant gap remains in the critical study of its failure modes. There is a paucity of empirical research investigating the latent architectural anti-patterns that emerge in practice and their consequential role in accruing crippling data debt the long-term costs of suboptimal data management decisions. This study addresses this gap by employing a constructivist grounded theory methodology to develop a substantive theory of operational decay within medallion implementations. We conducted in-depth, semi-structured interviews with thirty-two data architects, engineers, and platform leads across twelve organizations in the finance, retail, and technology sectors, all of whom had at least eighteen months of hands-on experience with active medallion lakehouse deployments on platforms like Databricks, Snowflake, or Apache Spark.
Our analysis, iterating between data collection and constant comparative analysis, led to the identification of four core categories of architectural anti-patterns that systematically deviate from the idealized medallion model. First, the Bypass and Contamination pattern, where urgent analytic demands lead producers or consumers to either write directly to the silver or gold layers, bypassing the cleansing tier, or to read directly from Bronze, thereby embedding raw data logic into downstream applications. This erodes the single source of truth and replicates data quality issues. Second, the Schema-on-Write Procrastination pattern, observed in teams misapplying schema-on-read flexibility in the bronze layer as a permanent fixture, leading to a proliferation of unstructured or semi-
structured dumps. This defers critical data modeling and validation, creating a "modeling cliff" at the Silver layer that
becomes a persistent bottleneck. Third, the Tiered Data Hoarding pattern, where a misinterpretation of data retention policies or fear of data loss results in the preservation of massive, incremental snapshots across all three layers without intelligent compaction or lifecycle management. This exponentially inflates storage costs and cripples query performance, directly countering the efficiency promises highlighted in early lakehouse discourse. Fourth, the Governance Decoupling pattern, where metadata management, lineage tracking, and access controls are implemented as a separate, batch-updated system rather than being intrinsically woven into the data transformation pipelines, leading to a loss of synchronization and trust.
The study further establishes a direct correlation between the persistence of these anti-patterns and the rapid accrual of three forms of data debt: Process Debt, manifesting as spiraling pipeline maintenance costs and declining engineer velocity; Quality Debt, seen in eroding trust in Gold-layer datasets and conflicting business metrics; and Platform Debt, evidenced by runaway cloud infrastructure costs and inability to upgrade underlying platform components. The emergent theory posits that the very prescriptive nature of the medallion layers, without corresponding adaptive governance and architectural guardrails, creates a false sense of completeness. This leads teams to focus on layer completion as a success metric rather than on the continuous fitness-for-purpose of the data ecosystem. The research concludes that mitigating this decay requires a shift from a purely layer-centric view to a contract-first, product-oriented approach. Recommendations include the formalization of inter-layer data contracts, the implementation of automated architectural fitness functions in CI/CD pipelines to detect anti-patterns, and the treatment of each data product in the Gold layer as a managed service with explicit SLAs. This study contributes a critical, reality-grounded counterpoint to the proliferating promotional literature on lakehouses. It provides practitioners with a diagnostic framework to assess their own deployments and offers academia a rich, empirically derived taxonomy of systemic failure modes in modern, layered data architectures, challenging purely schematic evaluations of their efficacy.
Files
EJAET-11-1-90-100.pdf
Files
(689.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:b69461b28a3fa48fc90387077bdea414
|
689.7 kB | Preview Download |
Additional details
References
- [1]. M. Armbrust et al., "Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics," in Proc. CIDR, 2021. [Online]. Available: https://www.cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf
- [2]. V. Heller, A. Kravchenko, and A. G. G. C., "Data Engineering Zoomcamp: An Open-Source Curriculum for Modern Data Platforms," in Proc. 2022 IEEE/ACM 44th Int. Conf. Softw. Eng., Softw. Eng. Educ. Training (ICSE-SEET), 2022, pp. 273–274, doi: 10.1145/3510456.3513652.
- [3]. Z. Dehghani, Data Mesh: Delivering Data-Driven Value at Scale. Sebastopol, CA, USA: O'Reilly Media, 2022.
- [4]. D. Sculley et al., "Hidden Technical Debt in Machine Learning Systems," in Proc. Adv. Neural Inf. Process. Syst. (NIPS), vol. 28, 2015, pp. 2503–2511.
- [5]. T. G. J. de Oliveira, M. A. Gerosa, and F. Kon, "Data Engineering: Who Does What? Findings from a Multi-Method Study," in Proc. 2023 IEEE/ACM 45th Int. Conf. Softw. Eng., Softw. Eng. Pract. (ICSE-SEIP), 2023, pp. 312–323, doi: 10.1109/ICSE-SEIP58684.2023.00032.