Published August 31, 2020 | Version v1

Enterprise-Scale Data Quality Improvement Using Machine Learning: Frameworks, Validation Strategies, and Operational Insights

Authors/Creators

Description

Enterprises operating at scale increasingly depend on accurate, consistent, and trustworthy data to support analytics, regulatory reporting, operational execution, and strategic decision making, yet traditional rule-based data quality programs have struggled to keep pace with rising data volumes, heterogeneous data sources, and rapidly changing business conditions. This study addresses the persistent gap between static data validation techniques and the dynamic error patterns that emerge within modern enterprise platforms by developing and evaluating machine learning driven frameworks designed to identify, classify, and prioritize data quality issues with higher precision and adaptability. Using a mixed methodological approach that integrates architectural analysis, quantitative experimentation, and scenario-oriented evaluation, the research examines how supervised learning models, probabilistic classifiers, and feature engineered validation pipelines can outperform conventional rule sets in detecting anomalies, missing values, inconsistent records, and entity level conflicts across complex datasets. Findings demonstrate that machine learning models achieve substantial gains in accuracy, recall, and operational robustness, while also improving the interpretability of validation decisions when embedded within structured enterprise workflows. The study presents innovative validation strategies that combine predictive modeling with governance-oriented feedback loops, enabling organizations to respond to evolving data behaviors and reduce downstream operational risks. By contributing reference architecture, empirical benchmarks, and applied insights, the research advances both academic understanding and industry practice in enterprise data quality management. The results indicate that machine learning enabled validation provides a scalable and future ready foundation for enterprises seeking to strengthen trust in data assets and enhance the resilience of data dependent operations.

Files

EJAET-7-8-138-149.pdf

Files (763.0 kB)

Name Size Download all
md5:1e3c70a5294a91b8c66a6f1aa9e7857b
763.0 kB Preview Download

Additional details

References

  • [1]. Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–34. 10.1080/07421222.1996.11518099
  • [2]. Shankaranarayanan, G., & Cai, Y. (2006). Supporting data quality management in decision-making. Decision Support Systems, 42(1), 302–317. 10.1016/j.dss.2004.12.006
  • [3]. Even, A., & Shankaranarayanan, G. (2009). Dual assessment of data quality in customer databases. Journal of Data and Information Quality, 1(3), 15. 10.1145/1659225.1659228
  • [4]. Chen, H., Hailey, D., Wang, N., & Yu, P. (2014). A review of data quality assessment methods for public health information systems. International Journal of Environmental Research and Public Health, 11(5), 5170–5207. 10.3390/ijerph110505170
  • [5]. Cai, L., & Zhu, Y. (2015). The challenges of data quality and data quality assessment in the big data era. Data Science Journal, 14(2), 1–10. 10.5334/dsj-2015-002
  • [6]. Ranshous, S., Shen, S., Koutra, D., Harenberg, S., Faloutsos, C., & Samatova, N. F. (2015). Anomaly detection in dynamic networks: A survey. WIREs Data Mining and Knowledge Discovery, 5(6), 353–375. 10.1002/wics.1347
  • [7]. Akoglu, L., Tong, H., & Koutra, D. (2015). Graph based anomaly detection and description: A survey. Data Mining and Knowledge Discovery, 29(3), 626–688. 10.1007/s10618-014-0365-y
  • [8]. Emani, C. K., Cullot, N., & Nicolle, C. (2015). Understandable big data: A survey. Computer Science Review, 17, 70–81. 10.1016/j.cosrev.2015.05.002
  • [9]. Gao, J., Xie, S., Tao, X., & Gao, Y. (2016). Big data validation and quality assurance: Issues, challenges and needs. In 2016 IEEE 10th International Symposium on Service-Oriented System Engineering (SOSE) (pp. 433–441). 10.1109/SOSE.2016.63
  • [10]. Sadiq, S., & Papotti, P. (2016). Big data quality: Whose problem is it? In 2016 IEEE 32nd International Conference on Data Engineering (ICDE) (pp. 1446–1447). 10.1109/ICDE.2016.7498367
  • [11]. Zhang, P., Xiong, F., Gao, J., & Wang, J. (2017). Data quality in big data processing: Issues, solutions and open problems. In 2017 IEEE SmartWorld, Ubiquitous Intelligence & Computing, Advanced & Trusted Computing, Scalable Computing & Communications, Cloud & Big Data Computing, Internet of People and Smart City Innovation (pp. 1–8). 10.1109/UIC-ATC.2017.8397554
  • [12]. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (Big Data) (pp. 1123–1132). 10.1109/BigData.2017.8258038
  • [13]. Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2018). Data lifecycle challenges in production machine learning: A survey. SIGMOD Record, 47(2), 17–28. 10.1145/3299887.3299891
  • [14]. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781–1794. 10.14778/3229863.3229867
  • [15]. Sudhir Vishnubhatla. (2017). Migrating Legacy Information Management Systems to AWS and GCP: Challenges, Hybrid Strategies, and a Dual-Cloud Readiness Playbook. In International Journal of Scientific Research & Engineering Trends (Vol. 3, Number 6). Zenodo. 10.5281/zenodo.17298069
  • [16]. Taleb, I., Serhani, M. A., & Dssouli, R. (2018). Big data quality: A survey. In 2018 IEEE International Congress on Big Data (BigData Congress) (pp. 166–173). 10.1109/BigDataCongress.2018.00029
  • [17]. Kranthi Kumar Routhu. (2018). Seamless HR Finance Interoperability: A Unified Framework through Oracle Integration Cloud. In International Journal of Science, Engineering and Technology (Vol. 6, Number 1). Zenodo. 10.5281/zenodo.17292100
  • [18]. Salehi, M., & Rashidi, L. (2018). A survey on anomaly detection in evolving data: With application to forest fire risk prediction. ACM SIGKDD Explorations Newsletter, 20(1), 13–23. 10.1145/3229329.3229332
  • [19]. Padur, S. K. R. (2016). Online patching and beyond: A practical blueprint for Oracle EBS R12.2 upgrades. International Journal of Scientific Research in Science, Engineering and Technology, 2(3), 1028–1039. 10.32628/IJSRSET1848864
  • [20]. Parasa, M. (2019). A modern recruitment intelligence framework using predictive scoring and adaptive talent pooling in SAP SuccessFactors. International Journal of Science, Engineering and Technology, 7(4). 10.5281/zenodo.17695684
  • [21]. Chu, X., Ilyas, I. F., Krishnan, S., & Wang, J. (2016). Data cleaning: Overview and emerging challenges. In Proceedings of the 2016 ACM SIGMOD International Conference on Management of Data (pp. 2201–2206). 10.1145/2882903.2912574
  • [22]. Ilyas, I. F., & Chu, X. (2018). Data quality: The role of empiricism. ACM Journal of Data and Information Quality, 9(2), 8. 10.1145/3186549.3186559
  • [23]. Mirzaie, M., Behkamal, B., & Paydar, S. (2019). Big data quality: A systematic literature review and future research directions. arXiv preprint arXiv:1904.05353. 10.48550/arXiv.1904.05353
  • [24]. Timmerman, Y., & Bronselaer, A. (2019). Measuring data quality in information systems research. Decision Support Systems, 126, 113138. 10.1016/j.dss.2019.113138