Published June 9, 2026
| Version v1
Journal article
Open
COMPARATIVE ANALYSIS AND IMPLEMENTATION OF MULTI-LEVEL NORMALIZATION TECHNIQUES IN NETWORK DATABASES
Authors/Creators
- 1. Tashkent State University of Economics
Description
Modern network databases integrate data from heterogeneous sources, which differ significantly in format, scale, encoding, and semantic structure. The challenge of normalizing such data across multiple levels — record, field, and value — remains a critical step in data integration pipelines. This paper presents a comparative analysis and practical implementation of multi-level normalization techniques across twelve distinct data types encountered in network databases: numeric/decimal, integer, text, categorical, temporal, boolean, binary/file, spatial, monetary, array/set, JSON/XML, and null values. For each data type, we systematically compare candidate normalization methods — including min-max scaling, Z-score standardization, one-hot encoding, label encoding, Unicode normalization, ISO 8601 formatting, and log-scaling — identifying the conditions under which each approach is most appropriate. We highlight the trade-offs between methods in terms of distribution preservation, semantic consistency, and computational cost. A Python-based implementation is evaluated on a synthetic dataset of 1,000 records, demonstrating that applying type-aware multi-level normalization reduces numerical attribute dispersion by approximately 70% and achieves 100% structural consistency in text and semi-structured fields. The proposed framework supports downstream tasks including ontological mapping, semantic integration, and 3D data visualization.
Files
V4I315.pdf
Files
(751.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:76f8329329ace8c2b06904d756b400b0
|
751.2 kB | Preview Download |
Additional details
References
- Dong, Y., Dragut, E. C., & Meng, W. (2019). Normalization of duplicate records from multiple sources. IEEE Transactions on Knowledge and Data Engineering, 31(4), 769-782. https://doi.org/10.1109/TKDE.2018.2875227.
- Little, R. J. A., & Rubin, D. B. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley.
- Mitchell, T. M. (1997). Machine Learning. McGraw-Hill.
- Dragut, E. C., & Meng, W. (2006). Meaningful labeling of integrated query interfaces. In Proceedings of the 32nd International Conference on Very Large Data Bases (VLDB '06) (pp. 679-690). VLDB Endowment.
- Tejada, S., Knoblock, C. A., & Minton, S. (2001). Learning object identification rules for information integration. Information Systems, 26(8), 607-633. https://doi.org/10.1016/S0306-4379(01)00033-4
- Wick, M. L., Rohanimanesh, K., Schultz, K., & McCallum, A. (2008). A unified aproach for schema matching, coreference and canonicalization. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '08) (pp. 722-730). ACM. https://doi.org/10.1145/1401890.1401977
- Ballé, J., Laparra, V., & Simoncelli, E. P. (2015). Density modeling of images using a generalized normalization transformation.
- Culotta, A., Mccallum, A., & Betz, J. (2007). Canonicalization of database records using adaptive edit distance. In Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '07) (pp. 104–113). ACM Press. https://doi.org/10.1145/1281192.1281207.
- Han, J., Kamber, M., & Pei, J. (2011). Data Mining: Concepts and Techniques (3rd ed.). Morgan Kaufmann.
- Abadi, D. J. (2018). The Design and Implementation of Modern Column-Oriented Database Systems. Springer.
- Han, J., Kamber, M., & Pei, J. (2011). Data Mining: Concepts and Techniques (3rd ed.). Morgan Kaufmann Publishers.
- Bishop, C. M. Pattern Recognition and Machine Learning. Springer.
- Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
- Shekhar, S., Chawla, S., Zhang, P., & Lazarz, A. (2015). Spatial Databases: A Tour (2nd ed.). Pearson Education.
- Sankpal, K. A., & Metre, K. V. (2020). A review on data normalization techniques. International Journal of Engineering Research & Technology (IJERT), 9(6), 885–889. https://doi.org/10.17577/IJERTV9IS060915.
- Benjelloun, O., Garcia-Molina, H., Menestrina, D., Su, Q., Whang, S. E., & Widom, J. (2009). Swoosh: A generic approach to entity resolution. The VLDB Journal, 18(1), 255-276. https://doi.org/10.1007/s00778-008-0098-x.
- Han, J., Kamber, M., & Pei, J. (2012). Data Mining: Concepts and Techniques (3rd ed.). Morgan Kaufmann Publishers
- Pang-Ning Tan et al. (2019). Introduction to Data Mining. Pearson Education.