Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R Software
Authors/Creators
- 1. Department of Basic Sciences, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.
Contributors
Contact person:
Researcher (3):
- 1. Department of Basic Sciences, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.
- 2. Department of Management Studies, N.S Raju Institute of Technology, Autonomous, Visakhapatnam (Andhra Pradesh), India.
- 3. Department of Computer Science and Engineering Lendi Institute of Technology, Autonomous Vizianagaram (Andhra Pradesh), India.
Description
Abstract: Predicting customer churn is essential for improving retention and supporting long-term business growth. In this study, we compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R. Our approach included data preprocessing, exploratory analysis, model development, performance evaluation, and further analysis. We developed and evaluated four classification algorithms: Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting, using an 80:20 train-test split. We assessed each model's accuracy, precision, recall, and F1-score. Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset. Feature importance analysis indicated that contract type, customer tenure, monthly charges, total charges, and internet service were the key factors influencing churn. These results suggest that explainable machine learning offers both strong predictive performance and greater transparency. The R-based framework we present provides a practical, reproducible approach to support customer retention strategies and help managers make evidence-based decisions in customer relationship management.
Files
L189312120826.pdf
Files
(1.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:328387ec3195ed75ab0874615facbc2a
|
1.0 MB | Preview Download |
Additional details
Identifiers
- DOI
- 10.35940/ijmh.L1893.12120826
- EISSN
- 2394-0913
Dates
- Accepted
-
2026-08-15Manuscript received on 28 July 2026 | First Revised Manuscript received on 01 August 2026 | Second Revised Manuscript received on 10 August 2026 | Manuscript Accepted on 15 August 2026 | Manuscript published on 30 August 2026.
References
- Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. DOI: 10.1109/ACCESS.2018.2870052
- Alloghani, M., Al-Jumeily, D., Hussain, A., Aljaaf, A. J., Mustafina, J., & Petrov, E. (2020). Decision tree algorithms: A survey. In D. Al Jumeily et al. (Eds.), Artificial Intelligence Review. Springer. https://www.academia.edu/97048115/
- Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. DOI:10.1016/j.inffus.2019.12.012
- Bzdok, D., Altman, N., & Krzywinski, M. (2018). Statistics versus machine learning. Nature Methods, 15(4), 233–234. DOI:10.1038/nmeth.4642
- Carvalho, D. V., Pereira, E. M., & Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics, 8(8), Article 832. DOI: 10.3390/electronics8080832
- Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). DOI: 10.1145/2939672.2939785
- Chen, T., He, T., Benesty, M., Khotilovich, V., Tang, Y., Cho, H., Mitchell, R., Cano, I., Zhou, T., Li, M., Xie, J., Lin, M., Geng, Y., & Li, Y. (2024). xgboost: Extreme Gradient Boosting (R package Version 1.7.10.1). https://CRAN.R-project.org/
- Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21, Article 6. DOI: 10.1186/s12864-019-6413-7
- Duan, Y., Edwards, J. S., & Dwivedi, Y. K. (2019). Artificial intelligence for decision making in the era of big data—Evolution, challenges and research agenda. International Journal of Information Management, 48, 63–71. DOI: 10.1016/j.ijinfomgt.2019.01.021
- Géron, A. (2023). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly Media.
- Hastie, T., Tibshirani, R., & Friedman, J. (2021). The elements of statistical learning: Data mining, inference, and prediction (2nd corrected ed.). Springer. DOI: 10.1007/978-0-387-84858-7
- IBM. (2013). Telco customer churn [Data set]. Kaggle. https://www.kaggle.com. Work remains significant, see the declaration
- James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. DOI: 10.1007/978-1-0716-1418-1
- Kotu, V., & Deshpande, B. (2019). Data science: Concepts and practice (2nd ed.). Morgan Kaufmann. DOI: 10.1016/C2018-0-01769-6
- Kumar, V., & Reinartz, W. (2016). Creating enduring customer value. Journal of Marketing, 80(6), 36–68. 10.1509/jm.15.0414
- Lemon, K. N., & Verhoef, P. C. (2016). Understanding customer experience throughout the customer journey. Journal of Marketing, 80(6), 69–96. DOI: 10.1509/jm.15.0420
- Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30. https://proceedings.neurips.cc/paper/2017/
- Manzoor, A., Qureshi, M. A., Kidney, E., & Longo, L. (2024). A review on machine learning methods for customer churn prediction and recommendations for business practitioners. IEEE Access, 12, 70434 70463. DOI: 10.1109/ACCESS.2024.3402092
- Molnar, C. (2022). Interpretable machine learning (2nd ed.). Leanpub. https://christophm.github.io/interpretable-ml-book/
- Probst, P., Wright, M. N., & Boulesteix, A.-L. (2019). Hyperparameters and tuning strategies for random forest. WIREs Data Mining and Knowledge Discovery, 9(3), e1301. DOI: 10.1002/widm.1301
- R Core Team. (2025). R: A language and environment for statistical computing (Version 4.5.1) [Computer software]. R Foundation for Statistical Computing. https://www.R-project.org/
- Sarker, I. H. (2022). Machine learning: Algorithms, real-world applications and research directions. SN Computer Science, 3, Article 160. DOI: 10.1007/s42979-021-00865-8
- Therneau, T., & Atkinson, B. (2023). rpart: Recursive partitioning and regression trees (R package Version 4.1-21). https://CRAN.R project.org/package=rpart