Logistic Regression for Customer Churn Analysis: The Role of Data Proportion and Class Balancing in Model Performance

Authors

  • Huimin Pan

DOI:

https://doi.org/10.61173/x28wbd81

Keywords:

Customer churn prediction, logistic regression, machine learning

Abstract

Customer churn prediction has become a critical challenge for the banking industry, as retaining existing clients is often more cost-effective than acquiring new ones. This study applies logistic regression to the Bank Customer Churn dataset (Kaggle, 2017), which contains 10,000 records and 13 features. After preprocessing, including removal of irrelevant identifiers, categorical encoding, and feature selection, seven key variables were retained. The research explores the impact of training set proportions (50%, 60%, 70%, 80%) and class balancing techniques (oversampling vs. no balancing) on model performance. Results show that model accuracy remains stable at approximately 0.81 across different training ratios, suggesting that increasing training size does not yield significant gains. In contrast, balancing the dataset reduced overall accuracy (0.72 vs. 0.81), reflecting the trade-off between accuracy and minority class recall in imbalanced classification. Logistic regression coefficients further revealed interpretable patterns: customers in Germany had higher churn odds, while active membership, longer tenure, and multiple product ownership reduced churn likelihood. These findings contribute to understanding how data preprocessing choices affect churn modeling outcomes and provide actionable insights for banking institutions seeking to strengthen retention strategies.

References

[1] Weiss GM, Provost F. Learning when training data are costly: The effect of class distribution on tree induction. J Artif Intell Res. 2003 Oct 1;19:315-54. Fangfang. Research on power load forecasting based on improved BP neural network [dissertation]. Harbin: Harbin Institute of Technology; 2011.

[2] Albisua I, Arbelaitz O, Gurrutxaga I, Lasarguren A, Muguerza J, Pérez JM. The quest for the optimal class distribution: an approach for enhancing the effectiveness of learning via resampling methods for imbalanced data sets. Prog Artif Intell. 2013 Mar;2(1):45-63.

[3] Tantithamthavorn C, Hassan AE, Matsumoto K. The impact of class rebalancing techniques on the performance and interpretation of defect prediction models. IEEE Trans Softw Eng. 2018 Oct 17;46(11):1200-19.

[4] Raja JB, Sandhya G, Peter SS, Karthik R, Femila F. Exploring effective feature selection methods for telecom churn prediction. Int J Innov Technol Explor Eng. 2020;9(3).

[5] Wu X, Meng S. E-commerce customer churn prediction based on improved SMOTE and AdaBoost. In: 2016 13th International Conference on Service Systems and Service Management (ICSSSM). IEEE; 2016. p. 1-5.

[6] Amin A, Rahim F, Ali I, Khan C, Anwar S. A comparison of two oversampling techniques (SMOTE vs MTDF) for handling class imbalance problem: A case study of customer churn prediction. In: New Contributions in Information Systems and Technologies. Vol 1. Cham: Springer International Publishing; 2015. p. 215-25.

[7] Harshini A, Nallagorla NSR, Bathula ST, Chandra S, Syed S. Improving customer churn prediction accuracy: A SMOTE- based approach. In: 2024 8th International Conference on Inventive Systems and Control (ICISC). IEEE; 2024. p. 215-22.

[8] KartikSaini18. Churn Bank Customer [Internet]. 2025. Available from: https://www.kaggle.com/datasets/kartiksaini18/ churn-bank-customer

[9] LaValley MP. Logistic regression. Circulation. 2008;117(18):2395-9.

[10] Hosmer DW, Lemeshow S, Sturdivant RX. Applied logistic regression. 3rd ed. Hoboken: Wiley; 2013.

Downloads

Published

2025-10-23