Detection of Breast Cancer Patient Mortality Status Using Machine Learning with SMOTE-Based Class Imbalance Treatment
Main Article Content
Abstract
Breast cancer is a disease with a high mortality rate, making early detection of patients at risk of death crucial for supporting medical decision-making. This study aims to develop a model for detecting patients at risk of death by comparing several machine learning algorithms. The research process included data collection, exploratory data analysis, data preprocessing, feature engineering, oversampling, modeling, and model evaluation. Data balancing was performed using SMOTE, and model optimization was conducted through hyperparameter tuning using GridSearchCV. The results show that Random Forest combined with GridSearchCV delivers the best performance in prediction, with an accuracy of 0.7491, precision of 0.2953, recall of 0.4634, an F1-score of 0.3608, and an ROC-AUC of 0.7030. This study demonstrates that the combination of Random Forest, SMOTE, and GridSearchCV is capable of optimizing the performance of the prediction model.
Article Details

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
References
World Health Organization, “Breast cancer.” 2024.
N. Harbeck and M. Gnant, “Breast cancer,” Lancet, vol. 389, no. 10074, pp. 1134–1150, 2017.
F. Bray et al., “Global cancer statistics 2022: {GLOBOCAN} estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA. Cancer J. Clin., vol. 74, no. 3, pp. 229–263, 2024.
K. Kourou, T. P. Exarchos, K. P. Exarchos, M. V Karamouzis, and D. I. Fotiadis, “Machine learning applications in cancer prognosis and prediction,” Comput. Struct. Biotechnol. J., vol. 13, pp. 8–17, 2015.
F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.
D. Delen, G. Walker, and A. Kadam, “Predicting breast cancer survivability: A comparison of three data mining methods,” Artif. Intell. Med., vol. 34, no. 2, pp. 113–127, 2005.
M. D. Ganggayah, N. A. Taib, Y. C. Har, P. Lio, and S. K. Dhillon, “Predicting factors for survival of breast cancer patients using machine learning techniques,” BMC Med. Inform. Decis. Mak., vol. 19, p. 48, 2019.
L. Breiman, “Random forests,” Mach. Learn., vol. 45, pp. 5–32, 2001.
T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” Proc. ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., vol. 13-17-Augu, pp. 785–794, 2016.
H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, 2009.
N. V Chawla, K. W. Bowyer, and L. O. Hall, “SMOTE : Synthetic Minority Over-sampling Technique,” vol. 16, pp. 321–357, 2002.
A. Fernández, S. Garcia, F. Herrera, and N. V Chawla, “{SMOTE} for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary,” J. Artif. Intell. Res., vol. 61, pp. 863–905, 2018.
D. R. Cox, “The regression analysis of binary sequences,” J. R. Stat. Soc. Ser. B, vol. 20, no. 2, pp. 215–242, 1958.
T. M. Cover and P. E. Hart, “Nearest neighbor pattern classification,” IEEE Trans. Inf. Theory, vol. 13, no. 1, pp. 21–27, 1967.
T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006.