Predicting life expectancy of lung cancer patients after thoracic surgery using SMOTE and machine learning approaches
Abstract
. Lung cancer is a life-threatening condition characterized by the uncontrolled growth and spread of abnormal cells in the lungs. Thoracic surgery is a commonly employed diagnostic and treatment procedure for lung cancer. The objective of this study is to utilize machine learning techniques to predict the life expectancy of lung cancer patients one year after thoraric surgery. The study utilizes the Thoraric Surgery Data Set, consisting of 454 data, with 385 data representing surviving patients and 69 data representing patients who passed away. Due to an imbalance in the data, the Synthetic Minority Oversampling Technique (SMOTE) process is applied to balance the dataset. Multiple machine learning algorithms, including Random Forest (RF), K-Nearest Neighbor (KNN), and Support Vector Machine (SVM), are employed for prediction. Validation is performed using 5-fold cross validation, repeated three times. The results indicate that the KNN model achieves the highest mean accuracy of 84.80% before the SMOTE process, although all models exhibit a low mean F1-score. Following the SMOTE process, the RF model attains the highest mean accuracy of 79.52%, while the KNN model demonstrates the highest mean F1-score of 26.54%. This research contributes valuable insights to clinicians in making informed decisions and improving patient outcomes.
Keywords
References
Zhang, G.; Liu, Z.; Qin, S.;Li, K. 2015. Decreased expression of SIRT6 promotes tumor cell growth correlates closely with poor prognosis of ovarian cancer. Eur. J. Gynaecol. Oncol. 36(6): 629-632.
Ali, A.; Shamsuddin, S. M.; & Ralescu, A. L. 2013. Classification with class imbalance problem. Int. J. Advance Soft Compu. Appl. 5(3): 1-30
Global Cancer Observatory: Cancer Today. Lyon, France: International Agency for Research on Cancer. gco.iarc.fr/today/fact-sheets-cancers
Susan, S.; Kumar, A. 2021. The balancing trick: Optimized sampling of imbalanced datasets—A brief survey of the recent State of the Art. Engineering Reports, 3(4), e12298. DOI: 10.1002/eng2.12298
Van Hulse, J.; Khoshgoftaar, T. 2009. Knowledge discovery from imbalanced and noisy data. Data Knowl. Eng. 68(12): 1513-1542. DOI: 10.1016/j.datak.2009.08.005
Li, J.; Zhu, Q.; Wu, Q.; Fan, Z. 2021. A Novel oversampling technique for class imbalanced learning based on SMOTE and natural neighbors. Inf. Sci. 565: 438–455. DOI: 10.1016/j.ins.2021.03.041
Liu, Q.; Xue, Y.; Li, G.; Qiu, D.; Zhang, W.; Guo, Z.; & Li, Z. 2023. Application of KM-SMOTE for rockburst intelligent prediction. unn. Undergr. Space Technol. 138: 105180. DOI: 10.1016/j.tust.2023.105180
Sáez, J. A.; Luengo, J.; Stefanowski, J.; & Herrera, F. 2015. SMOTE–IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering. Inf. Sci. 291: 184-203. DOI: 10.1016/j.ins.2014.08.051
Amaral, J. L.; Lopes, A. J.; Jansen, J. M.; Faria, A. C.; Melo, P. L. 2012. Machine learning algorithms and forced oscillation measurements applied to the automatic identification of chronic obstructive pulmonary disease. Comput. Methods Programs Biomed. 105(3): 183-193. DOI: 10.1016/j.cmpb.2011.09.009
Arslan, H.; Arslan, H. 2021. A new COVID-19 detection method from human genome sequences using CpG island features and KNN classifier. Eng. Sci. Technol Int. J. 24(4): 839-847. DOI: 10.1016/j.jestch.2020.12.026
Gray, K. R.; Aljabar, P.; Heckemann, R. A.; Hammers, A.; Rueckert, D.; Alzheimer's Disease Neuroimaging Initiative. 2013. Random forest-based similarity measures for multi-modal classification of Alzheimer's disease. NeuroImage 65, 167-175. DOI: 10.1016/j.neuroimage.2012.09.065
Gao, Y.; Zhu, Z.; & Sun, F. 2022. Increasing prediction performance of colorectal cancer disease status using random forests classification based on metagenomic shotgun sequencing data. Synth. Syst. Biotechnol. 7(1): 574-585. DOI: 10.1016/j.synbio.2022.01.005
Gerardin, E.; Chupin, M.; Cuingnet, R.; Dubois, B.; Lehéricy, S.; Garnero, L.; Colliot, O. 2009. SVM classification of patients with Alzheimer's disease and mild cognitive impairment using hippocampal shape features. NeuroImage 47: S57. DOI: 10.1016/S1053-8119(09)70228-9
Karthik, B. U.; Muthupandi, G. 2023. SVM and CNN based skin tumour classification using WLS smoothing filter. Optik 272: 170337. DOI: 10.1016/j.ijleo.2022.170337
Retnoningsih, E.; Pramudita, R. 2020. Mengenal Machine Learning dengan Teknik Supervised dan Unsupervised Learning Menggunakan Python. Bina Insani ICT J. 7(2): 156–165. DOI: 10.51211/biict.v7i2.1422
Kelleher, J.D.; Mac Namee, B.; D'Arcy, A. 2015. Fundamentals of Machine Learning for Predictive Data Analytics : Algorithms, Worked Examples, and Case Studies. MIT Press.
Fan, Z.; Xie, J. K.; Wang, Z. Y.; Liu, P. C.; Qu, S. J.; Huo, L. 2021. Image Classification Method Based on Improved KNN Algorithm. IOP Publishing, 1930(1), 012009. DOI: 10.1088/1742-6596/1930/1/012009
Bhat, A. D.; Acharya, H. R.; & Srikanth, H. R. 2019. A novel solution to the curse of dimensionality in using KNNs for image classification. In 2019 2nd International Conference on Intelligent Autonomous Systems (ICoIAS), 32-36. DOI: 10.1109/ICoIAS.2019.00012
Qi, Y. 2012. Random forest for bioinformatics. In Ensemble machine learning; Boston: Springer. DOI: 10.1007/978-1-4419-9326-7_11
Hernández, B.; Parnell, A.; Pennington, S. R. 2014. Why have so few proteomic biomarkers “survived” validation? (Sample size and independent validation considerations). Proteomics 14(13-14): 1587-1592. DOI: 10.1002/pmic.201300377
Kharis, S.A.A; Hadi, I; Hasanah, K., A. 2019. Multiclass Classification of Brain Cancer with Multiple Multiclass Artificial Bee Colony Feature Selection and Support Vector Machine. J.Phys.: Conf. Ser., 1417 012015. DOI: 10.1088/1742-6596/1417/1/012015
Pisner, D. A.; Schnyer, D. M. 2020. Support vector machine. In Machine learning. Academic Press; p. 101-121. DOI: 10.1016/B978-0-12-815739-8.00006-7
Razzak, I.; Zafar, K.; Imran, M.; Xu, G. 2020. Randomized nonlinear one-class support vector machines with bounded loss function to detect of outliers for large scale IoT data. Future Gener. Comput. Syst. 112: 715-723. DOI: 10.1016/j.future.2020.05.045
Marek, L.; Konrad, P.; Adam, R.; Jerzy, Z. 2013. Thoracic Surgery Data Set. National Science Foundation. https://archive.ics.uci.edu/ml/datasets/Thoracic+Sur-gery+Data.
Oni, S.; Chen, Z.; Hoban, S.; Jademi, O. 2019. A Comparative Study of Data Cleaning Tools. Int. J. Data Warehous. Min. (IJDWM) 15(4): 48-65. DOI: 10.4018/IJDWM.2019100103
Singh, D.; Singh, B. 2020. Investigating the impact of data normalization on classification performance. Appl. Soft Comput. 97 105524. DOI: 10.1016/j.asoc.2019.105524
Patro, S.G.K.; Sahu, K.K. 2015. Normalization: A Preprocessing Stage. Int. Adv. Res. J. Sci. Eng. Technol. 2(3): 20-22. DOI: 10.17148/IARJSET.2015.2305
Han, J.; Kamber, M.; Pei, J. 2006. Data Mining Concepts and Techniques. San Francisco: Diane Cerra.
Siringoringo, R. 2018. Klasifikasi data tidak seimbang menggunakan algoritma SMOTE dan k-nearest neighbor. J. Inf. Sys. Dev. (ISD) 3(1): 162-173.
Mujumdar, A.; Vaidehi, V. 2019. Diabetes prediction using machine learning algorithms. Procedia Comput. Sci.165: 292-299. DOI: 10.1016/j.procs.2020.01.047.
Powers, D. M. 2020. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv preprint arXiv:2010.16061.
Chicco, D.; Jurman, G. 2020. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21: 1-13. DOI: 10.1186/s12864-019-6413-7.
Rupapara, V.; Rustam, F.; Shahzad, H. F.; Mehmood, A.; Ashraf, I.; Choi, G. S. 2021. Impact of SMOTE on imbalanced text features for toxic comments classification using RVVC model. IEEE Access, 9, 78621-78634. DOI: 10.1109/ACCESS.2021.3083638.
Allouche, O.; Tsoar, A.; & Kadmon, R. 2006. Assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS). J. Appl. Ecol. 43(6): 1223-1232. DOI: 10.1111/j.1365-2664.2006.01214.x.
Danquah, R. A. 2020. Handling Imbalanced data: A case study for binary class problems. arXiv preprint arXiv:2010.04326.
Jeni, L. A.; Cohn, J. F.; De La Torre, F. 2013. Facing imbalanced data--recommendations for the use of performance metrics. In 2013 Humaine association conference on affective computing and intelligent interaction, 245-251. DOI: 10.1109/ACII.2013.47.
Anis, M.; & Ali, M. 2017. Investigating the performance of smote for class imbalanced learning: a case study of credit scoring datasets. Eur. Sci. J. 13(33): 340-353. DOI: 10.19044/esj.2017.v13n33p340.
DOI: 10.24815/jn.v23i3.29144
Refbacks
- There are currently no refbacks.


Universitas Syiah Kuala (recognizedly abbreviated as