Predicting life expectancy of lung cancer patients after thoracic surgery using SMOTE and machine learning approaches

SELLY ANASTASSIA AMELLIA KHARIS, ARMAN HAQQI ANNA ZILI

Abstract


. Lung cancer is a life-threatening condition characterized by the uncontrolled growth and spread of abnormal cells in the lungs. Thoracic surgery is a commonly employed diagnostic and treatment procedure for lung cancer. The objective of this study is to utilize machine learning techniques to predict the life expectancy of lung cancer patients one year after thoraric surgery. The study utilizes the Thoraric  Surgery Data Set, consisting of 454 data, with 385 data representing surviving patients and 69 data representing patients who passed away. Due to an imbalance in the data, the Synthetic Minority Oversampling Technique (SMOTE) process is applied to balance the dataset. Multiple machine learning algorithms, including Random Forest (RF), K-Nearest Neighbor (KNN), and Support Vector Machine (SVM), are employed for prediction. Validation is performed using 5-fold cross validation, repeated three times. The results indicate that the KNN model achieves the highest mean accuracy of 84.80% before the SMOTE process, although all models exhibit a low mean F1-score. Following the SMOTE process, the RF model attains  the highest mean accuracy of 79.52%, while the KNN model demonstrates  the highest mean F1-score of 26.54%. This research contributes valuable insights to clinicians in making informed decisions and improving patient outcomes.


Keywords


machine learning, lung cancer, SMOTE, thoracic surgery

References


Zhang, G.; Liu, Z.; Qin, S.;Li, K. 2015. Decreased expression of SIRT6 promotes tumor cell growth correlates closely with poor prognosis of ovarian cancer. Eur. J. Gynaecol. Oncol. 36(6): 629-632.

Ali, A.; Shamsuddin, S. M.; & Ralescu, A. L. 2013. Classification with class imbalance problem. Int. J. Advance Soft Compu. Appl. 5(3): 1-30

Global Cancer Observatory: Cancer Today. Lyon, France: International Agency for Research on Cancer. gco.iarc.fr/today/fact-sheets-cancers

Susan, S.; Kumar, A. 2021. The balancing trick: Optimized sampling of imbalanced datasets—A brief survey of the recent State of the Art. Engineering Reports, 3(4), e12298. DOI: 10.1002/eng2.12298

Van Hulse, J.; Khoshgoftaar, T. 2009. Knowledge discovery from imbalanced and noisy data. Data Knowl. Eng. 68(12): 1513-1542. DOI: 10.1016/j.datak.2009.08.005

Li, J.; Zhu, Q.; Wu, Q.; Fan, Z. 2021. A Novel oversampling technique for class imbalanced learning based on SMOTE and natural neighbors. Inf. Sci. 565: 438–455. DOI: 10.1016/j.ins.2021.03.041

Liu, Q.; Xue, Y.; Li, G.; Qiu, D.; Zhang, W.; Guo, Z.; & Li, Z. 2023. Application of KM-SMOTE for rockburst intelligent prediction. unn. Undergr. Space Technol. 138: 105180. DOI: 10.1016/j.tust.2023.105180

Sáez, J. A.; Luengo, J.; Stefanowski, J.; & Herrera, F. 2015. SMOTE–IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering. Inf. Sci. 291: 184-203. DOI: 10.1016/j.ins.2014.08.051

Amaral, J. L.; Lopes, A. J.; Jansen, J. M.; Faria, A. C.; Melo, P. L. 2012. Machine learning algorithms and forced oscillation measurements applied to the automatic identification of chronic obstructive pulmonary disease. Comput. Methods Programs Biomed. 105(3): 183-193. DOI: 10.1016/j.cmpb.2011.09.009

Arslan, H.; Arslan, H. 2021. A new COVID-19 detection method from human genome sequences using CpG island features and KNN classifier. Eng. Sci. Technol Int. J. 24(4): 839-847. DOI: 10.1016/j.jestch.2020.12.026

Gray, K. R.; Aljabar, P.; Heckemann, R. A.; Hammers, A.; Rueckert, D.; Alzheimer's Disease Neuroimaging Initiative. 2013. Random forest-based similarity measures for multi-modal classification of Alzheimer's disease. NeuroImage 65, 167-175. DOI: 10.1016/j.neuroimage.2012.09.065

Gao, Y.; Zhu, Z.; & Sun, F. 2022. Increasing prediction performance of colorectal cancer disease status using random forests classification based on metagenomic shotgun sequencing data. Synth. Syst. Biotechnol. 7(1): 574-585. DOI: 10.1016/j.synbio.2022.01.005

Gerardin, E.; Chupin, M.; Cuingnet, R.; Dubois, B.; Lehéricy, S.; Garnero, L.; Colliot, O. 2009. SVM classification of patients with Alzheimer's disease and mild cognitive impairment using hippocampal shape features. NeuroImage 47: S57. DOI: 10.1016/S1053-8119(09)70228-9

Karthik, B. U.; Muthupandi, G. 2023. SVM and CNN based skin tumour classification using WLS smoothing filter. Optik 272: 170337. DOI: 10.1016/j.ijleo.2022.170337

Retnoningsih, E.; Pramudita, R. 2020. Mengenal Machine Learning dengan Teknik Supervised dan Unsupervised Learning Menggunakan Python. Bina Insani ICT J. 7(2): 156–165. DOI: 10.51211/biict.v7i2.1422

Kelleher, J.D.; Mac Namee, B.; D'Arcy, A. 2015. Fundamentals of Machine Learning for Predictive Data Analytics : Algorithms, Worked Examples, and Case Studies. MIT Press.

Fan, Z.; Xie, J. K.; Wang, Z. Y.; Liu, P. C.; Qu, S. J.; Huo, L. 2021. Image Classification Method Based on Improved KNN Algorithm. IOP Publishing, 1930(1), 012009. DOI: 10.1088/1742-6596/1930/1/012009

Bhat, A. D.; Acharya, H. R.; & Srikanth, H. R. 2019. A novel solution to the curse of dimensionality in using KNNs for image classification. In 2019 2nd International Conference on Intelligent Autonomous Systems (ICoIAS), 32-36. DOI: 10.1109/ICoIAS.2019.00012

Qi, Y. 2012. Random forest for bioinformatics. In Ensemble machine learning; Boston: Springer. DOI: 10.1007/978-1-4419-9326-7_11

Hernández, B.; Parnell, A.; Pennington, S. R. 2014. Why have so few proteomic biomarkers “survived” validation? (Sample size and independent validation considerations). Proteomics 14(13-14): 1587-1592. DOI: 10.1002/pmic.201300377

Kharis, S.A.A; Hadi, I; Hasanah, K., A. 2019. Multiclass Classification of Brain Cancer with Multiple Multiclass Artificial Bee Colony Feature Selection and Support Vector Machine. J.Phys.: Conf. Ser., 1417 012015. DOI: 10.1088/1742-6596/1417/1/012015

Pisner, D. A.; Schnyer, D. M. 2020. Support vector machine. In Machine learning. Academic Press; p. 101-121. DOI: 10.1016/B978-0-12-815739-8.00006-7

Razzak, I.; Zafar, K.; Imran, M.; Xu, G. 2020. Randomized nonlinear one-class support vector machines with bounded loss function to detect of outliers for large scale IoT data. Future Gener. Comput. Syst. 112: 715-723. DOI: 10.1016/j.future.2020.05.045

Marek, L.; Konrad, P.; Adam, R.; Jerzy, Z. 2013. Thoracic Surgery Data Set. National Science Foundation. https://archive.ics.uci.edu/ml/datasets/Thoracic+Sur-gery+Data.

Oni, S.; Chen, Z.; Hoban, S.; Jademi, O. 2019. A Comparative Study of Data Cleaning Tools. Int. J. Data Warehous. Min. (IJDWM) 15(4): 48-65. DOI: 10.4018/IJDWM.2019100103

Singh, D.; Singh, B. 2020. Investigating the impact of data normalization on classification performance. Appl. Soft Comput. 97 105524. DOI: 10.1016/j.asoc.2019.105524

Patro, S.G.K.; Sahu, K.K. 2015. Normalization: A Preprocessing Stage. Int. Adv. Res. J. Sci. Eng. Technol. 2(3): 20-22. DOI: 10.17148/IARJSET.2015.2305

Han, J.; Kamber, M.; Pei, J. 2006. Data Mining Concepts and Techniques. San Francisco: Diane Cerra.

Siringoringo, R. 2018. Klasifikasi data tidak seimbang menggunakan algoritma SMOTE dan k-nearest neighbor. J. Inf. Sys. Dev. (ISD) 3(1): 162-173.

Mujumdar, A.; Vaidehi, V. 2019. Diabetes prediction using machine learning algorithms. Procedia Comput. Sci.165: 292-299. DOI: 10.1016/j.procs.2020.01.047.

Powers, D. M. 2020. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv preprint arXiv:2010.16061.

Chicco, D.; Jurman, G. 2020. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21: 1-13. DOI: 10.1186/s12864-019-6413-7.

Rupapara, V.; Rustam, F.; Shahzad, H. F.; Mehmood, A.; Ashraf, I.; Choi, G. S. 2021. Impact of SMOTE on imbalanced text features for toxic comments classification using RVVC model. IEEE Access, 9, 78621-78634. DOI: 10.1109/ACCESS.2021.3083638.

Allouche, O.; Tsoar, A.; & Kadmon, R. 2006. Assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS). J. Appl. Ecol. 43(6): 1223-1232. DOI: 10.1111/j.1365-2664.2006.01214.x.

Danquah, R. A. 2020. Handling Imbalanced data: A case study for binary class problems. arXiv preprint arXiv:2010.04326.

Jeni, L. A.; Cohn, J. F.; De La Torre, F. 2013. Facing imbalanced data--recommendations for the use of performance metrics. In 2013 Humaine association conference on affective computing and intelligent interaction, 245-251. DOI: 10.1109/ACII.2013.47.

Anis, M.; & Ali, M. 2017. Investigating the performance of smote for class imbalanced learning: a case study of credit scoring datasets. Eur. Sci. J. 13(33): 340-353. DOI: 10.19044/esj.2017.v13n33p340.


Full Text: PDF

DOI: 10.24815/jn.v23i3.29144

Refbacks

  • There are currently no refbacks.