Implementation of Winsorizing and random oversampling on data containing outliers and unbalanced data with the random forest classification method

FAHREZAL ZUBEDI, BAGUS SARTONO, KHAIRIL ANWAR NOTODIPUTRO

Abstract


Many researchers conduct research using the classification method, to find out the best method for predicting the class of an observation. Some of these studies explain that random forest is the best method. However, the classification of data containing outliers and unbalanced data is a complicated problem. Many researchers are also conducting research to deal with these problems. In this study, we propose a winsorizing to deal with outliers by replacing the outlier values with the upper and lower limit values obtained from the interquartile range method and random oversampling to balance the data. It is also known that cases of the Human Development Index (HDI) in regencies/cities in eastern Indonesia vary widely, so cases of HDI in these areas can be used as case studies of data containing outliers and unbalanced data. The purpose of this study was to compare the performance of the random forest before and after the data were applied to the winsorizing and random oversampling to predict HDI in districts/cities in eastern Indonesia. Classification method random forest after handling data containing outliers and unbalanced data has better performance in terms of accuracy and kappa values, which are 96.43% and 93.41%, respectively. The variables of expenditure per capita and the mean years of schooling are the most important.

Keywords


Winsorizing, Random oversampling, Interquartile range, Random forest, HDI

References


Johnson, R. A.; Wichern, D. W. 2007. Applied Multivariate Statistical Analysis 6th ed. (London : Pearson Education)

Nnamoko, N.; Korkontzelos, I. 2020. Efficient treatment of outliers and class imbalance for diabetes prediction. Artif. Intell. Med. 104 1–12.

Ghosh, D.; Vogt, A. 2012. Outliers: An Evaluation of Methodologies. Proc. Am. Stat. Assoc. 2012 3455-3460.

Ferdowsi, H.; Jagannathan, S.; Zawodniok, M. 2014. An online outlier identification and removal scheme for improving fault detection performance. IEEE Trans. Neural Netw. Learn. Syst. 25 908-919 DOI : 10.1109/TNNLS.2013.2283456.

Zhu, C.; Idemudia, C. U.; Feng, W. 2019. Improved logistic regression model for diabetes prediction by integrating PCA and K-means techniques. Inform. Med. Unlocked 17 100179 DOI : 10.1016/j.imu.2019.100179.

Xu, X.; Liu, H.; Li, L.; Yao, M. 2018. A comparison of outlier detection techniques for high-dimensional data. Int. J. Comput. Intell. Syst. 11 652–662.

Reifman, A.; Garrett, K. 2010. Winsorize. In: N J Salkind, editors.. Encyclopedia of Research Design. Thousand Oaks (CA): Sage Publishing. p. 1636–1638.

Skryjomski, P.; Krawczyk, B. 2017. Influence of minority class instance types on SMOTE imbalanced data oversampling. Proc. Mach. Learn. Res. 74 7–21.

Whaley, D. L. 2005. The Interquartile Range: Theory and Estimation [Master's Thesis, East Tennessee State University]. Electronic Theses and Dissertations. https://dc.etsu.edu/etd/1030

Mohammed, R.; Rawashdeh, J.; Abdullah, M. 2020. Machine Learning with Oversampling and Undersampling Techniques: Overview Study and Experimental Results. Int. Conf. Inf. Commun. Syst. 2020 243-248 DOI : 10.1109/ICICS49469.2020.239556.

Batista, G. E. A. P. A.; Pati, R. C.; Monard, M. C. 2004. A study of the behavior of several methods for balancing machine learning training data. ACM SIGKDD Explor. News. 6 20-29.

Santoso, B.; Wijayanto, H.; Notodiputro, K. A.; Sartono, B. 2018. A comparative study of synthetic over-sampling method to improve the classification of poor households in yogyakarta province. IOP Conf. Ser. Earth Environ. Sci. 187 012048.

Han, J.; Micheline, K.; Pei, J. 2014. Data mining concepts and techniques 3rd ed. (USA : Morgan Kaufmann)

Pratiwi, A.; Notodiputro, K. A.; Wijayanto, H. 2018. Pemodelan Loyalitas Konsumen Susu Pertumbuhan dalam Mengikuti Program Rewards Menggunakan Metode Random Forest dan Neural Network. Xplore: J. Stat. 2 41-48 DOI : 10.29244/xplore.v2i2.104.

BPS. 2014. Indeks Pembangunan Manusia 2013. (Jakarta : Badan Pusat Statistik RI)

Breiman, L. 2001. Random Forests. Mach. Learn. 45 5–32 DOI : 10.1023/A:1010950718922.

Triscowati, D. W.; Sartono, B.; Kurnia, A.; Domiri, D. D.; Wijayanto, A. W. 2019. Classification of Rice-Plant Growth Phase using Supervised Random Forest Method Based On Landsat-8 Multitemporal Data. Int. J. Remote Sens. Earth Sci. 16 1–11.

Sandri, M.; Zuccolotto, P. 2006. Variable Selection Using Random Forests. In: Zani S. Cerioli A. Riani M. Vichi M, editors. Data Analysis, Classification and the Forward Search. Studies in Classification, Data Analysis, and Knowledge Organization. Berlin: Springer Publishing. p. 263–270.

Belouafa, S.; Habti, F.; Benhar, S.; Belafkih, B.; Tayane, S.; Hamdouch, S.; Bennamara, A.; Abourriche, A. 2017. Statistical tools and approaches to validate analytical methods: Methodology and practical examples. Int. J. Metrol. Qual. Eng. 8 1-10 DOI : 10.1051/ijmqe/2016030.

Sokolova, M.; Lapalme, G. 2009. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 45 427-437 DOI : 10.1016/j.ipm.2009.03.002.

Trevethan, R. 2017. Sensitivity, Specificity, and Predictive Values: Foundations, Pliabilities, and Pitfalls in Research and Practice. Front. Public Health 5 1-7.

Brodersen, K. H.; Ong, C. S.; Stephan, K. E.; Buhmann, J. M. 2010. The balanced accuracy and its posterior distribution. Proc. Int. Conf. Pattern Recognit. 2010 3121-3124 DOI : 10.1109/ICPR.2010.764.

Kohavi, R. 1995. A study of cross-validation and bootstrap for accuracy estimation and model selection. Proc. Int. Joint Conf. Artif. Intell. 2 1137–1143.


Full Text: PDF

DOI: 10.24815/jn.v22i2.25499

Refbacks

  • There are currently no refbacks.