Classification of household poverty in West Java using the generalized mixed-effects trees model

FARDILLA RAHMAWATI, KHAIRIL ANWAR NOTODIPUTRO, KUSMAN SADIK

Abstract


Dealing with fixed effects and random effects can be accomplished by combining statistical modeling and machine learning techniques. This paper discusses the modeling of fixed effects and random effects using a statistical machine-learning approach. We used the generalized mixed-effects trees (GMET), a tree-based mixed-effect model for dealing with response variables that belong to the exponential family of distributions. In this study, both simulation and actual/empirical data utilized the GMET method to discover data conditions that were appropriate for employing this approach. The simulation data was generated using different response variable generations, as well as different values of the variance of random effect and fixed effect coefficients. The findings indicated that the GMET performs similarly for different response variable generation scenarios. However, it performed better when the fixed effect value and the variance of random effects were large. When applied to the empirical data, the GMET method describes fixed effects and random effects and classifies household poverty status quite well based on the area under curve (AUC) value. It has also revealed that important variables for poverty classification are the number of household members, owning land, the type of main fuel used for cooking, and the main source of water used for drinking. In order to address the socioeconomic disparity that leads to poverty, the government may become concerned about these factors. In addition to that information, the use of regional typology as a random effect in the model has also contributed to the variation of household poverty status. Based on research, the fixed effects in mixed models do not need to be linear and GMET may be employed in grouped data structures, giving the GMET technique the ability to compete with other approaches/methods.


Keywords


fixed effects; household per capita expenditure; machine learning; random effects; supervised learning

References


Gorunescu, F. 2011. Data mining: concepts, models and techniques (Berlin: Springer). DOI: 10.1007/978-3-642-19721-5.

Jacob, E. K. 2004. Classification and categorization: a difference that makes a difference. Libr. Trends. 52 (3) 515–540.

Rokach, L.; Maimon, O. 2005. Decision trees. Data Mining and Knowledge Discovery Handbook. 165–192. DOI: 10.1007/0-387-25465-X_9.

Jijo, B. T.; Abdulazeez, A. M. 2021. Classification based on decision tree algorithm for machine learning. J. Appl. Sci. Technol. Trends. 2 (1) 20–28. DOI: 10.38094/jastt20165.

Breiman, L. 1996. Bagging predictors. Mach. Learn. 24 (2) 123–140. DOI: 10.1007/BF00058655.

Freund, Y.; Schapire, R. E. 1997. A decision-theoretic generalization of on-line learning and an application to boosting. J. Comput. Syst. Sci. 55 (1) 119–139. DOI: 10.1006/jcss.1997.1504.

Breiman, L. 2001. Random forests. Mach. Learn. 45 (1) 5–32. DOI: 10.1023/A:1010933404324.

Fokkema, M.; Edbrooke-Childs, J.; Wolpert, M. 2021. Generalized linear mixed-model (GLMM) trees: a flexible decision-tree method for multilevel and longitudinal data. Psychother. Res. 31 (3) 329–341. DOI: 10.1080/10503307.2020.1785037.

Sela, R. J.; Simonoff, J. S. 2012. RE-EM trees: a data mining approach for longitudinal and clustered data. Mach. Learn. 86 (2) 169–207. DOI: 10.1007/s10994-011-5258-3.

Fontana, L.; Masci, C.; Ieva, F.; Paganoni, A. M. 2021. Performing learning analytics via generalised mixed-effects trees. Data. 6 (74) 1-31. DOI: 10.3390/data6070074.

Anderson, C.; Verkuilen, J.; Johnson, T. 2012. Applied generalized linear mixed models: continuous and discrete data (for the social and behavioral sciences) (Illinois: Springer).

Agresti, A. 2019. An introduction to categorical analysis, 3rd edition (Hoboken: John Wiley & Sons, Inc.).

West, B. T.; Welch, K. B.; Galecki, A. T. 2015. Linear mixed models: a practical guide using statistical software, second edition (New York: Chapman & Hall).

Hajjem, A. 2010. Mixed effects trees and forests for clustered data mixed effects trees and forests for clustered data. (Dissertation, Canada: University of Montreal, 2010). [online] Available at: http://biblos.hec.ca/biblio/theses/003002.PDF [accessed 7 Jun. 2022].

Pellagatti, M.; Masci, C.; Ieva, F.; Paganoni, A. M. 2021. Generalized mixed‐effects random forest: a flexible approach to predict university student dropout. Stat. Anal. Data Min. ASA Data Sci. J. 14 (3) 241–257. DOI: 10.1002/sam.11505.

Gottard, A.; Vannucci, G.; Grilli, L.; Rampichini, C. 2022. Mixed-effect models with trees. Adv. Data Anal. Classif. 17 (2) 431–461. DOI: 10.1007/s11634-022-00509-3.

Arfiani, D. 2020. Berantas kemiskinan (Semarang: Alprin).

Ritonga, H.; Surbakti, I.; Sodikin; Marsisno, W.; Suryaningsih, T.; Rizal, N.; Widya, C.; Ariewidayanti, D.; Taufiq, N. 2012. Pendataan program perlindungan sosial (PPLS) 2011 (Jakarta: Badan Pusat Statistik).

Avenzora, A.; Karyono, Y. 2008. Analisis dan perhitungan tingkat kemiskinan tahun 2008 (Jakarta: Badan Pusat Statistik).

Susanti, S. 2013. Pengaruh produk domestik regional bruto, pengangguran dan indeks pembangunan manusia terhadap kemiskinan di jawa barat dengan menggunakan analisis data panel. J. Mat. Integr. 9 (1) 1–18. DOI: 10.24198/jmi.v9.n1.9374.1-18.

[BPS] Badan Pusat Statistik. 2020. Statistik Indonesia 2020 (Jakarta: Badan Pusat Statistik).

Nanga, M.; HW, E. F.; Rahayuningsih, D.; Dinayanti, E.; Aulia, F. M.; Rismalasari, M.; Hafid, M.; Wahyu, R.; Putra, R.R.; Kartika, V.; Widaryatmo. 2018. Analisis wilayah dengan kemiskinan tinggi (Jakarta: Kedeputian Bidang Kependudukan dan Ketenagakerjaan Kementerian PPN/Bappenas).

Kusuma, M. E.; Muta’ali, L. 2019. Hubungan pembangunan infrastruktur dan perkembangan ekonomi wilayah Indonesia. J. Bumi Indones. DOI: oai:ojs.lib.geo.ugm.ac.id:article/1083.

[BPS] Badan Pusat Statistik. 2021. Statistik Indonesia 2021 (Jakarta: Badan Pusat Statistik).

Isdijoso, W.; Suryahadi, A.; Akhmadi. 2016. Penetapan kriteria dan variabel pendataan penduduk miskin yang komprehensif dalam rangka perlindungan penduduk miskin di kabupaten/kota. The SMERU Research Institute. 1-25. [online] Available at: http://www.smeru.or.id/sites/default/files/publication/cbms_criteria_ind.pdf [accessed 29 Sep. 2022].

Flatt, C.; Jacobs, R. L. 2019. Principle assumptions of regression analysis: testing, techniques, and statistical reporting of imperfect data sets. Adv. Dev. Hum. Resour. 21 (4) 484–502. DOI: 10.1177/1523422319869915.

Ainiyah, N.; Deliar, A.; Virtriana, R. 2016. The classical assumption test to driving factors of land cover change in the development region of northern part of west java. ISPRS - Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. XLI-B6 205–210. DOI: 10.5194/isprsarchives-XLI-B6-205-2016.

Chreiber-Gregory, D.; Bader, K. 2018. Logistic and linear regression assumptions: violation recognition and control. Southeast SAS Users Gr. Conf. 47 (3) 1–22. [online] Available at: https://www.lexjansen.com/sesug/2018/SESUG2018_Paper-247_Final_PDF.pdf [accessed 13 Oct. 2023].


Full Text: PDF

DOI: 10.24815/jn.v23i3.33079

Refbacks

  • There are currently no refbacks.