Comparative Analysis to Predict Reading Literacy based on PISA 2022 using Gradient Boosted Decision Trees and Extreme Gradient Boosting

Authors

  • Hary Susanto Universitas Amikom Yogyakarta, Indonesia
  • Ema Utami Universitas Amikom Yogyakarta, Indonesia

DOI:

https://doi.org/10.70609/gtech.v9i1.6257

Keywords:

GBDT, Literacy, PISA, XGBoost

Abstract

Reading is a fundamental skill essential for interdisciplinary understanding and serves as a crucial indikator of a nation’s educational quality. PISA provides an international evaluation of students' reading literacy across various countries, including Indonesia. This study compares the performance of Gradient Boosting Decision Trees (GBDT) and Extreme Gradient Boosting (XGBoost), two widely recognized machine learning algorithms for predicting reading literacy, utilizing PISA 2022 data from 12.853 Indonesian students and 59 variables from the Student Questionnaire Data File. GBDT achieved R² of 0.5106, with optimal parameters (n_estimators = 150, learning_rate = 0.2, max_depth = 3, subsample = 0.9). XGBoost reached a higher R² of 0.5247, with parameters (n_estimators = 1000, learning_rate = 0.01, max_depth = 7, colsample_bytree = 0.3, min_child_weight = 20, gamma = 1, alpha = 0), indicating XGBoost's superior performance in predicting reading literacy. Further analysis revealed that the most significant variables in the GBDT model included students' access to technology at home, extracurricular creative activities, socioeconomic status, school involvement in sustainable development, and problem-solving skills. In contrast, significant variables in the XGBoost model included family support, socioeconomic status, school belongingness, family environment's effectiveness in fostering creativity, and student imagination.

References

Ahmad, G. N., Fatima, H., Shafiullah, Salah Saidi, A., & Imdadullah. (2022). Efficient Medical Diagnosis of Human Heart Diseases Using Machine Learning Techniques with and Without GridSearchCV. IEEE Access, 10(August), 80151–80173. https://doi.org/10.1109/ACCESS.2022.3165792

Breiman. (2019). Classification and Regression Trees. In Sustainability (Switzerland) (Vol. 11, Issue 1). http://scioteca.caf.com/bitstream/handle/123456789/1091/RED2017-Eng-8ene.pdf?sequence=12&isAllowed=y%0Ahttp://dx.doi.org/10.1016/j.regsciurbeco.2008.06.005%0Ahttps://www.researchgate.net/publication/305320484_SISTEM_PEMBETUNGAN_TERPUSAT_STRATEGI_MELESTARI

Bu, Y., & Chen, F. (2023). What key contextual factors contribute to students’ reading literacy among top-performing countries and economies? Statistical and machine learning analyses. International Journal of Educational Research, 122(July), 102267. https://doi.org/10.1016/j.ijer.2023.102267

Dilekçi, A. (2022). Evaluation of Turkey’s PISA Reading Literacy Scores. International Journal of Education and Literacy Studies, 10(1), 138. https://doi.org/10.7575/aiac.ijels.v.10n.1p.138

Farhangi, F., Sadeghi-Niaraki, A., Razavi-Termeh, S. V., & Choi, S. M. (2021). Evaluation of tree-based machine learning algorithms for accident risk mapping caused by driver lack of alertness at a national scale. Sustainability (Switzerland), 13(18). https://doi.org/10.3390/su131810239

Guoqing Chen, T. Z. (2023). Combined Prediction Model of Gasoline Octane Loss Value Based on XGboost-GBDT Combined Prediction. Sustainability (Switzerland), 11(1), 1–14. http://scioteca.caf.com/bitstream/handle/123456789/1091/RED2017-Eng-8ene.pdf?sequence=12&isAllowed=y%0Ahttp://dx.doi.org/10.1016/j.regsciurbeco.2008.06.005%0Ahttps://www.researchgate.net/publication/305320484_SISTEM_PEMBETUNGAN_TERPUSAT_STRATEGI_MELESTARI

Su, Q., Chen, L., & Qian, L. (2024). Optimization of big data analysis resources supported by XGBoost algorithm: Comprehensive analysis of industry 5.0 and ESG performance. Measurement: Sensors, 36(October), 101310. https://doi.org/10.1016/j.measen.2024.101310.

Katagiri, K., & Fujii, T. (2022). Partitioned Path Loss Models Based on Coefficient of Determination. International Conference on Information Networking, 2022-Janua, 198–203. https://doi.org/10.1109/ICOIN53446.2022.9687163

Kuan, S. I., Kim, J., Kwon, O. H., & Song, H. J. (2022). Canopy�K-means Combined Collaborative Filtering Using RMSE-minimization. Proceedings - 2022 IEEE International Conference on Big Data and Smart Computing, BigComp 2022, 31–34. https://doi.org/10.1109/BigComp54360.2022.00016

Liu, H., Chen, X., & Liu, X. (2022). Factors influencing secondary school students’ reading literacy: An analysis based on XGBoost and SHAP methods. Frontiers in Psychology, 13(September). https://doi.org/10.3389/fpsyg.2022.948612

Mahdavi, M. (2024). Efficient Hardware Acceleration of Mean Squared Error Calculation Through In-Memory Computing. 2024 IEEE International Conference on Omni-Layer Intelligent Systems, COINS 2024, Imc, 1–6. https://doi.org/10.1109/COINS61597.2024.10622442

Moitra, A. (2024). MSBoost : Using Model Selection with Multiple Base Estimators for Gradient Boosting.

Robeson, S. M., & Willmott, C. J. (2023). Decomposition of the mean absolute error (MAE) into systematic and unsystematic components. PLoS ONE, 18(2 February), 1–8. https://doi.org/10.1371/journal.pone.0279774

Shen, E. Z. (2022). Short-time cab speed prediction model based on XGBoost. Research Square, 1, 0–13.

Smiti, S., Soui, M., & Ghedira, K. (2024). Tri-XGBoost model improved by BLSmote-ENN: an interpretable semi-supervised approach for addressing bankruptcy prediction. Knowledge and Information Systems, 66(7), 3883–3920. https://doi.org/10.1007/s10115-024-02067-w

State, T. (2022). PISA 2022 Result The State of Learning and Equity in Education: Vol. I (Issue 2).

Tağa, T. (2023). What does PISA Assess in Reading Literacy ? Misconceptions and Misuses. c.

Wang, S., & Ma, J. (2023). A novel GBDT-BiLSTM hybrid model on improving day-ahead photovoltaic prediction. Scientific Reports, 13(1), 1–13. https://doi.org/10.1038/s41598-023-42153-7

Wang, Y., King, R., Haw, J., & Leung, S. on. (2023). What explains Macau students’ achievement? An integrative perspective using a machine learning approach (¿Cuál es la explicación del rendimiento de los estudiantes macaenses? Una perspectiva integradora mediante la adopción del enfoque del aprendizaje aut. Infancia y Aprendizaje, 46(1), 71–108. https://doi.org/10.1080/02103702.2022.2149120

Wang, Y., Yan, Z., & Xing, L. (2021). A Movie Score Prediction Model Based on XGBoost Algorithm. Proceedings - 2021 International Conference on Culture-Oriented Science and Technology, ICCST 2021, 486–491. https://doi.org/10.1109/ICCST53801.2021.00108

Downloads

Published

2025-01-16

How to Cite

Comparative Analysis to Predict Reading Literacy based on PISA 2022 using Gradient Boosted Decision Trees and Extreme Gradient Boosting. (2025). G-Tech: Jurnal Teknologi Terapan, 9(1), 390-399. https://doi.org/10.70609/gtech.v9i1.6257

Most read articles by the same author(s)