Comparative Analysis to Predict Reading Literacy based on PISA 2022 using Gradient Boosted Decision Trees and Extreme Gradient Boosting
DOI:
https://doi.org/10.70609/gtech.v9i1.6257Keywords:
GBDT, Literacy, PISA, XGBoostAbstract
Reading is a fundamental skill essential for interdisciplinary understanding and serves as a crucial indikator of a nation’s educational quality. PISA provides an international evaluation of students' reading literacy across various countries, including Indonesia. This study compares the performance of Gradient Boosting Decision Trees (GBDT) and Extreme Gradient Boosting (XGBoost), two widely recognized machine learning algorithms for predicting reading literacy, utilizing PISA 2022 data from 12.853 Indonesian students and 59 variables from the Student Questionnaire Data File. GBDT achieved R² of 0.5106, with optimal parameters (n_estimators = 150, learning_rate = 0.2, max_depth = 3, subsample = 0.9). XGBoost reached a higher R² of 0.5247, with parameters (n_estimators = 1000, learning_rate = 0.01, max_depth = 7, colsample_bytree = 0.3, min_child_weight = 20, gamma = 1, alpha = 0), indicating XGBoost's superior performance in predicting reading literacy. Further analysis revealed that the most significant variables in the GBDT model included students' access to technology at home, extracurricular creative activities, socioeconomic status, school involvement in sustainable development, and problem-solving skills. In contrast, significant variables in the XGBoost model included family support, socioeconomic status, school belongingness, family environment's effectiveness in fostering creativity, and student imagination.
References
Ahmad, G. N., Fatima, H., Shafiullah, Salah Saidi, A., & Imdadullah. (2022). Efficient Medical Diagnosis of Human Heart Diseases Using Machine Learning Techniques with and Without GridSearchCV. IEEE Access, 10(August), 80151–80173. https://doi.org/10.1109/ACCESS.2022.3165792
Breiman. (2019). Classification and Regression Trees. In Sustainability (Switzerland) (Vol. 11, Issue 1). http://scioteca.caf.com/bitstream/handle/123456789/1091/RED2017-Eng-8ene.pdf?sequence=12&isAllowed=y%0Ahttp://dx.doi.org/10.1016/j.regsciurbeco.2008.06.005%0Ahttps://www.researchgate.net/publication/305320484_SISTEM_PEMBETUNGAN_TERPUSAT_STRATEGI_MELESTARI
Bu, Y., & Chen, F. (2023). What key contextual factors contribute to students’ reading literacy among top-performing countries and economies? Statistical and machine learning analyses. International Journal of Educational Research, 122(July), 102267. https://doi.org/10.1016/j.ijer.2023.102267
Dilekçi, A. (2022). Evaluation of Turkey’s PISA Reading Literacy Scores. International Journal of Education and Literacy Studies, 10(1), 138. https://doi.org/10.7575/aiac.ijels.v.10n.1p.138
Farhangi, F., Sadeghi-Niaraki, A., Razavi-Termeh, S. V., & Choi, S. M. (2021). Evaluation of tree-based machine learning algorithms for accident risk mapping caused by driver lack of alertness at a national scale. Sustainability (Switzerland), 13(18). https://doi.org/10.3390/su131810239
Guoqing Chen, T. Z. (2023). Combined Prediction Model of Gasoline Octane Loss Value Based on XGboost-GBDT Combined Prediction. Sustainability (Switzerland), 11(1), 1–14. http://scioteca.caf.com/bitstream/handle/123456789/1091/RED2017-Eng-8ene.pdf?sequence=12&isAllowed=y%0Ahttp://dx.doi.org/10.1016/j.regsciurbeco.2008.06.005%0Ahttps://www.researchgate.net/publication/305320484_SISTEM_PEMBETUNGAN_TERPUSAT_STRATEGI_MELESTARI
Su, Q., Chen, L., & Qian, L. (2024). Optimization of big data analysis resources supported by XGBoost algorithm: Comprehensive analysis of industry 5.0 and ESG performance. Measurement: Sensors, 36(October), 101310. https://doi.org/10.1016/j.measen.2024.101310.
Katagiri, K., & Fujii, T. (2022). Partitioned Path Loss Models Based on Coefficient of Determination. International Conference on Information Networking, 2022-Janua, 198–203. https://doi.org/10.1109/ICOIN53446.2022.9687163
Kuan, S. I., Kim, J., Kwon, O. H., & Song, H. J. (2022). Canopy�K-means Combined Collaborative Filtering Using RMSE-minimization. Proceedings - 2022 IEEE International Conference on Big Data and Smart Computing, BigComp 2022, 31–34. https://doi.org/10.1109/BigComp54360.2022.00016
Liu, H., Chen, X., & Liu, X. (2022). Factors influencing secondary school students’ reading literacy: An analysis based on XGBoost and SHAP methods. Frontiers in Psychology, 13(September). https://doi.org/10.3389/fpsyg.2022.948612
Mahdavi, M. (2024). Efficient Hardware Acceleration of Mean Squared Error Calculation Through In-Memory Computing. 2024 IEEE International Conference on Omni-Layer Intelligent Systems, COINS 2024, Imc, 1–6. https://doi.org/10.1109/COINS61597.2024.10622442
Moitra, A. (2024). MSBoost : Using Model Selection with Multiple Base Estimators for Gradient Boosting.
Robeson, S. M., & Willmott, C. J. (2023). Decomposition of the mean absolute error (MAE) into systematic and unsystematic components. PLoS ONE, 18(2 February), 1–8. https://doi.org/10.1371/journal.pone.0279774
Shen, E. Z. (2022). Short-time cab speed prediction model based on XGBoost. Research Square, 1, 0–13.
Smiti, S., Soui, M., & Ghedira, K. (2024). Tri-XGBoost model improved by BLSmote-ENN: an interpretable semi-supervised approach for addressing bankruptcy prediction. Knowledge and Information Systems, 66(7), 3883–3920. https://doi.org/10.1007/s10115-024-02067-w
State, T. (2022). PISA 2022 Result The State of Learning and Equity in Education: Vol. I (Issue 2).
Tağa, T. (2023). What does PISA Assess in Reading Literacy ? Misconceptions and Misuses. c.
Wang, S., & Ma, J. (2023). A novel GBDT-BiLSTM hybrid model on improving day-ahead photovoltaic prediction. Scientific Reports, 13(1), 1–13. https://doi.org/10.1038/s41598-023-42153-7
Wang, Y., King, R., Haw, J., & Leung, S. on. (2023). What explains Macau students’ achievement? An integrative perspective using a machine learning approach (¿Cuál es la explicación del rendimiento de los estudiantes macaenses? Una perspectiva integradora mediante la adopción del enfoque del aprendizaje aut. Infancia y Aprendizaje, 46(1), 71–108. https://doi.org/10.1080/02103702.2022.2149120
Wang, Y., Yan, Z., & Xing, L. (2021). A Movie Score Prediction Model Based on XGBoost Algorithm. Proceedings - 2021 International Conference on Culture-Oriented Science and Technology, ICCST 2021, 486–491. https://doi.org/10.1109/ICCST53801.2021.00108
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Hary Susanto, Ema Utami

This work is licensed under a Creative Commons Attribution 4.0 International License.









