Komparasi Kinerja Algoritma XGBoost dengan Reduksi Dimensi PCA pada Klasifikasi Diabetes

Authors

  • Rivale Belano Nusa Institut Teknologi dan Bisnis Asia Malang, Indonesia
  • Rina Dewi Indahsari Institut Teknologi dan Bisnis Asia Malang, Indonesia

DOI:

https://doi.org/10.70609/jusifor.v4i2.8563

Keywords:

PCA, XGBoost, Diabetes, Machine Learning, Klasifikasi, Classification

Abstract

Diabetes is one of the most prevalent chronic diseases worldwide and requires accurate early detection to prevent long-term complications. In the field of medical data analysis, the application of machine learning algorithms such as XGBoost has proven effective in classifying disease risk. This study aims to compare the performance of the XGBoost algorithm before and after applying Principal Component Analysis (PCA) in diabetes risk classification using the Early Stage Diabetes Risk Prediction Dataset. The research stages include data preprocessing involving missing value checking, label encoding, outlier removal, normalization, and followed by the application of PCA with a 90% variance retention threshold. The experimental results show that the XGBoost model without PCA achieved the highest accuracy of 99.04%, while the model with PCA achieved 98.08%. Although the application of PCA slightly reduced accuracy, this technique successfully decreased the number of features and improved computational efficiency without losing important information. Therefore, PCA is proven to be effective in simplifying data complexity while maintaining optimal model performance.

References

[1] I. D. Federation, “IDF Diabetes Atlas, 11th edition Global factsheet,” International Diabetes Federation, 2025. [Daring]. Tersedia pada: http://httpsdiabetesatlas.org

[2] W. H. Organization, “Urgent action needed as global diabetes cases increase four-fold over past decades,” 2024, World Health Organization. [Daring]. Tersedia pada: https://www.who.int/news/item/13-11-2024-urgent-action-needed-as-global-diabetes-cases-increase-four-fold-over-past-decades

[3] W. Li, Y. Peng, dan K. Peng, “Diabetes prediction model based on GA-XGBoost and stacking ensemble algorithm,” PLoS One, vol. 19, no. 9, hlm. e0311222, Sep 2024, [Daring]. Tersedia pada: https://doi.org/10.1371/journal.pone.0311222

[4] D. N. Jawza, M. I. Mazdadi, A. Farmadi, T. H. Saragih, D. Kartini, dan V. Abdullayev, “The Enhancing Diabetes Prediction Accuracy Using Random Forest and XGBoost with PSO and GA-Based Feature Selection ,” Journal of Electronics, Electromedical Engineering, and Medical Informatics, vol. 7, no. 2 SE-Electronics, Feb 2025, doi: 10.35882/jeeemi.v7i2.626.

[5] O. Iparraguirre-Villanueva, K. Espinola-Linares, R. O. Flores Castañeda, dan M. Cabanillas-Carbonell, “Application of Machine Learning Models for Early Detection and Accurate Classification of Type 2 Diabetes,” 2023. doi: 10.3390/diagnostics13142383.

[6] U. C. I. M. L. Repository, “Early Stage Diabetes Risk Prediction Dataset,” 2020, University of California, Irvine. [Daring]. Tersedia pada: https://archive.ics.uci.edu/dataset/529/early%2Bstage%2Bdiabetes%2Brisk%2Bprediction%2Bdataset

[7] U. E. Laila, K. Mahboob, A. W. Khan, F. Khan, dan W. Taekeun, “An Ensemble Approach to Predict Early-Stage Diabetes Risk Using Machine Learning: An Empirical Study,” 2022. doi: 10.3390/s22145247.

[8] M. Tantowen, K. Putra, M. Isnan, dan B. Pardamean, “Principal Component Analysis Implementation on Machine Learning in Diabetes Classification,” Communications in Mathematical Biology and Neuroscience, vol. 2024, hlm. 1–19, 2024, doi: 10.28919/cmbn/8492.

[9] F. Rahman, S. Hossain, J.-J. Tiang, dan A.-A. Nahid, “Diabetes Prediction Using Feature Selection Algorithms and Boosting-Based Machine Learning Classifiers,” 2025. doi: 10.3390/diagnostics15202622.

[10] “Principal Component Analysis for Prediabetes Prediction using Extreme Gradient Boosting (XGBoost),” Scientific Journal of Informatics, vol. 11, no. 3 SE-Articles, hlm. 863–872, doi: 10.15294/sji.v11i3.13416.

[11] R. Abdurrosyid, A. Teguh, dan W. Almais, “JEPIN (Jurnal Edukasi dan Penelitian Informatika) Deteksi Dini Diabetes menggunakan Machine Learning dengan Metode PCA dan XGBoost,” vol. 11, no. 1, hlm. 51–56, 2025.

[12] S. Manjula, N. H. Rajini, dan K. Chokkanathan, “Enhanced chronic kidney disease detection using XGBoost with improved brainstorm optimization for hyperparameter tuning,” Discover Applied Sciences, vol. 7, no. 10, hlm. 1181, 2025, doi: 10.1007/s42452-025-07633-7.

[13] J. Gupta, N. Sharma, dan S. Aggarwal, “Impact of Principal Component Analysis on the Performance of Machine Learning Models for the Prediction of Length of Stay of Patients,” EMITTER International Journal of Engineering Technology, vol. 12, no. 2 SE-Articles, Des 2024, doi: 10.24003/emitter.v12i2.835.

[14] I. T. Jolliffe dan J. Cadima, “Principal component analysis: A review and recent developments,” Philosophical Transactions of the Royal Society A, vol. 374, no. 2065, hlm. 20150202, 2016, doi: 10.1098/rsta.2015.0202.

[15] T. Chen dan C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” Mar 2016, doi: 10.48550/arXiv.1603.02754.

Published

2025-12-31

How to Cite

Komparasi Kinerja Algoritma XGBoost dengan Reduksi Dimensi PCA pada Klasifikasi Diabetes. (2025). JUSIFOR (Jurnal Sistem Informasi Dan Informatika), 4(2), 320-329. https://doi.org/10.70609/jusifor.v4i2.8563