Application of K-Means and Naïve Bayes Algorithms for Prediction Model of Student Interest Concentration (Case Study: Amikom University Yogyakarta)

Authors

  • Danang Eko Prayogo Universitas Amikom Yogyakarta, Indonesia
  • Kusrini Kusrini Universitas Amikom Yogyakarta, Indonesia

DOI:

https://doi.org/10.70609/gtech.v9i1.6523

Keywords:

Student Concentration, K-Means, Naïve Bayes, Amikom University Yogyakarta

Abstract

Amikom University Yogyakarta, has a Master of Informatics Engineering study program with three concentrations of specialization: Business Intelligence, Digital Information Intelligence, and Intelligence Animation. The choice of concentration by prospective students has been based on subjectivity, not on competence or work experience. To overcome this, this research proposes an algorithm-based concentration prediction and recommendation model to help prospective students choose the appropriate concentration. The dataset is obtained through questionnaires collected from active and inactive students. This research uses the K-Means algorithm for clustering raw data (unsupervised) in order to generate target classes, which are then classified using Naïve Bayes. The clustering process determines concentration labels such as Business Intelligence and others, while the SMOTE technique is used to balance the dataset to avoid data imbalance problems. This approach aims to produce more objective and accurate recommendations in determining student concentrations, reducing the tendency of subjectivity, and increasing the relevance of student competencies to the chosen field of specialization. From this research, the K-Means DBI score is 0.277 and the Naïve Bayes prediction accuracy score is 89%. This research aims to produce more objective and accurate recommendations in determining student concentrations, reducing subjectivity, and increasing the relevance of student competencies to the chosen field of specialization. The proposed model is expected to help universities in designing a more targeted admission strategy, as well as supporting students in making academic decisions that are in accordance with their abilities and interests, thereby increasing the effectiveness of the learning process and the suitability of graduates to the needs of the world of work.

References

Ali, Z. M., Hassoon, N. H., Ahmed, W. S., & Abed, H. N. (2020). The application of data mining for predicting academic performance using K-means clustering and Naïve Bayes classification. International Journal of Psychosocial Rehabilitation, 24(3), 2143–2151. https://doi.org/10.37200/ijpr/v24i3/pr200962

Amalina, N., Suhaimi, D., & Abas, H. (2020). A systematic literature review on supervised machine learning algorithms. Perintis eJournal, 10(1).

Anagora, R., Taufiq, R., Dedi Jubaedi, A., Wirawan, R., & Syah Putra, A. (2022). The classification of phishing websites using Naive Bayes classifier algorithm. International Journal of Science. Retrieved from http://ijstm.inarah.co.id

Byna, A., & Basit, M. (2020). Application of Adaboost method to optimize stroke disease prediction with Naïve Bayes algorithm. Sisfokom Journal (Information and Computer Systems), 9(3), 407–411. https://doi.org/10.32736/sisfokom.v9i3.1023

Fadhil, Z. M. (2021). Hybrid of K-means clustering and Naive Bayes classifier for predicting performance of an employee. Journal of Information Technology, 9(2), 799–807.

Farid, M., Wibowo, S., Puspitasari, N. F., & Satya, B. (2022). Application of data mining and Naïve Bayes algorithm for student concentration selection using classification method. Journal of Information Systems Management (JOISM), 3(2).

Febrianto, A., Achmadi, S., & Sasmito, A. P. (2021). Application of K-Means method for clustering visitors to ITN Malang library. Jurnal Mahasiswa Teknik Informatika, 5(1).

Gustientiedina, G., Adiya, M. H., & Desnelita, Y. (2019). Application of K-Means algorithm for drug data clustering. National Journal of Information Technology and Systems, 5(1), 17–24. https://doi.org/10.25077/teknosi.v5i1.2019.17-24

Harahap, F. (2021). Comparison of K-Means and K-Medoids algorithms for clustering classes of Tunagrahita students. Applied Informatics Nusantara, 2(4).

Hediyati, D., & Suartana, I. M. (2021). Application of principal component analysis (PCA) for dimension reduction in the clustering process of agricultural production data in Bojonegoro Regency.

Hubbs, C. D., Li, C., Sahinidis, N. V., Grossmann, I. E., & Wassick, J. M. (2020). A deep reinforcement learning approach for chemical production scheduling. Computers and Chemical Engineering, 141. https://doi.org/10.1016/j.compchemeng.2020.106982

Hutagalung, J., Ginantra, N. L. W. S. R., Bhawika, G. W., Parwita, W. G. S., Wanto, A., & Panjaitan, P. D. (2021). COVID-19 cases and deaths in Southeast Asia clustering using K-Means algorithm. Journal of Physics: Conference Series, 1783(1). https://doi.org/10.1088/1742-6596/1783/1/012027

Kotsiantis, S. B. (2007). Supervised machine learning: A review of classification techniques. Informatica, 31.

Liaw, L. C. M., Tan, S. C., Goh, P. Y., & Lim, C. P. (2025). A histogram SMOTE-based sampling algorithm with incremental learning for imbalanced data classification. Information Sciences, 686, 121193.

Lubis, A., Irawan, Y., Junadhi, J., & Defit, S. (2024). Leveraging K-Nearest Neighbors with SMOTE and boosting techniques for data imbalance and accuracy improvement. Journal of Applied Data Sciences, 5(4), 1625–1638.

Mega, W. (2015). Clustering using K-Means methods to determine literary nutrition status. National Journal of Information Technology, 15(2).

Nabila, Z., Rahman Isnain, A., & Abidin, Z. (2021). Data mining analysis for clustering COVID-19 cases in Lampung Province with K-Means algorithm. Journal of Information Technology and Systems (JTSI), 2(2), 100. Retrieved from http://jim.teknokrat.ac.id/index.php/JTSI

Singh, A. (2019, October 31). Build better and accurate clusters with Gaussian mixture models. Analytics Vidhya. Retrieved from https://www.analyticsvidhya.com/blog/2019/10/gaussian-mixture-models-clustering/

Wickramasinghe, I., & Kalutarage, H. (2021). Naive Bayes: Applications, variations and vulnerabilities: A review of literature with code snippets for implementation. Soft Computing, 25(3), 2277–2293. https://doi.org/10.1007/s00500-020-05297-6

Widodo, Y. B., Anggraeini, S. A., & Sutabri, T. (2021). Design of a web-based diabetes disease diagnosis expert system using the Naive Bayes algorithm. Journal of Informatics and Computer Technology, 7(1), 112–123. https://doi.org/10.37012/jtik.v7i1.507

Downloads

Published

2025-01-30

How to Cite

Application of K-Means and Naïve Bayes Algorithms for Prediction Model of Student Interest Concentration (Case Study: Amikom University Yogyakarta). (2025). G-Tech: Jurnal Teknologi Terapan, 9(1), 511-519. https://doi.org/10.70609/gtech.v9i1.6523

Most read articles by the same author(s)