Hoax News Detection with Website Comparison Using Deep Learning Approach BERT Algorithm
DOI:
https://doi.org/10.33379/gtech.v8i3.4541Keywords:
Hoax News, Text Mining, Deep Learning, BERT.Abstract
Hoax news is false and misleading information that can cause provocation and hatred for readers. With easy internet access, the spread of hoax news is getting more massive. Therefore, there needs to be a method that can detect hoax news. The research uses deep learning methods by integrating text mining to find information and news patterns related to hoaxes. By using a dataset from the kaggle site totaling around 2700 then text preprocessing is carried out so that the data is more structured for further processing. Then make feature engineering from BERT so that the data can be processed by machine learning with three classification methods namely BERT, SVM and random forest then testing and evaluation. In this study, the model that produces the highest performance is BERT with (accuracy = 0.99, ROC-AUC = 0.99) compared to traditional machine learning models.
References
Andrian, S., & Nur Asyikin, N. (2023). Literasi Digital Dalam Tindak Pidana Penyebaran Berita Bohong dan Menyesatkan Berdasarkan Undang-Undang Nomor 19 Tahun 2016 Tentang Informasi dan Transaksi Elektronik. Ameena Journal, 1(4), 340–350. https://ejournal.ymal.or.id/index.php/aij/article/view/38
Awalina, A., Fawaid, J., Krisnabayu, R. Y., & Yudistira, N. (2021, Juli 14). Indonesia’s Fake News Detection using Transformer Network. SIET ’21: 6th International Conference on Sustainable Information Engineering and Technology 2021. https://doi.org/10.1145/3479645.3479666
Daru Kusuma, P. (2020). Machine Learning Teori, Program, Dan Studi Kasus. Sleman : Deepublish.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019, Oktober 10). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. North American Chapter of the Association for Computational Linguistics.
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020, November 1). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. COLING 2020 - The 28th International Conference on Computational Linguistics.
Kurniawan, A. A., & Mustikasari, M. (2020). Implementasi Deep Learning Menggunakan Metode CNN dan LSTM untuk Menentukan Berita Palsu dalam Bahasa Indonesia. Jurnal Informatika Universitas Pamulang, 5(4), 2622–4615. https://doi.org/10.32493/informatika.v5i4.7760
Lin, W., Wu, Z., Lin, L., Wen, A., & Li, J. (2017). An ensemble random forest algorithm for insurance big data analysis. IEEE Access, 5, 16568–16575. https://doi.org/10.1109/ACCESS.2017.2738069
Mahesh, B. (2018). Machine Learning Algorithms-A Review. International Journal of Science and Research. https://doi.org/10.21275/ART20203995
Mohan, V. (2015). Preprocessing Techniques for Text Mining - An Overview. International Journal of Computer Science & Communication Networks, 5(1), 7–16. https://www.researchgate.net/publication/339529230_Preprocessing_Techniques_for_Text_Mining_-_An_Overview
Munawar, & Riadi Silitonga, Y. (2019). Sistem Pendeteksi Berita Hoax di Media Sosial dengan Teknik Data Mining Scikit Learn. Dalam Jurnal Ilmu Komputer (Vol. 4).
Munfarida, B. (2024, Maret 1). Dewan Pers: Baru 1.700 Media yang Sudah Terverifikasi . SindoNews. https://nasional.sindonews.com/read/1331869/15/dewan-pers-baru-1700-media-yang-sudah-terverifikasi-1709283799
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Asep Ripa'i, Firman Santoso, Farihin Lazim

This work is licensed under a Creative Commons Attribution 4.0 International License.









