Comparative Analysis of NLP Techniques for Hate Speech Classification in Online Communications
DOI:
https://doi.org/10.33379/gtech.v8i1.3959Keywords:
NLP, Hatefull Speech Detection, Word Embedding, TFIDF, Machine LearningAbstract
This research aimed to compare the effectiveness of two Natural Language Processing (NLP) techniques—SpaCy's word embeddings and Sklearn's TF-IDF vectorization—in identifying hate speech within online comments. Utilizing a balanced dataset, each model was meticulously assessed on its ability to classify comments as 'hateful' or 'non-hateful'. The evaluation metrics employed were precision, recall, F1-score, and overall accuracy. The model using SpaCy's word embeddings achieved an accuracy of 65%, with equal precision and recall for both classes. The Sklearn's TF-IDF vectorization model, however, demonstrated superior performance with an overall accuracy of 75% and an enhanced ability to correctly identify hateful comments, evidenced by a 77% recall rate. This suggests that the TF-IDF model is more adept at discerning nuanced expressions of hate speech. The study's findings highlight the critical role of vectorization methods in the field of automated content moderation and stress the importance of continued innovation and model adaptation to effectively manage the evolving nature of online hate speech.
References
Abardazzou, N. (2023). Unmasking implicit abuse: a data-centric approach to detect online abusive language.
Anansaringkarn, P., & Neo, R. (2021). How can state regulations over the online sphere continue to respect the freedom of expression? A case study of contemporary ‘fake news’ regulations in Thailand. Information & Communications Technology Law, 30(3), 283–303.
Borrego-D’iaz, J., & Galán-Páez, J. (2022). Explainable Artificial Intelligence in Data Science: From Foundational Issues Towards Socio-technical Considerations. Minds and Machines, 32(3), 485–531.
Çinar, N. (2020). The Rise Of Consumer Generated Content And Its Transformative Effect On Advertising. In Reimagining Communication: Mediation (pp. 193–209). Routledge.
Chen, J.-L., Dai, Y.-N., Grimaldi, N. S., Lin, J.-J., Hu, B.-Y., Wu, Y.-F., & Gao, S. (2022). Plantar Pressure-Based Insole Gait Monitoring Techniques for Diseases Monitoring and Analysis: A Review. Advanced Materials Technologies, 7(1), 2100566.
De Gregorio, G. (2020). Democratising online content moderation: A constitutional framework. Computer Law & Security Review, 36, 105374.
Duwairi, R., Hayajneh, A., & Quwaider, M. (2021). A deep learning framework for automatic detection of hate speech embedded in Arabic tweets. Arabian Journal for Science and Engineering, 46, 4001–4014.
Eusebius, S. (2020). Customer-based brand equity in a digital age: An analysis of brand associations in user-generated social media content. University of Otago.
Hassija, V., Chamola, V., Mahapatra, A., Singal, A., Goel, D., Huang, K., … Hussain, A. (2023). Interpreting black-box models: a review on explainable artificial intelligence. Cognitive Computation, 1–30.
Heath, R. (2020). Branding in a Digitally Empowered World: The Role of User-Generated Content. Auckland University of Technology.
Homayounfar, S. Z., & Andrew, T. L. (2020). Wearable sensors for monitoring human motion: a review on mechanisms, materials, and challenges. SLAS TECHNOLOGY: Translating Life Sciences Innovation, 25(1), 9–24.
Iosifidis, P., & Nicoli, N. (2020). Digital democracy, social media and disinformation. Routledge.
Kalra, V., Kashyap, I., & Kaur, H. (2022). Improving document classification using domain-specific vocabulary: hybridization of deep learning approach with TFIDF. International Journal of Information Technology, 14(5), 2451–2457.
Kiritchenko, S., Nejadgholi, I., & Fraser, K. C. (2021). Confronting abusive language online: A survey from the ethical and human rights perspective. Journal of Artificial Intelligence Research, 71, 431–478.
Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., … others. (2023). Summary of chatgpt-related research and perspective towards the future of large language models. Meta-Radiology, 100017.
Lukings, M., & Habibi Lashkari, A. (2022). Technical Complexities. In Understanding Cybersecurity Law in Data Sovereignty and Digital Governance: An Overview from a Legal Perspective (pp. 117–180). Springer.
Machova, K., Srba, I., Sarnovsk`y, M., Paralič, J., Kresnakova, V. M., Hrckova, A., … others. (2021). Addressing False Information and Abusive Language in Digital Space Using Intelligent Approaches. Towards Digital Intelligence Society: A Knowledge-Based Approach, 3–32.
Mann, B. L. (2020). Applying Internet Laws and Regulations to Educational Technology. IGI Global.
Markov, Č., & DJordjević, A. (2024). Becoming a Target: Journalists’ Perspectives on Anti-Press Discourse and Experiences with Hate Speech. Journalism Practice, 18(2), 283–300.
McDermid, J. A., Jia, Y., Porter, Z., & Habli, I. (2021). Artificial intelligence explainability: the technical and ethical dimensions. Philosophical Transactions of the Royal Society A, 379(2207), 20200363.
Miric, M., Jia, N., & Huang, K. G. (2023). Using supervised machine learning for large-scale classification in management research: The case for identifying artificial intelligence patents. Strategic Management Journal, 44(2), 491–519.
Nguyen, T. T., Huynh, T. T., Yin, H., Weidlich, M., Nguyen, T. T., Mai, T. S., & Nguyen, Q. V. H. (2023). Detecting rumours with latency guarantees using massive streaming data. The VLDB Journal, 32(2), 369–387.
Nodehi, I., Hassannataj Joloudari, J., Sharifrazi, D., Nematollahi, M., Marefat, A., Çifçi, M. A., & Hussain, S. (n.d.). Ocsvm-Cnn: Malicious Script Detection Using One-Class Support Vector Machine Combined with Convolutional Neural Network. Danial and Nematollahi, Mohammad and Marefat, Abdolreza and Çifçi, Mehmet Akif and Hussain, Sadiq, Ocsvm-Cnn: Malicious Script Detection Using One-Class Support Vector Machine Combined with Convolutional Neural Network.
Rodriguez, P. L., & Spirling, A. (2022). Word embeddings: What works, what doesn’t, and how to tell the difference for applied research. The Journal of Politics, 84(1), 101–115.
Shawkat, N. (2023). Evaluation of Different Machine Learning, Deep Learning and Text Processing Techniques for Hate Speech Detection.
Tworek, H. J. S. (2021). Fighting hate with speech Law: Media and German visions of democracy. The Journal of Holocaust Research, 35(2), 106–122.
Vidgen, B., & Yasseri, T. (2020). Detecting weak and strong Islamophobic hate speech on social media. Journal of Information Technology & Politics, 17(1), 66–78.
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., … others. (2023). The rise and potential of large language model based agents: A survey. ArXiv Preprint ArXiv:2309.07864.
Yin, W., & Zubiaga, A. (2021). Towards generalisable hate speech detection: a review on obstacles and solutions. PeerJ Computer Science, 7, e598.
Zhu, Y., Yuan, H., Wang, S., Liu, J., Liu, W., Deng, C., … Wen, J.-R. (2023). Large language models for information retrieval: A survey. ArXiv Preprint ArXiv:2308.07107.
Downloads
Published
Issue
Section
License
Copyright (c) 2023 Gregorius Airlangga

This work is licensed under a Creative Commons Attribution 4.0 International License.









