Impact of Preprocessing on Indonesian Extractive Summarization Using LexRank, TextRank, DivRank, and Cosine Similarity
DOI:
https://doi.org/10.70609/g-tech.v9i4.8306Keywords:
Natural Language Processing, Summarization, Extractive, TextRank, LexRank, DivRankAbstract
Extractive text summarization is a fundamental approach to tackle information overload, yet its quality is highly dependent on the pre-processing stage. Despite its crucial role, there is no consensus on the most optimal pre-processing scenario for the Indonesian language, which has a complex morphological structure. This study aims to fill this research gap by systematically analyzing the impact of seven pre-processing scenarios on four summarization methods: three graph-based methods (LexRank, TextRank, DivRank) and one topic-relevance method (Cosine Similarity against the title). Using a corpus of 3,000 Indonesian news articles and ROUGE evaluation metrics, the results show two key findings. First, the Cosine Similarity method significantly outperforms all graph-based methods, achieving the highest F1-Measure scores on ROUGE-1 (0.5073), ROUGE-2 (0.4018), and ROUGE-L (0.4574), which emphasizes the important role of the title in news texts. Second, a comprehensive pre-processing scenario involving Case Folding, Punctuation Removal, Tokenization, Normalization, Negation Handling, Stopword Removal and Stemming proves to be the most effective in improving the performance of all algorithms. These findings provide empirical evidence and practical recommendations that the combination of a title-relevancy approach with proper text normalization is the most effective strategy for optimizing extractive text summarization for the Indonesian language.
References
Alfin, M., Abidin, Z., & Basid, P. M. N. S. A. (2024). Peringkasan Multi Dokumen Berbahasa Indonesia Menggunakan Metode Recurrent Neural Network. Techno.Com, 23(1), 187–197. https://doi.org/10.62411/tc.v23i1.9605 DOI: https://doi.org/10.62411/tc.v23i1.9605
Ansor, B., Solichan, A., Al Amin, M. Z., Ramdani, A. P., Khaira, M., & Sari, N. C. (2024). Indonesian Document Text Summarization Based on Extractive Using Sentences Scoring and Fuzzy Logic (pp. 131–146). https://doi.org/10.2991/978-94-6463-480-8_11 DOI: https://doi.org/10.2991/978-94-6463-480-8_11
Azam, M., Khalid, S., Almutairi, S., Ali Khattak, H., Namoun, A., Ali, A., & Syed Muhammad Bilal, H. (2025). Current Trends and Advances in Extractive Text Summarization: A Comprehensive Review. IEEE Access, 13, 28150–28166. https://doi.org/10.1109/ACCESS.2025.3538886 DOI: https://doi.org/10.1109/ACCESS.2025.3538886
Chai, C. P. (2023). Comparison of text preprocessing methods. Natural Language Engineering, 29(3), 509–553. https://doi.org/10.1017/S1351324922000213 DOI: https://doi.org/10.1017/S1351324922000213
Deo, S., & Banik, D. (2022). Text Summarization using Textrank and Lexrank through Latent Semantic analysis. 2022 OITS International Conference on Information Technology (OCIT), 113–118. https://doi.org/10.1109/OCIT56763.2022.00031 DOI: https://doi.org/10.1109/OCIT56763.2022.00031
Ding, J., Chen, H., Kolapudi, S., Pobbathi, L., & Nguyen, H. (2023). Quality Evaluation of Summarization Models for Patent Documents. 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security (QRS), 250–259. https://doi.org/10.1109/QRS60937.2023.00033 DOI: https://doi.org/10.1109/QRS60937.2023.00033
Dursun, M. A., & Serttaş, S. (2024). A Multi-Metric Model for analyzing and comparing extractive text summarization approaches and algorithms on scientific papers. DÜMF Mühendislik Dergisi. https://doi.org/10.24012/dumf.1376978 DOI: https://doi.org/10.24012/dumf.1376978
Faisal, M., Halmahera, S., Cahyani, V. O., Aziz, A., Afrah, A. S., & Supriyono. (2024). Advanced Extractive Summarization of Indonesian Texts Using LSTM Models. 2024 8th International Conference on Information Technology, Information Systems and Electrical Engineering (ICITISEE), 58–63. https://doi.org/10.1109/ICITISEE63424.2024.10730507 DOI: https://doi.org/10.1109/ICITISEE63424.2024.10730507
Fauzi, D., Abidin, Z., & Fatchurrochman, F. (2024). Peringkasan Teks Multi Dokumen Berbahasa Indonesia Menggunakan Sentence Scoring dan SVM. Techno.Com, 23(1), 233–242. https://doi.org/10.62411/tc.v23i1.9648 DOI: https://doi.org/10.62411/tc.v23i1.9648
Girsang, A. S., & Amadeus, F. J. (2023). Extractive Text Summarization for Indonesian News Article Using Ant System Algorithm. Journal of Advances in Information Technology, 14(2), 295–301. https://doi.org/10.12720/jait.14.2.295-301 DOI: https://doi.org/10.12720/jait.14.2.295-301
Hao, R., Li, Y., Feng, Y., & Chen, Z. (2023). Are duplicates really harmful? An empirical study on bug report summarization techniques. Journal of Software: Evolution and Process, 35(11). https://doi.org/10.1002/smr.2424 DOI: https://doi.org/10.1002/smr.2424
Juna, M. F., & Hayaty, M. (2023). The observed preprocessing strategies for doing automatic text summarizing. Computer Science and Information Technologies, 4(2), 119–126. https://doi.org/10.11591/csit.v4i2.p119-126 DOI: https://doi.org/10.11591/csit.v4i2.pp119-126
Lin, N., Li, J., & Jiang, S. (2022). A simple but effective method for Indonesian automatic text summarisation. Connection Science, 34(1), 29–43. https://doi.org/10.1080/09540091.2021.1937942 DOI: https://doi.org/10.1080/09540091.2021.1937942
Liu, W., Sun, Y., Yu, B., Wang, H., Peng, Q., Hou, M., Guo, H., Wang, H., & Liu, C. (2024). Automatic Text Summarization Method Based on Improved TextRank Algorithm and K-Means Clustering. Knowledge-Based Systems, 287, 111447. https://doi.org/10.1016/j.knosys.2024.111447 DOI: https://doi.org/10.1016/j.knosys.2024.111447
Meidelfi, D., Yulherniwati, -, Rahmayuni, I., Hidayat, T., & Chandra, D. (2021). TF-IDF Implementation for Similarity Checker on The Final Project Title. International Journal of Advanced Science Computing and Engineering, 3(1), 40–52. https://doi.org/10.62527/ijasce.3.1.3 DOI: https://doi.org/10.30630/ijasce.3.1.3
Muharam, A. F., Gerhana, Y. A., Maylawati, D. S., Ramdhani, M. A., & Rahman, T. K. A. (2025). Enhancing Abstractive Multi-Document Summarization with Bert2Bert Model for Indonesian Language. JISKA (Jurnal Informatika Sunan Kalijaga), 10(1), 110–121. https://doi.org/10.14421/jiska.2025.10.1.110-121 DOI: https://doi.org/10.14421/jiska.2025.10.1.110-121
Nugraha, D. S. (2024). Some Notes on Indonesian Word Formation: A Study Based on the Derivational Morphology Approach. South Asian Research Journal of Humanities and Social Sciences, 6(01), 20–31. https://doi.org/10.36346/sarjhss.2024.v06i01.004 DOI: https://doi.org/10.36346/sarjhss.2024.v06i01.004
Park, K., Hong, J. S., & Kim, W. (2020). A Methodology Combining Cosine Similarity with Classifier for Text Classification. Applied Artificial Intelligence, 34(5), 396–411. https://doi.org/10.1080/08839514.2020.1723868 DOI: https://doi.org/10.1080/08839514.2020.1723868
Thakkar, A., Mungra, D., Agrawal, A., & Chaudhari, K. (2022). Improving the Performance of Sentiment Analysis Using Enhanced Preprocessing Technique and Artificial Neural Network. IEEE Transactions on Affective Computing, 13(4), 1771–1782. https://doi.org/10.1109/TAFFC.2022.3206891 DOI: https://doi.org/10.1109/TAFFC.2022.3206891
Wijaya, J., & Suganda Girsang, A. (2024). Indonesian News Extractive Summarization using Lexrank and YAKE Algorithm. Statistics, Optimization & Information Computing, 12(6), 1973–1983. https://doi.org/10.19139/soic-2310-5070-1976 DOI: https://doi.org/10.19139/soic-2310-5070-1976
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Andri Setiawan, Zainal Abidin, Mochamad Imamudin

This work is licensed under a Creative Commons Attribution 4.0 International License.








