Impact of Preprocessing on Indonesian Extractive Summarization Using LexRank, TextRank, DivRank, and Cosine Similarity

Authors

  • Andri Setiawan Universitas Islam Negeri Maulana Malik Ibrahim Malang, Indonesia
  • Zainal Abidin Universitas Islam Negeri Maulana Malik Ibrahim Malang, Indonesia
  • Mochamad Imamudin Universitas Islam Negeri Maulana Malik Ibrahim Malang, Indonesia

DOI:

https://doi.org/10.70609/g-tech.v9i4.8306

Keywords:

Natural Language Processing, Summarization, Extractive, TextRank, LexRank, DivRank

Abstract

Extractive text summarization is a fundamental approach to tackle information overload, yet its quality is highly dependent on the pre-processing stage. Despite its crucial role, there is no consensus on the most optimal pre-processing scenario for the Indonesian language, which has a complex morphological structure. This study aims to fill this research gap by systematically analyzing the impact of seven pre-processing scenarios on four summarization methods: three graph-based methods (LexRank, TextRank, DivRank) and one topic-relevance method (Cosine Similarity against the title). Using a corpus of 3,000 Indonesian news articles and ROUGE evaluation metrics, the results show two key findings. First, the Cosine Similarity method significantly outperforms all graph-based methods, achieving the highest F1-Measure scores on ROUGE-1 (0.5073), ROUGE-2 (0.4018), and ROUGE-L (0.4574), which emphasizes the important role of the title in news texts. Second, a comprehensive pre-processing scenario involving Case Folding, Punctuation Removal, Tokenization, Normalization, Negation Handling, Stopword Removal and Stemming proves to be the most effective in improving the performance of all algorithms. These findings provide empirical evidence and practical recommendations that the combination of a title-relevancy approach with proper text normalization is the most effective strategy for optimizing extractive text summarization for the Indonesian language.

References

Alfin, M., Abidin, Z., & Basid, P. M. N. S. A. (2024). Peringkasan Multi Dokumen Berbahasa Indonesia Menggunakan Metode Recurrent Neural Network. Techno.Com, 23(1), 187–197. https://doi.org/10.62411/tc.v23i1.9605 DOI: https://doi.org/10.62411/tc.v23i1.9605

Ansor, B., Solichan, A., Al Amin, M. Z., Ramdani, A. P., Khaira, M., & Sari, N. C. (2024). Indonesian Document Text Summarization Based on Extractive Using Sentences Scoring and Fuzzy Logic (pp. 131–146). https://doi.org/10.2991/978-94-6463-480-8_11 DOI: https://doi.org/10.2991/978-94-6463-480-8_11

Azam, M., Khalid, S., Almutairi, S., Ali Khattak, H., Namoun, A., Ali, A., & Syed Muhammad Bilal, H. (2025). Current Trends and Advances in Extractive Text Summarization: A Comprehensive Review. IEEE Access, 13, 28150–28166. https://doi.org/10.1109/ACCESS.2025.3538886 DOI: https://doi.org/10.1109/ACCESS.2025.3538886

Chai, C. P. (2023). Comparison of text preprocessing methods. Natural Language Engineering, 29(3), 509–553. https://doi.org/10.1017/S1351324922000213 DOI: https://doi.org/10.1017/S1351324922000213

Deo, S., & Banik, D. (2022). Text Summarization using Textrank and Lexrank through Latent Semantic analysis. 2022 OITS International Conference on Information Technology (OCIT), 113–118. https://doi.org/10.1109/OCIT56763.2022.00031 DOI: https://doi.org/10.1109/OCIT56763.2022.00031

Ding, J., Chen, H., Kolapudi, S., Pobbathi, L., & Nguyen, H. (2023). Quality Evaluation of Summarization Models for Patent Documents. 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security (QRS), 250–259. https://doi.org/10.1109/QRS60937.2023.00033 DOI: https://doi.org/10.1109/QRS60937.2023.00033

Dursun, M. A., & Serttaş, S. (2024). A Multi-Metric Model for analyzing and comparing extractive text summarization approaches and algorithms on scientific papers. DÜMF Mühendislik Dergisi. https://doi.org/10.24012/dumf.1376978 DOI: https://doi.org/10.24012/dumf.1376978

Faisal, M., Halmahera, S., Cahyani, V. O., Aziz, A., Afrah, A. S., & Supriyono. (2024). Advanced Extractive Summarization of Indonesian Texts Using LSTM Models. 2024 8th International Conference on Information Technology, Information Systems and Electrical Engineering (ICITISEE), 58–63. https://doi.org/10.1109/ICITISEE63424.2024.10730507 DOI: https://doi.org/10.1109/ICITISEE63424.2024.10730507

Fauzi, D., Abidin, Z., & Fatchurrochman, F. (2024). Peringkasan Teks Multi Dokumen Berbahasa Indonesia Menggunakan Sentence Scoring dan SVM. Techno.Com, 23(1), 233–242. https://doi.org/10.62411/tc.v23i1.9648 DOI: https://doi.org/10.62411/tc.v23i1.9648

Girsang, A. S., & Amadeus, F. J. (2023). Extractive Text Summarization for Indonesian News Article Using Ant System Algorithm. Journal of Advances in Information Technology, 14(2), 295–301. https://doi.org/10.12720/jait.14.2.295-301 DOI: https://doi.org/10.12720/jait.14.2.295-301

Hao, R., Li, Y., Feng, Y., & Chen, Z. (2023). Are duplicates really harmful? An empirical study on bug report summarization techniques. Journal of Software: Evolution and Process, 35(11). https://doi.org/10.1002/smr.2424 DOI: https://doi.org/10.1002/smr.2424

Juna, M. F., & Hayaty, M. (2023). The observed preprocessing strategies for doing automatic text summarizing. Computer Science and Information Technologies, 4(2), 119–126. https://doi.org/10.11591/csit.v4i2.p119-126 DOI: https://doi.org/10.11591/csit.v4i2.pp119-126

Lin, N., Li, J., & Jiang, S. (2022). A simple but effective method for Indonesian automatic text summarisation. Connection Science, 34(1), 29–43. https://doi.org/10.1080/09540091.2021.1937942 DOI: https://doi.org/10.1080/09540091.2021.1937942

Liu, W., Sun, Y., Yu, B., Wang, H., Peng, Q., Hou, M., Guo, H., Wang, H., & Liu, C. (2024). Automatic Text Summarization Method Based on Improved TextRank Algorithm and K-Means Clustering. Knowledge-Based Systems, 287, 111447. https://doi.org/10.1016/j.knosys.2024.111447 DOI: https://doi.org/10.1016/j.knosys.2024.111447

Meidelfi, D., Yulherniwati, -, Rahmayuni, I., Hidayat, T., & Chandra, D. (2021). TF-IDF Implementation for Similarity Checker on The Final Project Title. International Journal of Advanced Science Computing and Engineering, 3(1), 40–52. https://doi.org/10.62527/ijasce.3.1.3 DOI: https://doi.org/10.30630/ijasce.3.1.3

Muharam, A. F., Gerhana, Y. A., Maylawati, D. S., Ramdhani, M. A., & Rahman, T. K. A. (2025). Enhancing Abstractive Multi-Document Summarization with Bert2Bert Model for Indonesian Language. JISKA (Jurnal Informatika Sunan Kalijaga), 10(1), 110–121. https://doi.org/10.14421/jiska.2025.10.1.110-121 DOI: https://doi.org/10.14421/jiska.2025.10.1.110-121

Nugraha, D. S. (2024). Some Notes on Indonesian Word Formation: A Study Based on the Derivational Morphology Approach. South Asian Research Journal of Humanities and Social Sciences, 6(01), 20–31. https://doi.org/10.36346/sarjhss.2024.v06i01.004 DOI: https://doi.org/10.36346/sarjhss.2024.v06i01.004

Park, K., Hong, J. S., & Kim, W. (2020). A Methodology Combining Cosine Similarity with Classifier for Text Classification. Applied Artificial Intelligence, 34(5), 396–411. https://doi.org/10.1080/08839514.2020.1723868 DOI: https://doi.org/10.1080/08839514.2020.1723868

Thakkar, A., Mungra, D., Agrawal, A., & Chaudhari, K. (2022). Improving the Performance of Sentiment Analysis Using Enhanced Preprocessing Technique and Artificial Neural Network. IEEE Transactions on Affective Computing, 13(4), 1771–1782. https://doi.org/10.1109/TAFFC.2022.3206891 DOI: https://doi.org/10.1109/TAFFC.2022.3206891

Wijaya, J., & Suganda Girsang, A. (2024). Indonesian News Extractive Summarization using Lexrank and YAKE Algorithm. Statistics, Optimization & Information Computing, 12(6), 1973–1983. https://doi.org/10.19139/soic-2310-5070-1976 DOI: https://doi.org/10.19139/soic-2310-5070-1976

Downloads

Published

2025-10-30

How to Cite

Impact of Preprocessing on Indonesian Extractive Summarization Using LexRank, TextRank, DivRank, and Cosine Similarity. (2025). G-Tech: Jurnal Teknologi Terapan, 9(4), 2311-2321. https://doi.org/10.70609/g-tech.v9i4.8306

Most read articles by the same author(s)

<< < 1 2