Paradigma Epistemologis Kompresi Data Teks: Huffman, Arithmetic, dan Neural Language Model

Authors

  • Luqman Affandi Universitas Negeri Malang, Indonesia
  • Didik Dwi Prasetya Universitas Negeri Malang, Indonesia
  • Syaad Patmanthara Universitas Negeri Malang, Indonesia

DOI:

https://doi.org/10.70609/jusifor.v4i2.8384

Keywords:

Kompresi teks, Epistemologi, Huffman coding, Arithmetic coding, Neural language model

Abstract

This study explores text data compression as an epistemological paradigm through a comparative analysis of three fundamental approaches: traditional methods (Huffman Coding + LZW), bit-based methods (Arithmetic Coding), and machine learning approaches (Neural Language Models). Using the Project Gutenberg dataset comprising 15,000 classical literary works with a total size of 8.5 GB and 2.1-billion-word tokens, the evaluation is conducted based on compression ratio, execution time, and memory usage. The results reveal fundamental trade-offs among the paradigms. Traditional methods achieve the fastest execution (8.3 seconds/GB, 482 MB/s, 52 MB) with a compression ratio of 3.2:1. Arithmetic coding attains near-optimal performance (99.5% of the Shannon bound) with a compression ratio of 3.8:1. Neural language models yield the highest compression ratio of 4.6:1 but require substantially higher execution time and memory. The epistemological analysis highlights distinct conceptions of information—mechanistic, mathematically optimal, and semantic-aware—and provides a conceptual framework for developing adaptive compression systems.

References

[1] Abdulmonim, D. A., & Muhamad, Z. H. (2024). Improvement of Lossless Text Compression Methods using a Hybrid Method by the Integration of RLE, LZW and Huffman Coding Algorithms. International Journal of Software Engineering and Applications, 15(5), 2.

[2] Ahuja, N. A., Datta, P., Kanzariya, B., Somayazulu, V., & Tickoo, O. (2023). Neural rate estimator and unsupervised learning for efficient distributed image analytics in Split-DNN models. In 2023 IEEE CVPR (pp. 00201).

[3] Astsatryan, H., Lalayan, A., Kocharyan, A., & Hagimont, D. (2021). Performance-efficient recommendation and prediction service for Big Data frameworks focusing on data compression and in-memory data storage indicators. Scientific Programming, 22(4), 1945.

[4] Baidoo, P. K. (2023). Comparative analysis of the compression of text data using Huffman, arithmetic, run-length, and Lempel Ziv Welch coding algorithms. Journal of Advances in Mathematics and Computer Science, 38(9), 1812.

[5] Eldstål-Ahrens, A., Arelakis, A., & Sourdis, I. (2022). L2C: Combining lossy and lossless compression on memory and I/O. ACM Transactions on Architecture and Code Optimization, 19(1), 3481641.

[6] Leiderman, T., & Ben-Ezra, Y. (2024). Information Bottleneck driven deep video compression. Entropy, 26(10), 836.

[7] Liu, J., & Gulisano, V. (2025). On-demand memory compression of stream aggregates through reinforcement learning. In 2025 ACM SIGMOD (pp. 3676151).

[8] Mafmudin, M., & Harnaningrum, L. N. (2025). Improving memory efficiency on Android: Leveraging data structures for optimal performance. Jurnal Ilmiah Komputer, 9(3), 2131.

[9] Nguyen-Tang, T., & Choi, J. (2017). Markov information bottleneck to improve information flow in stochastic neural networks. Entropy, 21(10), 976.

[10] Rahman, M., & Hamada, M. (2020). Burrows-Wheeler Transform Based Lossless Text Compression Using Keys and Huffman Coding. Symmetry, 12(10), 1654.

[11] Saidutta, Y. M., Abdi, A., & Fekri, F. (2021). Analog joint source-channel coding for distributed functional compression using deep neural networks. In 2021 IEEE ISIT (pp. 9517797).

[12] Senthil, S., & Robert, L. (2011). Text compression algorithms - a comparative study. International Journal of Computer Theory and Engineering, 3(1), 1-7.

[13] Sharma, K., & Gupta, K. (2017). Lossless data compression techniques and their performance. In 2017 IEEE CCAA (pp. 8229810).

[14] Sharmiladevi, S., More, S., & Bose, H. (2025). Enhancing data compression techniques for optimization: A novel integration of Burrows-Wheeler Transform, Lempel-Ziv-Welch, run-length encoding and Huffman coding. In 2025 IEEE INCIP (pp. 11020095).

[15] Soflaei, M., Zhang, R., Guo, H., Al-Bashabsheh, A., & Mao, Y. (2023). Information bottleneck and aggregated learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 3302150.

[16] Tunnicliffe, M., & Hunter, G. (2025). The classical model of type-token systems compared with items from the Standardized Project Gutenberg Corpus. Analytics, 4(2), 16.

[17] Venkatesh, V., Poorna, N., Vinay, C. S., Reddy, K. A., Raj, S., & Anushiadevi. (2024). Modified arithmetic coding to increase the compression rate of text data. In 2024 IEEE InC460750 (pp. 10649229).

[18] Wan, L., Alpcan, T., Kuijper, M., & Viterbo, E. (2024). Lightweight conceptual dictionary learning for text classification using information compression. IEEE Transactions on Knowledge and Data Engineering, 36(7), 3421255.

[19] Wang, Z., Lin, J., Aly, M., Young, S. I., Chandrasekhar, V., & Girod, B. (2021). Rate-distortion optimized coding for efficient CNN compression. In 2021 DCC (pp. 00033).

[20] Yang, Y., Mandt, S., & Theis, L. (2022). An introduction to neural data compression. Foundations and Trends® in Machine Learning, 15(2), 1-123. doi:10.1561/0600000107

Published

2025-12-31

How to Cite

Paradigma Epistemologis Kompresi Data Teks: Huffman, Arithmetic, dan Neural Language Model. (2025). JUSIFOR (Jurnal Sistem Informasi Dan Informatika), 4(2), 299-308. https://doi.org/10.70609/jusifor.v4i2.8384