Compressing Large Language Models (LLMs) using Knowledge Distillation for Optimizing Inference Time and Model Size
DOI:
https://doi.org/10.70609/gtech.v9i2.6749Keywords:
Knowledge Distillation, Model Compression, Large Language ModelsAbstract
Large Language Models (LLMs) contain a vast number of parameters and are significantly large in size. For instance, the DeepSeek-V3 model consists of approximately 671 billion parameters and has a file size of up to 720GB. The sheer number of parameters in LLMs reflects their high complexity, which can serve as both an advantage and a drawback, particularly when deployed in environments with limited computational resources. This study focuses on compressing a custom-built lightweight model using knowledge distillation techniques applied to LLMs. The results indicate that the model’s parameters can be reduced by up to 94.18%, its file size by up to 71.00%, and its inference time by up to 1.13%. Notably, despite these reductions, the model remains capable of performing specialized tasks with satisfactory accuracy. This finding underscores the potential of knowledge distillation as an effective method for reducing model size while maintaining operational efficiency, particularly in scenarios where computational constraints lead to mismatched capabilities. Efficiency in knowledge distillation is achieved through a combination of model size reduction and the alignment of computational capacity with task-specific requirements.
References
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., ... & McGrew, B. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., ... & Zhu, T. (2023). Qwen technical report. arXiv preprint arXiv:2309.16609.
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., ... & Zaremba, W. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
Chen, W., Peng, L., Huang, Y., Jing, M., & Zeng, X. (2021). Knowledge Distillation for U-Net Based Image Denoising. 2021 IEEE 14th International Conference on ASIC (ASICON), 1–4.
Cho, J., & Hariharan, B. (2019). On the Efficacy of Knowledge Distillation. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 4793–4801.
Cui, J., Tian, Z., Zhong, Z., Qi, X., Yu, B., & Zhang, H. (2023). Decoupled Kullback-Leibler Divergence Loss. ArXiv, abs/2305.13948.
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186).
Gou, J., Yu, B., Maybank, S. J., & Tao, D. (2021). Knowledge distillation: A survey. International Journal of Computer Vision, 129(6), 1789–1819.
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., ... & He, Y. (2025). Deepseek-r1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948.
Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
L. Wang and K.-J. Yoon. (2022). Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6), 3048–3068.
Naveed, H., Khan, A., Qiu, S., Saqib, M., Anwar, S., Usman, M., Barnes, N., & Mian, A. (2023). A Comprehensive Overview of Large Language Models. ArXiv, abs/2307.06435.
Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A., & Wilson, A. (2021). Does Knowledge Distillation Really Work?. ArXiv, abs/2106.05945.
Tang, J., Shivanna, R., Zhao, Z., Lin, D., Singh, A., Chi, E., & Jain, S. (2020). Understanding and Improving Knowledge Distillation. ArXiv, abs/2002.03532.
Tao, Z., Xia, Q., Cheng, S., & Li, Q. (2023). An Efficient and Robust Cloud-Based Deep Learning With Knowledge Distillation. IEEE Transactions on Cloud Computing, 11, 1733–1745.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., ... & Lample, G. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
Wu, L., Zheng, Z., Qiu, Z., Wang, H., Gu, H., Shen, T., Qin, C., Zhu, C., Zhu, H., Liu, Q., Xiong, H., & Chen, E. (2023). A Survey on Large Language Models for Recommendation. ArXiv, abs/2305.19860.
Yang, C., Zhu, Y., Lu, W., Wang, Y., Chen, Q., Gao, C., ... & Chen, Y. (2024). Survey on knowledge distillation for large language models: methods, evaluation, and application. ACM Transactions on Intelligent Systems and Technology.
Yu, Y., & Kim, N. (2024). Heterogeneous Knowledge Distillation Using Conceptual Learning. IEEE Access, 12, 52803–52814.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., ... & Wen, J. R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2).
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Rachmad Imam Tarecha, Priska Choirina

This work is licensed under a Creative Commons Attribution 4.0 International License.









