Return to Issue Details Compressing Large Language Models (LLMs) using Knowledge Distillation for Optimizing Inference Time and Model Size Download Download PDF