Tailored knowledge distillation with automated loss function learning
Sheng Ran, Tao Huang, Wuyue Yang
Knowledge Distillation (KD) is one of the most effective and widely used methods for model compression of large models. It has achieved significant success with the meticulous development of distillation losses. However, most state-of-the-art KD losses are manually crafted and task-specific, raising …
cn, au (code pays fourni par la source)