Decoupled Knowledge Distillation

Развязанная передача знаний (Decoupled Knowledge Distillation)
Quan Cui, Borui Zhao, Renjie Song, Yiyu Qiu, Jiajun Liang
2022-06-01

CIFAR-100Decoupled Knowledge DistillationImageNetMS-COCOimage classificationlogit distillationnon-target class knowledge distillation (NCKD)object detectiontarget class knowledge distillation (TCKD)
State-of-the-art distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. To provide a novel viewpoint to study logit distillation, we re-formulate the classical KD loss into two parts, i.e., target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). We empirically investigate and prove the effects of the two parts: TCKD transfers knowledge concerning the “difficulty” of training samples, while NCKD is the prominent reason why logit distillation works. More importantly, we reveal that the classical KD loss is a coupled formulation, which (1) suppresses the effectiveness of NCKD and (2) limits the flexibility to balance these two parts. To address these issues, we present Decoupled Knowledge Distillation (DKD), enabling TCKD and NCKD to play their roles more efficiently and flexibly. Compared with complex feature-based methods, our DKD achieves comparable or even better results and has better training efficiency on CIFAR-100, ImageNet, and MS-COCO datasets for image classification and object detection tasks. This paper proves the great potential of logit distillation, and we hope it will be helpful for future research. The code is available at https://github.com/megviiresearch/mdistiller.
1
Classical knowledge distillation (KD) loss can be re-formulated into two parts: target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD).
2
DKD achieves comparable or better results than complex feature-based methods and offers better training efficiency on CIFAR-100, ImageNet, and MS-COCO for classification and object detection.
3
Decoupled Knowledge Distillation (DKD) separates TCKD and NCKD, enabling each to operate more efficiently and flexibly.
4
TCKD transfers knowledge about the "difficulty" of training samples, while NCKD is the main reason logit distillation is effective.
5
The classical KD loss is a coupled formulation that suppresses NCKD effectiveness and limits flexibility in balancing TCKD and NCKD.

Logit-based knowledge distillation process between teacher and student neural networks

Decoupling classical KD loss into target-class knowledge distillation (TCKD) and non-target-class knowledge distillation (NCKD), analyzing their distinct roles and proposing Decoupled Knowledge Distillation (DKD) to improve effectiveness and flexibility of logit distillation

Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
956
Access Type
Author Information
Authors
Quan Cui
Borui Zhao
Renjie Song
Yiyu Qiu
Jiajun Liang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%