Decoupled Knowledge Distillation
Развязанная передача знаний (Decoupled Knowledge Distillation)
2022-06-01
SCID: 54.1/vu8bsajp
Discuss with AI
CIFAR-100Decoupled Knowledge DistillationImageNetMS-COCOimage classificationlogit distillationnon-target class knowledge distillation (NCKD)object detectiontarget class knowledge distillation (TCKD)
Figures from the paper
Abstract (AI)
State-of-the-art distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. To provide a novel viewpoint to study logit distillation, we re-formulate the classical KD loss into two parts, i.e., target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). We empirically investigate and prove the effects of the two parts: TCKD transfers knowledge concerning the “difficulty” of training samples, while NCKD is the prominent reason why logit distillation works. More importantly, we reveal that the classical KD loss is a coupled formulation, which (1) suppresses the effectiveness of NCKD and (2) limits the flexibility to balance these two parts. To address these issues, we present Decoupled Knowledge Distillation (DKD), enabling TCKD and NCKD to play their roles more efficiently and flexibly. Compared with complex feature-based methods, our DKD achieves comparable or even better results and has better training efficiency on CIFAR-100, ImageNet, and MS-COCO datasets for image classification and object detection tasks. This paper proves the great potential of logit distillation, and we hope it will be helpful for future research. The code is available at https://github.com/megviiresearch/mdistiller.
Key Findings
1
Classical knowledge distillation (KD) loss can be re-formulated into two parts: target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD).
2
DKD achieves comparable or better results than complex feature-based methods and offers better training efficiency on CIFAR-100, ImageNet, and MS-COCO for classification and object detection.
3
Decoupled Knowledge Distillation (DKD) separates TCKD and NCKD, enabling each to operate more efficiently and flexibly.
4
TCKD transfers knowledge about the "difficulty" of training samples, while NCKD is the main reason logit distillation is effective.
5
The classical KD loss is a coupled formulation that suppresses NCKD effectiveness and limits flexibility in balancing TCKD and NCKD.
Research Object
Logit-based knowledge distillation process between teacher and student neural networks
Research Subject
Decoupling classical KD loss into target-class knowledge distillation (TCKD) and non-target-class knowledge distillation (NCKD), analyzing their distinct roles and proposing Decoupled Knowledge Distillation (DKD) to improve effectiveness and flexibility of logit distillation
Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
956
Access Type
Author Information
Download PDF
Subscribe to digest