Sequence-Level Knowledge Distillation

Дистилляция знаний на уровне последовательностей
Alexander M. Rush, Yoon Kim
2016-01-01

BLEUbeam search eliminationknowledge distillationneural machine translationsequence-level knowledge distillation
Neural machine translation (NMT) offers a novel alternative formulation of translation that is potentially simpler than statistical approaches.However to reach competitive performance, NMT models need to be exceedingly large.In this paper we consider applying knowledge distillation approaches (Bucila et al., 2006;Hinton et al., 2015) that have proven successful for reducing the size of neural models in other domains to the problem of NMT.We demonstrate that standard knowledge distillation applied to word-level prediction can be effective for NMT, and also introduce two novel sequence-level versions of knowledge distillation that further improve performance, and somewhat surprisingly, seem to eliminate the need for beam search (even when applied on the original teacher model).Our best student model runs 10 times faster than its state-of-the-art teacher with little loss in performance.It is also significantly better than a baseline model trained without knowledge distillation: by 4.2/1.7 BLEU with greedy decoding/beam search.Applying weight pruning on top of knowledge distillation results in a student model that has 13× fewer parameters than the original teacher model, with a decrease of 0.4 BLEU.
1
Applying weight pruning on top of knowledge distillation yields a student with 13 times fewer parameters than the teacher and only a 0.4 BLEU decrease.
2
Sequence-level distillation can eliminate the need for beam search, even when applied to the original teacher model.
3
Standard word-level knowledge distillation is effective for neural machine translation (NMT).
4
The best student model runs 10 times faster than the state-of-the-art teacher with little loss in performance.
5
The distilled student outperforms a baseline trained without distillation by 4.2 BLEU with greedy decoding and 1.7 BLEU with beam search.
6
Two novel sequence-level knowledge distillation methods further improve NMT performance beyond word-level distillation.

Neural machine translation models (teacher and student NMT models)

Effectiveness of knowledge distillation methods—word-level and two sequence-level variants—on student model performance, decoding behavior (eliminating beam search), speed, size, and BLEU scores

Publication Details
Publication Date
2016-01-01
Journal
Publisher
ISSN
Cited by
788
Access Type
Author Information
Authors
Alexander M. Rush
Yoon Kim
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%