TinyBERT: Distilling BERT for Natural Language Understanding

TinyBERT: дистилляция BERT для задач понимания естественного языка
Xiao Dong Chen, Xin Jiang, Qun Liu, Lifeng Shang, Xiaoqi Jiao, Yichun Yin, Linlin Li, Fang Wang, Qun Liu
2020-01-01

TinyBERTTransformer distillationknowledge distillationpretraining-stage distillationtask-specific distillation
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks.However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices.To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models.By leveraging this new KD method, the plenty of knowledge encoded in a large "teacher" BERT can be effectively transferred to a small "student" Tiny-BERT.Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages.This framework ensures that TinyBERT can capture the general-domain as well as the task-specific knowledge in BERT.TinyBERT 4 1 with 4 layers is empirically effective and achieves more than 96.8% the performance of its teacher BERT BASE on GLUE benchmark, while being 7.5x smaller and 9.4x faster on inference.TinyBERT 4 is also significantly better than 4-layer state-of-the-art baselines on BERT distillation, with only ∼28% parameters and ∼31% inference time of them.Moreover, TinyBERT 6 with 6 layers performs on-par with its teacher BERT BASE .
1
A novel Transformer distillation method tailored for knowledge distillation of Transformer-based models is proposed.
2
A two-stage learning framework is introduced that applies Transformer distillation during both pretraining and task-specific learning.
3
The proposed distillation method effectively transfers knowledge from a large teacher BERT to a smaller student TinyBERT.
4
The two-stage framework enables TinyBERT to capture both general-domain and task-specific knowledge from BERT.
5
TinyBERT achieves reduced model size and accelerated inference while maintaining accuracy on natural language understanding tasks.

Tiny-BERT (a distilled, smaller Transformer-based student model derived from BERT)

Knowledge distillation of Transformer-based models to accelerate inference and reduce model size while preserving general-domain and task-specific performance (Transformer distillation applied at pretraining and task-specific stages)

Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Cited by
1716
Access Type
Author Information
Authors
Xiao Dong Chen
Xin Jiang
Qun Liu
Lifeng Shang
Xiaoqi Jiao
Yichun Yin
Linlin Li
Fang Wang
Qun Liu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%