TinyBERT: Distilling BERT for Natural Language Understanding
TinyBERT: дистилляция BERT для задач понимания естественного языка
2020-01-01
SCID: 54.1/trsj9tc5
Discuss with AI
TinyBERTTransformer distillationknowledge distillationpretraining-stage distillationtask-specific distillation
Figures from the paper
Abstract (AI)
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks.However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices.To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models.By leveraging this new KD method, the plenty of knowledge encoded in a large "teacher" BERT can be effectively transferred to a small "student" Tiny-BERT.Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages.This framework ensures that TinyBERT can capture the general-domain as well as the task-specific knowledge in BERT.TinyBERT 4 1 with 4 layers is empirically effective and achieves more than 96.8% the performance of its teacher BERT BASE on GLUE benchmark, while being 7.5x smaller and 9.4x faster on inference.TinyBERT 4 is also significantly better than 4-layer state-of-the-art baselines on BERT distillation, with only ∼28% parameters and ∼31% inference time of them.Moreover, TinyBERT 6 with 6 layers performs on-par with its teacher BERT BASE .
Key Findings
1
A novel Transformer distillation method tailored for knowledge distillation of Transformer-based models is proposed.
2
A two-stage learning framework is introduced that applies Transformer distillation during both pretraining and task-specific learning.
3
The proposed distillation method effectively transfers knowledge from a large teacher BERT to a smaller student TinyBERT.
4
The two-stage framework enables TinyBERT to capture both general-domain and task-specific knowledge from BERT.
5
TinyBERT achieves reduced model size and accelerated inference while maintaining accuracy on natural language understanding tasks.
Research Object
Tiny-BERT (a distilled, smaller Transformer-based student model derived from BERT)
Research Subject
Knowledge distillation of Transformer-based models to accelerate inference and reduce model size while preserving general-domain and task-specific performance (Transformer distillation applied at pretraining and task-specific stages)
Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Cited by
1716
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Distilling the Knowledge in a Neural Network2015
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank2013
Sequence-Level Knowledge Distillation2016