Distilling the Knowledge in a Neural Network
Дистилляция знаний в нейронной сети
2015-03-09
SCID: 54.1/sx9q4xyw
Discuss with AI
MNISTensemble distillationknowledge distillationmodel compressionspecialist models
Figures from the paper
Abstract (AI)
A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using a different compression technique. We achieve some surprising results on MNIST and we show that we can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model. We also introduce a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse. Unlike a mixture of experts, these specialist models can be trained rapidly and in parallel.
Key Findings
1
A new ensemble type combining one or more full models with many specialist models is introduced; specialists learn to distinguish fine-grained classes confused by full models.
2
Distillation produced surprising performance improvements on MNIST compared to typical single-model training.
3
Distilling an ensemble into a single model significantly improved the acoustic model of a heavily used commercial system.
4
Ensembles of models improve performance but are expensive to deploy as predictions require multiple large neural nets.
5
Knowledge from an ensemble can be compressed into a single model using a distillation/compression technique, enabling easier deployment.
6
Specialist models can be trained rapidly and in parallel, unlike traditional mixtures of experts.
Research Object
Knowledge distilled from an ensemble of neural network models into a single neural network (for tasks such as MNIST classification and acoustic modeling)
Research Subject
The efficacy of model compression/distillation: transferring ensemble predictive behavior into a single model to retain performance while reducing computational cost, including use of specialist models and evaluation on MNIST and commercial acoustic systems
Publication Details
Publication Date
2015-03-09
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai4
Cited by20
Emerging Properties in Self-Supervised Vision Transformers2021
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks2023
Text Data Augmentation for Deep Learning2021
Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence2023
A survey of uncertainty in deep neural networks2023
Efficient Transformers: A Survey2022
Decoupled Knowledge Distillation2022
Connecting the dots in trustworthy Artificial Intelligence: From AI principles, ethics, and key requirements to responsible AI systems and regulation2023
Efficient Acceleration of Deep Learning Inference on Resource-Constrained Edge Devices: A Review2022
Estimation of continuous valence and arousal levels from faces in naturalistic conditions2021
VGSG: Vision-Guided Semantic-Group Network for Text-Based Person Search2023
ATVITSC: A Novel Encrypted Traffic Classification Method Based on Deep Learning2024
Interpretability of machine learning‐based prediction models in healthcare2020
Convergence of Edge Computing and Deep Learning: A Comprehensive Survey2020
TinyBERT: Distilling BERT for Natural Language Understanding2020
Deep Learning for Generic Object Detection: A Survey2019
Correlation Congruence for Knowledge Distillation2019
Machine Learning Interpretability: A Survey on Methods and Metrics2019
Relational Knowledge Distillation2019