An Empirical Study of Training Self-Supervised Vision Transformers

Эмпирическое исследование обучения самоконтролируемых Vision Transformer
Kaiming He, Saining Xie, Xinlei Chen
2021-10-01

MoCo v3ViTVision Transformersself-supervised learningtraining recipes
This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know baseline given the recent progress in computer vision: self-supervised learning for Vision Transformers (ViT). While the training recipes for standard convolutional networks have been highly mature and robust, the recipes for ViT are yet to be built, especially in the self-supervised scenarios where training becomes more challenging. In this work, we go back to basics and investigate the effects of several fundamental components for training self-supervised ViT. We observe that instability is a major issue that degrades accuracy, and it can be hidden by apparently good results. We reveal that these results are indeed partial failure, and they can be improved when training is made more stable. We benchmark ViT results in MoCo v3 and several other self-supervised frameworks, with ablations in various aspects. We discuss the currently positive evidence as well as challenges and open questions. We hope that this work will provide useful data points and experience for future research.
1
Benchmarking ViT with MoCo v3 and other self-supervised frameworks plus ablations identifies important components affecting performance and stability.
2
Instability during training is a major issue for self-supervised ViT that degrades accuracy and can mask partial training failures.
3
Stabilizing training procedures improves apparently good but partial-failure results, yielding better final performance for self-supervised ViT.
4
The study documents positive evidence, challenges, and open questions, providing empirical data points to guide future self-supervised ViT research.
5
Training self-supervised Vision Transformers (ViT) requires distinct recipes; convolutional network recipes do not directly transfer to ViT in self-supervised settings.

Vision Transformers (ViT) models trained with self-supervised learning

Effects of fundamental training components and stability issues on the performance of self-supervised ViT (including benchmarks in MoCo v3 and other self-supervised frameworks, and ablation studies)

Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
1542
Access Type
Author Information
Authors
Kaiming He
Saining Xie
Xinlei Chen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%