An Empirical Study of Training Self-Supervised Vision Transformers
Эмпирическое исследование обучения самоконтролируемых Vision Transformer
2021-10-01
SCID: 54.1/xngrgc2c
Discuss with AI
MoCo v3ViTVision Transformersself-supervised learningtraining recipes
Figures from the paper
Abstract (AI)
This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know baseline given the recent progress in computer vision: self-supervised learning for Vision Transformers (ViT). While the training recipes for standard convolutional networks have been highly mature and robust, the recipes for ViT are yet to be built, especially in the self-supervised scenarios where training becomes more challenging. In this work, we go back to basics and investigate the effects of several fundamental components for training self-supervised ViT. We observe that instability is a major issue that degrades accuracy, and it can be hidden by apparently good results. We reveal that these results are indeed partial failure, and they can be improved when training is made more stable. We benchmark ViT results in MoCo v3 and several other self-supervised frameworks, with ablations in various aspects. We discuss the currently positive evidence as well as challenges and open questions. We hope that this work will provide useful data points and experience for future research.
Key Findings
1
Benchmarking ViT with MoCo v3 and other self-supervised frameworks plus ablations identifies important components affecting performance and stability.
2
Instability during training is a major issue for self-supervised ViT that degrades accuracy and can mask partial training failures.
3
Stabilizing training procedures improves apparently good but partial-failure results, yielding better final performance for self-supervised ViT.
4
The study documents positive evidence, challenges, and open questions, providing empirical data points to guide future self-supervised ViT research.
5
Training self-supervised Vision Transformers (ViT) requires distinct recipes; convolutional network recipes do not directly transfer to ViT in self-supervised settings.
Research Object
Vision Transformers (ViT) models trained with self-supervised learning
Research Subject
Effects of fundamental training components and stability issues on the performance of self-supervised ViT (including benchmarks in MoCo v3 and other self-supervised frameworks, and ablation studies)
Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
1542
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai9
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
ImageNet: A large-scale hierarchical image database2009
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Learning Multiple Layers of Features from Tiny Images2024
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Non-local Neural Networks2018
Emerging Properties in Self-Supervised Vision Transformers2021
Representation Learning with Contrastive Predictive Coding2018