Emerging Properties in Self-Supervised Vision Transformers
Появляющиеся свойства в самообучающихся Vision Transformer
2021-10-01
SCID: 54.1/e5xddv6x
Discuss with AI
DINO (self-distillation without labels)k-NN classification on ImageNetmomentum encoderself-supervised Vision Transformersemantic segmentation emergence
Figures from the paper
Abstract (AI)
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [16] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [26], multi-crop training [9], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
Key Findings
1
DINO, a simple self-supervised method interpreted as label-free self-distillation, synergizes with ViTs to reach 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
2
Momentum encoder, multi-crop training, and using small patches are important components for strong self-supervised ViT performance.
3
Self-supervised ViT features are strong k-NN classifiers, achieving 78.3% top-1 on ImageNet with a small ViT.
4
Self-supervised ViT features explicitly encode semantic segmentation information, more clearly than supervised ViTs or convnets.
Research Object
Self-supervised Vision Transformer (ViT) models and their learned feature representations
Research Subject
Emergent properties of the learned features including explicit semantic segmentation information, k-NN classification performance, and factors enabling these properties (momentum encoder, multi-crop training, small patches) demonstrated via the DINO self-supervised method
Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
5537
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai7
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Neural Machine Translation by Jointly Learning to Align and Translate2014
Distilling the Knowledge in a Neural Network2015
YFCC100M2016
Deep Residual Learning for Image Recognition2016
Cited by10
UNETR: Transformers for 3D Medical Image Segmentation2022
Attention mechanisms in computer vision: A survey2022
An Empirical Study of Training Self-Supervised Vision Transformers2021
Multimodal Learning With Transformers: A Survey2023
Scaling Vision Transformers2022
Self-Supervised Speech Representation Learning: A Review2022
SSAST: Self-Supervised Audio Spectrogram Transformer2022
Multimodal Foundation Models: From Specialists to General-Purpose Assistants2024
Materials science in the era of large language models: a perspective2024
smDeepFLUOR: single-molecule deep learning fluorescence classification2026