Emerging Properties in Self-Supervised Vision Transformers

Появляющиеся свойства в самообучающихся Vision Transformer
Hugo Touvron, Hervé Jeǵou, Armand Joulin, Ishan Misra, Julien Mairal, Mathilde Caron, Piotr Bojanowski
2021-10-01

DINO (self-distillation without labels)k-NN classification on ImageNetmomentum encoderself-supervised Vision Transformersemantic segmentation emergence
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [16] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [26], multi-crop training [9], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
1
DINO, a simple self-supervised method interpreted as label-free self-distillation, synergizes with ViTs to reach 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
2
Momentum encoder, multi-crop training, and using small patches are important components for strong self-supervised ViT performance.
3
Self-supervised ViT features are strong k-NN classifiers, achieving 78.3% top-1 on ImageNet with a small ViT.
4
Self-supervised ViT features explicitly encode semantic segmentation information, more clearly than supervised ViTs or convnets.

Self-supervised Vision Transformer (ViT) models and their learned feature representations

Emergent properties of the learned features including explicit semantic segmentation information, k-NN classification performance, and factors enabling these properties (momentum encoder, multi-crop training, small patches) demonstrated via the DINO self-supervised method

Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
5537
Access Type
Author Information
Authors
Hugo Touvron
Hervé Jeǵou
Armand Joulin
Ishan Misra
Julien Mairal
Mathilde Caron
Piotr Bojanowski
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%