An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Изображение стоит 16×16 слов: трансформеры для распознавания изображений в масштабе
Jakob Uszkoreit, Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas Beyer, Alexey Dosovitskiy, Dirk Weissenborn, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly
2020-10-22

ImageNetVision Transformerimage classificationimage patchespre-training on large datasets
While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.
1
A pure Transformer applied directly to sequences of image patches (Vision Transformer, ViT) can perform very well on image classification tasks.
2
ViT does not require convolutional neural networks; reliance on CNNs is unnecessary for strong vision performance.
3
ViT requires substantially fewer computational resources to train than state-of-the-art convolutional networks when achieving competitive performance.
4
When pre-trained on large datasets and transferred to mid-sized or small benchmarks (ImageNet, CIFAR-100, VTAB), ViT attains excellent results compared to state-of-the-art convolutional networks.

Vision Transformer (a pure Transformer model applied to sequences of image patches for image classification)

Performance and effectiveness of a pure Transformer applied to image patch sequences for image recognition when pre-trained at scale and transferred to mid-sized and small benchmarks (e.g., ImageNet, CIFAR-100, VTAB), including training resource efficiency compared to state-of-the-art convolutional networks

Publication Details
Publication Date
2020-10-22
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Jakob Uszkoreit
Xiaohua Zhai
Alexander Kolesnikov
Neil Houlsby
Lucas Beyer
Alexey Dosovitskiy
Dirk Weissenborn
Thomas Unterthiner
Mostafa Dehghani
Matthias Minderer
Georg Heigold
Sylvain Gelly
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%