An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Изображение стоит 16×16 слов: трансформеры для распознавания изображений в масштабе
2020-10-22
SCID: 54.1/n8ujjgcs
Discuss with AI
ImageNetVision Transformerimage classificationimage patchespre-training on large datasets
Figures from the paper
Abstract (AI)
While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.
Key Findings
1
A pure Transformer applied directly to sequences of image patches (Vision Transformer, ViT) can perform very well on image classification tasks.
2
ViT does not require convolutional neural networks; reliance on CNNs is unnecessary for strong vision performance.
3
ViT requires substantially fewer computational resources to train than state-of-the-art convolutional networks when achieving competitive performance.
4
When pre-trained on large datasets and transferred to mid-sized or small benchmarks (ImageNet, CIFAR-100, VTAB), ViT attains excellent results compared to state-of-the-art convolutional networks.
Research Object
Vision Transformer (a pure Transformer model applied to sequences of image patches for image classification)
Research Subject
Performance and effectiveness of a pure Transformer applied to image patch sequences for image recognition when pre-trained at scale and transferred to mid-sized and small benchmarks (e.g., ImageNet, CIFAR-100, VTAB), including training resource efficiency compared to state-of-the-art convolutional networks
Publication Details
Publication Date
2020-10-22
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai7
Learning Multiple Layers of Features from Tiny Images2024
Learning Deep Transformer Models for Machine Translation2019
Non-local Neural Networks2018
ImageNet classification with deep convolutional neural networks2017
Deep Residual Learning for Image Recognition2016
Adam: A Method for Stochastic Optimization2015
ImageNet: A large-scale hierarchical image database2009
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Emerging Properties in Self-Supervised Vision Transformers2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
SwinIR: Image Restoration Using Swin Transformer2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
UNETR: Transformers for 3D Medical Image Segmentation2022
Point Transformer2021
Swin Transformer V2: Scaling Up Capacity and Resolution2022
PVT v2: Improved baselines with pyramid vision transformer2022
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks2023
Uformer: A General U-Shaped Transformer for Image Restoration2022
TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios2021
Pre-Trained Image Processing Transformer2021
Video Swin Transformer2022
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
LoFTR: Detector-Free Local Feature Matching with Transformers2021
An Empirical Study of Training Self-Supervised Vision Transformers2021