Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Swin Transformer: Иерархический визуальный трансформер с использова́нием сдвинутых окон
2021-10-01
SCID: 54.1/7fdha76b
Discuss with AI
Hierarchical TransformerShifted windowsSwin Transformerimage classification (ImageNet-1K)object detection and instance segmentation (COCO)
Figures from the paper
Abstract (AI)
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at https://github.com/microsoft/Swin-Transformer.
Key Findings
1
For COCO object detection and instance segmentation, Swin attains 58.7 box AP and 51.1 mask AP on test-dev.
2
For semantic segmentation on ADE20K val, Swin achieves 53.5 mIoU.
3
Shifted windowing limits self-attention to non-overlapping local windows for efficiency while enabling cross-window connections.
4
Swin Transformer achieves 87.3% top-1 accuracy on ImageNet-1K image classification.
5
Swin Transformer is a hierarchical vision Transformer backbone using Shifted windows to compute representations.
6
Swin outperforms previous state-of-the-art by +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K.
7
The architecture is hierarchical, models multiple scales, and has linear computational complexity with respect to image size.
8
The hierarchical design and shifted window approach also benefit all-MLP architectures.
Research Object
Swin Transformer (hierarchical vision Transformer backbone using shifted windows)
Research Subject
Design and evaluation of a hierarchical shifted-window self-attention architecture for vision backbones, focusing on efficiency, multi-scale representation, linear complexity w.r.t. image size, and performance on image classification, object detection, and semantic segmentation
Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
32720
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai16
ImageNet classification with deep convolutional neural networks2017
Adam: A Method for Stochastic Optimization2014
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
ImageNet: A large-scale hierarchical image database2009
Gradient-based learning applied to document recognition1998
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Aggregated Residual Transformations for Deep Neural Networks2017
Non-local Neural Networks2018
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
mixup: Beyond Empirical Risk Minimization2017
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Transformer in Transformer2021
Deep Residual Learning for Image Recognition2016
Cited by20
SwinIR: Image Restoration Using Swin Transformer2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
UNETR: Transformers for 3D Medical Image Segmentation2022
Are Transformers Effective for Time Series Forecasting?2023
Attention mechanisms in computer vision: A survey2022
Swin Transformer V2: Scaling Up Capacity and Resolution2022
PVT v2: Improved baselines with pyramid vision transformer2022
Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks2023
Uformer: A General U-Shaped Transformer for Image Restoration2022
TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios2021
Video Swin Transformer2022
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
Vision Transformers for Single Image Dehazing2023
Efficient Transformers: A Survey2022
Multimodal Learning With Transformers: A Survey2023
Vision Transformer with Deformable Attention2022
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
CMT: Convolutional Neural Networks Meet Vision Transformers2022
Transformers in medical image analysis2022
Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation2023