Swin Transformer V2: Scaling Up Capacity and Resolution
Swin Transformer V2: увеличение ёмкости и разрешения
2022-06-01
SCID: 54.1/96hbndrf
Discuss with AI
Swin Transformer V2high-resolution training (1536x1536)log-spaced continuous position biasmemory-efficient implementationresidual post normalizationscaled cosine attentionscaling up capacityself-supervised pre-trainingstate-of-the-art results on ImageNet-V2 COCO ADE20K Kinetics-400transfer from low to high resolution
Figures from the paper
Abstract (AI)
We present techniques for scaling Swin Transformer [35] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on ImageNet- V2 image classification, 63.1 / 54.4 box / mask mAP on COCO object detection, 59.9 mIoU on ADE20K semantic segmentation, and 86.8% top-1 accuracy on Kinetics-400 video action classification. We tackle issues of training instability, and study how to effectively transfer models pre-trained at low resolutions to higher resolution ones. To this aim, several novel technologies are proposed: 1) a residual post normalization technique and a scaled cosine attention approach to improve the stability of large vision models; 2) a log-spaced continuous position bias technique to effectively transfer models pre-trained at low-resolution images and windows to their higher-resolution counterparts. In addition, we share our crucial implementation details that lead to significant savings of GPU memory consumption and thus make it feasi-ble to train large vision models with regular GPUs. Using these techniques and self-supervised pre-training, we suc-cessfully train a strong 3 billion Swin Transformer model and effectively transfer it to various vision tasks involving high-resolution images or windows, achieving the state-of-the-art accuracy on a variety of benchmarks. Code is avail-able at https://github.com/microsoft/Swin-Transformer.
Key Findings
1
Demonstrated that self-supervised pre-training plus proposed techniques enables effective transfer of the 3B Swin Transformer to high-resolution vision tasks achieving state-of-the-art accuracy.
2
Introduced techniques to scale Swin Transformer up to 3 billion parameters and train with images up to 1536x1536 resolution.
3
Proposed log-spaced continuous position bias to effectively transfer models pre-trained at low resolution and window sizes to higher resolutions.
4
Proposed residual post normalization and scaled cosine attention to improve training stability of large vision models.
5
Provided implementation details that substantially reduce GPU memory consumption, enabling training of large vision models on regular GPUs.
6
Set new records on benchmarks: 84.0% top-1 on ImageNet-V2, 63.1/54.4 box/mask mAP on COCO, 59.9 mIoU on ADE20K, 86.8% top-1 on Kinetics-400.
Research Object
Swin Transformer V2 large-scale vision transformer models (up to 3 billion parameters) trained for high-resolution image and video tasks
Research Subject
Techniques and behaviors for scaling capacity and resolution including training stability (residual post normalization, scaled cosine attention), resolution-transfer via log-spaced continuous position bias, GPU memory–efficient implementations, and their impact on performance across vision benchmarks
Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
2389
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai8
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Video Swin Transformer2022
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019