PVT v2: Improved baselines with pyramid vision transformer
PVT v2: Улучшённые базовые модели с Pyramid Vision Transformer
2022-03-16
SCID: 54.1/4fmjckqw
Discuss with AI
PVT v2Pyramid Vision Transformerconvolutional feed-forward networklinear complexity attentionoverlapping patch embedding
Figures from the paper
Abstract (AI)
Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT .
Key Findings
1
PVT v2 achieves comparable or better performance than recent transformers such as the Swin transformer.
2
PVT v2 introduces three design changes: a linear complexity attention layer, overlapping patch embedding, and a convolutional feed-forward network.
3
PVT v2 provides significant performance improvements on classification, detection, and segmentation tasks.
4
The authors provide code for PVT v2 to facilitate further transformer research in computer vision.
5
The linear complexity attention reduces PVT v1's computational complexity to linearity.
Research Object
Pyramid Vision Transformer (PVT) architecture (PVT v2)
Research Subject
Architectural improvements and their impact on computational complexity and performance for vision tasks—specifically linear-complexity attention, overlapping patch embedding, and convolutional feed-forward network reducing complexity to linearity and improving classification, detection, and segmentation
Publication Details
Publication Date
2022-03-16
Journal
Publisher
ISSN
Cited by
2368
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai13
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
Transformer in Transformer2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
mixup: Beyond Empirical Risk Minimization2017
Aggregated Residual Transformations for Deep Neural Networks2017
Deep Residual Learning for Image Recognition2016
Rethinking the Inception Architecture for Computer Vision2016
ImageNet: A large-scale hierarchical image database2009
Proceedings of the 24th international conference on Machine learning2007