PVT v2: Improved baselines with pyramid vision transformer

PVT v2: Улучшённые базовые модели с Pyramid Vision Transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lü, Ping Luo, Ling Shao
2022-03-16

PVT v2Pyramid Vision Transformerconvolutional feed-forward networklinear complexity attentionoverlapping patch embedding
Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT .
1
PVT v2 achieves comparable or better performance than recent transformers such as the Swin transformer.
2
PVT v2 introduces three design changes: a linear complexity attention layer, overlapping patch embedding, and a convolutional feed-forward network.
3
PVT v2 provides significant performance improvements on classification, detection, and segmentation tasks.
4
The authors provide code for PVT v2 to facilitate further transformer research in computer vision.
5
The linear complexity attention reduces PVT v1's computational complexity to linearity.

Pyramid Vision Transformer (PVT) architecture (PVT v2)

Architectural improvements and their impact on computational complexity and performance for vision tasks—specifically linear-complexity attention, overlapping patch embedding, and convolutional feed-forward network reducing complexity to linearity and improving classification, detection, and segmentation

Publication Details
Publication Date
2022-03-16
Journal
Publisher
ISSN
Cited by
2368
Access Type
Author Information
Authors
Wenhai Wang
Enze Xie
Xiang Li
Deng-Ping Fan
Kaitao Song
Ding Liang
Tong Lü
Ping Luo
Ling Shao
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%