Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Pyramid Vision Transformer: универсальный бэкоун для плотного предсказания без свёрток
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lü, Ping Luo, Ling Shao
2021-10-01

PVTPyramid Vision TransformerVision Transformerconvolution-free backbonedense prediction
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research.
1
PVT combines advantages of CNNs and Transformers and can serve as a direct replacement for CNN backbones across vision tasks.
2
PVT employs a progressive shrinking pyramid to reduce computation and memory for large feature maps.
3
PVT improves downstream task performance: PVT+RetinaNet achieves 40.4 AP on COCO versus ResNet50+RetinaNet at 36.3 AP, a 4.1 AP absolute gain.
4
PVT uses training on dense image partitions to produce high-resolution outputs important for dense prediction.
5
Pyramid Vision Transformer (PVT) is a convolution-free backbone designed specifically to support dense prediction tasks.

Pyramid Vision Transformer (PVT) backbone network for dense prediction tasks

Design and evaluation of a convolution-free, pyramid Transformer backbone that produces high-resolution outputs with reduced computation and memory, and its effectiveness as a unified replacement for CNN backbones on dense prediction tasks (object detection, instance and semantic segmentation)

Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
4919
Access Type
Author Information
Authors
Wenhai Wang
Enze Xie
Xiang Li
Deng-Ping Fan
Kaitao Song
Ding Liang
Tong Lü
Ping Luo
Ling Shao
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%