Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
Pyramid Vision Transformer: универсальный бэкоун для плотного предсказания без свёрток
2021-10-01
SCID: 54.1/8v636w28
Discuss with AI
PVTPyramid Vision TransformerVision Transformerconvolution-free backbonedense prediction
Figures from the paper
Abstract (AI)
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research.
Key Findings
1
PVT combines advantages of CNNs and Transformers and can serve as a direct replacement for CNN backbones across vision tasks.
2
PVT employs a progressive shrinking pyramid to reduce computation and memory for large feature maps.
3
PVT improves downstream task performance: PVT+RetinaNet achieves 40.4 AP on COCO versus ResNet50+RetinaNet at 36.3 AP, a 4.1 AP absolute gain.
4
PVT uses training on dense image partitions to produce high-resolution outputs important for dense prediction.
5
Pyramid Vision Transformer (PVT) is a convolution-free backbone designed specifically to support dense prediction tasks.
Research Object
Pyramid Vision Transformer (PVT) backbone network for dense prediction tasks
Research Subject
Design and evaluation of a convolution-free, pyramid Transformer backbone that produces high-resolution outputs with reduced computation and memory, and its effectiveness as a unified replacement for CNN backbones on dense prediction tasks (object detection, instance and semantic segmentation)
Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
4919
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai16
ImageNet classification with deep convolutional neural networks2017
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
ImageNet: A large-scale hierarchical image database2009
Gradient-based learning applied to document recognition1998
Rethinking the Inception Architecture for Computer Vision2016
Squeeze-and-Excitation Networks2018
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
A survey on Image Data Augmentation for Deep Learning2019
Aggregated Residual Transformations for Deep Neural Networks2017
Non-local Neural Networks2018
mixup: Beyond Empirical Risk Minimization2017
PVT v2: Improved baselines with pyramid vision transformer2022
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Transformer in Transformer2021
Deep Residual Learning for Image Recognition2016
Cited by17
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
UNETR: Transformers for 3D Medical Image Segmentation2022
Attention mechanisms in computer vision: A survey2022
Swin Transformer V2: Scaling Up Capacity and Resolution2022
PVT v2: Improved baselines with pyramid vision transformer2022
Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks2023
Uformer: A General U-Shaped Transformer for Image Restoration2022
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
Vision Transformers for Single Image Dehazing2023
Vision Transformer with Deformable Attention2022
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
CMT: Convolutional Neural Networks Meet Vision Transformers2022
Scaling Vision Transformers2022
MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection2022
RenAIssance: A Survey Into AI Text-to-Image Generation in the Era of Large Model2024