Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
Переосмысление семантической сегментации с точки зрения sequence-to-sequence на основе трансформеров
2021-06-01
SCID: 54.1/vxbtezdr
Discuss with AI
ADE20K mIoUSEgmentation TRansformer (SETR)semantic segmentationsequence-to-sequence predictiontransformer encoder
Figures from the paper
Abstract (AI)
Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for segmentation, the latest efforts have been focused on increasing the receptive field, through either dilated/atrous convolutions or inserting attention modules. However, the encoder-decoder based FCN architecture remains unchanged. In this paper, we aim to provide an alternative perspective by treating semantic segmentation as a sequence-to-sequence prediction task. Specifically, we deploy a pure transformer (i.e., without convolution and resolution reduction) to encode an image as a sequence of patches. With the global context modeled in every layer of the transformer, this encoder can be combined with a simple decoder to provide a powerful segmentation model, termed SEgmentation TRansformer (SETR). Extensive experiments show that SETR achieves new state of the art on ADE20K (50.28% mIoU), Pascal Context (55.83% mIoU) and competitive results on Cityscapes. Particularly, we achieve the first position in the highly competitive ADE20K test server leaderboard on the day of submission.
Key Findings
1
At time of submission, SETR ranked first on the ADE20K test server leaderboard
2
Introduces SEgmentation TRansformer (SETR), combining a transformer-based encoder with a simple decoder for segmentation
3
Proposes treating semantic segmentation as a sequence-to-sequence prediction task using a pure transformer encoder without convolutions or resolution reduction
4
SETR achieves state-of-the-art on ADE20K with 50.28% mIoU
5
SETR achieves state-of-the-art on Pascal Context with 55.83% mIoU and shows competitive results on Cityscapes
6
Transformer encoder models global context in every layer by encoding images as sequences of patches
Research Object
Semantic segmentation model using a pure transformer encoder treating an image as a sequence of patches (SETR)
Research Subject
Effectiveness of sequence-to-sequence transformer encoding (without convolutions or resolution reduction) combined with a simple decoder for modeling global context and improving segmentation performance (mIoU) on benchmarks such as ADE20K, Pascal Context, and Cityscapes
Publication Details
Publication Date
2021-06-01
Journal
Publisher
ISSN
Cited by
3673
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai7
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Non-local Neural Networks2018
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Deep Residual Learning for Image Recognition2016
Cited by13
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
SwinIR: Image Restoration Using Swin Transformer2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation2021
UNETR: Transformers for 3D Medical Image Segmentation2022
Attention mechanisms in computer vision: A survey2022
Swin Transformer V2: Scaling Up Capacity and Resolution2022
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
Remote Sensing Image Change Detection With Transformers2021
Transformer in Transformer2021
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
CMT: Convolutional Neural Networks Meet Vision Transformers2022
RenAIssance: A Survey Into AI Text-to-Image Generation in the Era of Large Model2024