Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers

Переосмысление семантической сегментации с точки зрения sequence-to-sequence на основе трансформеров
Hengshuang Zhao, Philip H. S. Torr, Jianfeng Feng, Tao Xiang, Sixiao Zheng, Jiachen Lu, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Li Zhang
2021-06-01

ADE20K mIoUSEgmentation TRansformer (SETR)semantic segmentationsequence-to-sequence predictiontransformer encoder
Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for segmentation, the latest efforts have been focused on increasing the receptive field, through either dilated/atrous convolutions or inserting attention modules. However, the encoder-decoder based FCN architecture remains unchanged. In this paper, we aim to provide an alternative perspective by treating semantic segmentation as a sequence-to-sequence prediction task. Specifically, we deploy a pure transformer (i.e., without convolution and resolution reduction) to encode an image as a sequence of patches. With the global context modeled in every layer of the transformer, this encoder can be combined with a simple decoder to provide a powerful segmentation model, termed SEgmentation TRansformer (SETR). Extensive experiments show that SETR achieves new state of the art on ADE20K (50.28% mIoU), Pascal Context (55.83% mIoU) and competitive results on Cityscapes. Particularly, we achieve the first position in the highly competitive ADE20K test server leaderboard on the day of submission.
1
At time of submission, SETR ranked first on the ADE20K test server leaderboard
2
Introduces SEgmentation TRansformer (SETR), combining a transformer-based encoder with a simple decoder for segmentation
3
Proposes treating semantic segmentation as a sequence-to-sequence prediction task using a pure transformer encoder without convolutions or resolution reduction
4
SETR achieves state-of-the-art on ADE20K with 50.28% mIoU
5
SETR achieves state-of-the-art on Pascal Context with 55.83% mIoU and shows competitive results on Cityscapes
6
Transformer encoder models global context in every layer by encoding images as sequences of patches

Semantic segmentation model using a pure transformer encoder treating an image as a sequence of patches (SETR)

Effectiveness of sequence-to-sequence transformer encoding (without convolutions or resolution reduction) combined with a simple decoder for modeling global context and improving segmentation performance (mIoU) on benchmarks such as ADE20K, Pascal Context, and Cityscapes

Publication Details
Publication Date
2021-06-01
Journal
Publisher
ISSN
Cited by
3673
Access Type
Author Information
Authors
Hengshuang Zhao
Philip H. S. Torr
Jianfeng Feng
Tao Xiang
Sixiao Zheng
Jiachen Lu
Xiatian Zhu
Zekun Luo
Yabiao Wang
Yanwei Fu
Li Zhang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%