Uformer: A General U-Shaped Transformer for Image Restoration

Uformer: Общая U-образная трансформерная архитектура для восстановления изображений
Jianmin Bao, Wengang Zhou, Houqiang Li, Zhendong Wang, Xiaodong Cun, Jianzhuang Liu
2022-06-01

U-shaped TransformerUformerdefocus deblurringderainingimage denoisingimage restorationlearnable multi-scale restoration modulatorlocally-enhanced window (LeWin) Transformer blockmotion deblurringmulti-scale spatial biasnon-overlapping window-based self-attention
In this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core designs. First, we introduce a novel locally-enhanced window (LeWin) Transformer block, which performs non-overlapping window-based self-attention instead of global self-attention. It significantly reduces the computational complexity on high resolution feature map while capturing local context. Second, we propose a learnable multi-scale restoration modulator in the form of a multi-scale spatial bias to adjust features in multiple layers of the Uformer decoder. Our modulator demonstrates superior capability for restoring details for various image restoration tasks while introducing marginal extra parameters and computational cost. Powered by these two designs, Uformer enjoys a high capability for capturing both local and global dependencies for image restoration. To evaluate our approach, extensive experiments are conducted on several image restoration tasks, including image denoising, motion deblurring, defocus deblurring and deraining. Without bells and whistles, our Uformer achieves superior or comparable performance compared with the state-of-the-art algorithms. The code and models are available at https://github.com/ZhendongWang6/Uformer.
1
A learnable multi-scale restoration modulator, implemented as a multi-scale spatial bias in the decoder, improves detail restoration across multiple layers with marginal extra parameters and compute.
2
Combining LeWin blocks and the multi-scale modulator enables Uformer to capture both local and global dependencies effectively for image restoration.
3
Extensive experiments on image denoising, motion deblurring, defocus deblurring, and deraining show Uformer achieves superior or comparable performance to state-of-the-art methods without additional bells and whistles.
4
The LeWin (locally-enhanced window) Transformer block uses non-overlapping window-based self-attention to capture local context while significantly reducing computational complexity on high-resolution feature maps.
5
Uformer is a hierarchical encoder-decoder Transformer architecture specifically designed for image restoration tasks.

Uformer Transformer-based hierarchical encoder-decoder architecture for image restoration

Design and evaluation of locally-enhanced window (LeWin) Transformer blocks and a learnable multi-scale spatial-bias restoration modulator to capture local and global dependencies and improve performance on image restoration tasks (denoising, deblurring, deraining)

Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
2219
Access Type
Author Information
Authors
Jianmin Bao
Wengang Zhou
Houqiang Li
Zhendong Wang
Xiaodong Cun
Jianzhuang Liu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%