Pre-Trained Image Processing Transformer

Предобученный трансформер для обработки изображений
Chunjing Xu, Yunhe Wang, Chao Xu, Chang Xu, Tianyu Guo, Hanting Chen, Yiping Deng, Zhenhua Liu, Siwei Ma, Wen Gao
2021-06-01

ImageNet corrupted image pairsPre-Trained Image Processing Transformercontrastive learningdenoisingderainingfine-tuningimage processing transformer (IPT)low-level computer visionmulti-heads and multi-tailssuper-resolution
As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at https://github.com/huawei-noah/Pretrained-IPT and https://gitee.com/mindspore/mindspore/tree/master/model_zoo/research/cv/IPT
1
A single pre-trained IPT model can be fine-tuned and efficiently employed for various low-level vision tasks.
2
Contrastive learning is incorporated into IPT to improve adaptation across different image processing tasks.
3
IPT is trained on large-scale corrupted image pairs generated from ImageNet, leveraging multi-heads and multi-tails architecture for multiple image processing tasks.
4
IPT outperforms current state-of-the-art methods on multiple low-level vision benchmarks.
5
Introduced Image Processing Transformer (IPT), a pre-trained transformer model for low-level vision tasks like denoising, super-resolution, and deraining.

Image Processing Transformer (IPT) pre-trained model for low-level image restoration tasks

Pre-training and adaptation of a transformer-based model using large-scale corrupted ImageNet pairs and contrastive learning to improve performance on low-level image processing tasks (denoising, super-resolution, deraining) after fine-tuning

Publication Details
Publication Date
2021-06-01
Journal
Publisher
ISSN
Cited by
2101
Access Type
Author Information
Authors
Chunjing Xu
Yunhe Wang
Chao Xu
Chang Xu
Tianyu Guo
Hanting Chen
Yiping Deng
Zhenhua Liu
Siwei Ma
Wen Gao
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%