Pre-Trained Image Processing Transformer
Предобученный трансформер для обработки изображений
2021-06-01
SCID: 54.1/6etjv6mr
Discuss with AI
ImageNet corrupted image pairsPre-Trained Image Processing Transformercontrastive learningdenoisingderainingfine-tuningimage processing transformer (IPT)low-level computer visionmulti-heads and multi-tailssuper-resolution
Figures from the paper
Abstract (AI)
As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at https://github.com/huawei-noah/Pretrained-IPT and https://gitee.com/mindspore/mindspore/tree/master/model_zoo/research/cv/IPT
Key Findings
1
A single pre-trained IPT model can be fine-tuned and efficiently employed for various low-level vision tasks.
2
Contrastive learning is incorporated into IPT to improve adaptation across different image processing tasks.
3
IPT is trained on large-scale corrupted image pairs generated from ImageNet, leveraging multi-heads and multi-tails architecture for multiple image processing tasks.
4
IPT outperforms current state-of-the-art methods on multiple low-level vision benchmarks.
5
Introduced Image Processing Transformer (IPT), a pre-trained transformer model for low-level vision tasks like denoising, super-resolution, and deraining.
Research Object
Image Processing Transformer (IPT) pre-trained model for low-level image restoration tasks
Research Subject
Pre-training and adaptation of a transformer-based model using large-scale corrupted ImageNet pairs and contrastive learning to improve performance on low-level image processing tasks (denoising, super-resolution, deraining) after fine-tuning
Publication Details
Publication Date
2021-06-01
Journal
Publisher
ISSN
Cited by
2101
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai10
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
ImageNet: A large-scale hierarchical image database2009
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Non-local Neural Networks2018
Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising2017
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Learning Deep Transformer Models for Machine Translation2019
Cited by11
SwinIR: Image Restoration Using Swin Transformer2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
Attention mechanisms in computer vision: A survey2022
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Uformer: A General U-Shaped Transformer for Image Restoration2022
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
Remote Sensing Image Change Detection With Transformers2021
Transformers in Time Series: A Survey2023
Transformer in Transformer2021
Multimodal Learning With Transformers: A Survey2023
CMT: Convolutional Neural Networks Meet Vision Transformers2022