Vision Transformers for Single Image Dehazing
Vision Transformers для удаления тумана с одиночного изображения
2023-01-01
SCID: 54.1/pjd7kvep
Discuss with AI
DehazeFormerPSNRSOTS indoorSwin TransformerVision Transformersactivation functionimage dehazingnormalization layerremote sensing dehazing datasetspatial information aggregation
Figures from the paper
Abstract (AI)
Image dehazing is a representative low-level vision task that estimates latent haze-free images from hazy images. In recent years, convolutional neural network-based methods have dominated image dehazing. However, vision Transformers, which has recently made a breakthrough in high-level vision tasks, has not brought new dimensions to image dehazing. We start with the popular Swin Transformer and find that several of its key designs are unsuitable for image dehazing. To this end, we propose DehazeFormer, which consists of various improvements, such as the modified normalization layer, activation function, and spatial information aggregation scheme. We train multiple variants of DehazeFormer on various datasets to demonstrate its effectiveness. Specifically, on the most frequently used SOTS indoor set, our small model outperforms FFA-Net with only 25% #Param and 5% computational cost. To the best of our knowledge, our large model is the first method with the PSNR over 40 dB on the SOTS indoor set, dramatically outperforming the previous state-of-the-art methods. We also collect a large-scale realistic remote sensing dehazing dataset for evaluating the method's capability to remove highly non-homogeneous haze. We share our code and dataset at https://github.com/IDKiro/DehazeFormer.
Key Findings
1
A large DehazeFormer model achieves PSNR over 40 dB on the SOTS indoor set, substantially surpassing previous state-of-the-art.
2
A new large-scale realistic remote sensing dehazing dataset was collected to evaluate removal of highly non-homogeneous haze; code and dataset are publicly released.
3
A small DehazeFormer variant outperforms FFA-Net on the SOTS indoor set while using only 25% of parameters and 5% of computational cost.
4
DehazeFormer introduces modified normalization, activation, and spatial information aggregation tailored for dehazing.
5
Vision Transformers can be adapted for single image dehazing by modifying Swin Transformer design elements unsuitable for dehazing.
Research Object
Vision Transformer-based models for single-image dehazing (DehazeFormer variants)
Research Subject
Design modifications and performance of Transformer-based architectures for single-image dehazing, including normalization, activation, spatial information aggregation, efficiency (parameters and computation), and restoration quality (PSNR) on benchmark and remote-sensing dehazing datasets
Publication Details
Publication Date
2023-01-01
Journal
Publisher
ISSN
Cited by
1149
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
SwinIR: Image Restoration Using Swin Transformer2021
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022