Vision Transformers for Remote Sensing Image Classification

Визуальные трансформеры для классификации изображений дистанционного зондирования
Yakoub Bazi, Laila Bashmal, Mohamad Mahmoud Al Rahhal, Reham Al-Dayil, Naif Al Ajlan
2021-02-01

Vision Transformersdata augmentationmultihead attentionnetwork pruningremote sensing image classification
In this paper, we propose a remote-sensing scene-classification method based on vision transformers. These types of networks, which are now recognized as state-of-the-art models in natural language processing, do not rely on convolution layers as in standard convolutional neural networks (CNNs). Instead, they use multihead attention mechanisms as the main building block to derive long-range contextual relation between pixels in images. In a first step, the images under analysis are divided into patches, then converted to sequence by flattening and embedding. To keep information about the position, embedding position is added to these patches. Then, the resulting sequence is fed to several multihead attention layers for generating the final representation. At the classification stage, the first token sequence is fed to a softmax classification layer. To boost the classification performance, we explore several data augmentation strategies to generate additional data for training. Moreover, we show experimentally that we can compress the network by pruning half of the layers while keeping competing classification accuracies. Experimental results conducted on different remote-sensing image datasets demonstrate the promising capability of the model compared to state-of-the-art methods. Specifically, Vision Transformer obtains an average classification accuracy of 98.49%, 95.86%, 95.56% and 93.83% on Merced, AID, Optimal31 and NWPU datasets, respectively. While the compressed version obtained by removing half of the multihead attention layers yields 97.90%, 94.27%, 95.30% and 93.05%, respectively.
1
A Vision Transformer method classifies remote-sensing scenes using patch embeddings, positional encodings, and multihead attention instead of convolutional layers.
2
Data augmentation strategies are used to generate additional training data and improve classification performance.
3
Pruning half of the multihead attention layers preserves competitive performance, yielding accuracies of 97.90%, 94.27%, 95.30%, and 93.05% on the respective datasets.
4
The full Vision Transformer achieves average accuracies of 98.49% on Merced, 95.86% on AID, 95.56% on Optimal31, and 93.83% on NWPU.
5
The model captures long-range contextual relationships between image pixels through stacked multihead attention mechanisms.

remote-sensing images and their scene-classification datasets

vision-transformer-based scene-classification performance, including the effects of multihead attention, data augmentation, and layer pruning

Publication Details
Publication Date
2021-02-01
Journal
Publisher
ISSN
Cited by
637
Access Type
Author Information
Authors
Yakoub Bazi
Laila Bashmal
Mohamad Mahmoud Al Rahhal
Reham Al-Dayil
Naif Al Ajlan
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%