Vision Transformers for Remote Sensing Image Classification
Визуальные трансформеры для классификации изображений дистанционного зондирования
2021-02-01
SCID: 54.1/9ajj5hfw
Discuss with AI
Vision Transformersdata augmentationmultihead attentionnetwork pruningremote sensing image classification
Figures from the paper
Abstract (AI)
In this paper, we propose a remote-sensing scene-classification method based on vision transformers. These types of networks, which are now recognized as state-of-the-art models in natural language processing, do not rely on convolution layers as in standard convolutional neural networks (CNNs). Instead, they use multihead attention mechanisms as the main building block to derive long-range contextual relation between pixels in images. In a first step, the images under analysis are divided into patches, then converted to sequence by flattening and embedding. To keep information about the position, embedding position is added to these patches. Then, the resulting sequence is fed to several multihead attention layers for generating the final representation. At the classification stage, the first token sequence is fed to a softmax classification layer. To boost the classification performance, we explore several data augmentation strategies to generate additional data for training. Moreover, we show experimentally that we can compress the network by pruning half of the layers while keeping competing classification accuracies. Experimental results conducted on different remote-sensing image datasets demonstrate the promising capability of the model compared to state-of-the-art methods. Specifically, Vision Transformer obtains an average classification accuracy of 98.49%, 95.86%, 95.56% and 93.83% on Merced, AID, Optimal31 and NWPU datasets, respectively. While the compressed version obtained by removing half of the multihead attention layers yields 97.90%, 94.27%, 95.30% and 93.05%, respectively.
Key Findings
1
A Vision Transformer method classifies remote-sensing scenes using patch embeddings, positional encodings, and multihead attention instead of convolutional layers.
2
Data augmentation strategies are used to generate additional training data and improve classification performance.
3
Pruning half of the multihead attention layers preserves competitive performance, yielding accuracies of 97.90%, 94.27%, 95.30%, and 93.05% on the respective datasets.
4
The full Vision Transformer achieves average accuracies of 98.49% on Merced, 95.86% on AID, 95.56% on Optimal31, and 93.83% on NWPU.
5
The model captures long-range contextual relationships between image pixels through stacked multihead attention mechanisms.
Research Object
remote-sensing images and their scene-classification datasets
Research Subject
vision-transformer-based scene-classification performance, including the effects of multihead attention, data augmentation, and layer pruning
Publication Details
Publication Date
2021-02-01
Journal
Publisher
ISSN
Cited by
637
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest