Remote Sensing Image Change Detection With Transformers
Обнаружение изменений на дистанционно зондируемых изображениях с использованием трансформеров
2021-07-20
SCID: 54.1/pqj69q5f
Discuss with AI
bitemporal image transformer (BIT)deep feature differencingnonlocal self-attentionremote sensing change detectiontoken-based spatial-temporal modeling
Figures from the paper
Abstract (AI)
Modern change detection (CD) has achieved remarkable success by the powerful discriminative ability of deep convolutions. However, high-resolution remote sensing CD remains challenging due to the complexity of objects in the scene. Objects with the same semantic concept may show distinct spectral characteristics at different times and spatial locations. Most recent CD pipelines using pure convolutions are still struggling to relate long-range concepts in space-time. Nonlocal self-attention approaches show promising performance via modeling dense relationships among pixels, yet are computationally inefficient. Here, we propose a bitemporal image transformer (BIT) to efficiently and effectively model contexts within the spatial-temporal domain. Our intuition is that the high-level concepts of the change of interest can be represented by a few visual words, that is, semantic tokens. To achieve this, we express the bitemporal image into a few tokens and use a transformer encoder to model contexts in the compact token-based space-time. The learned context-rich tokens are then fed back to the pixel-space for refining the original features via a transformer decoder. We incorporate BIT in a deep feature differencing-based CD framework. Extensive experiments on three CD datasets demonstrate the effectiveness and efficiency of the proposed method. Notably, our BIT-based model significantly outperforms the purely convolutional baseline using only three times lower computational costs and model parameters. Based on a naive backbone (ResNet18) without sophisticated structures (e.g., feature pyramid network (FPN) and UNet), our model surpasses several state-of-the-art CD methods, including better than four recent attention-based methods in terms of efficiency and accuracy. Our code is available athttps://github.com/justchenhao/BIT_CD.
Key Findings
1
BIT operates in a compact token-based space-time to learn context-rich tokens, which are fed back to pixel-space to refine original features via a transformer decoder.
2
BIT-based model significantly outperforms a purely convolutional baseline while using three times lower computational cost and fewer model parameters.
3
Incorporated BIT into a deep feature differencing change detection framework and demonstrated effectiveness and efficiency on three CD datasets.
4
Proposed a Bitemporal Image Transformer (BIT) that models spatial-temporal contexts by converting bitemporal images into a few semantic tokens and using a transformer encoder-decoder pipeline.
5
With a naive ResNet18 backbone (no FPN or UNet), the BIT model surpasses several state-of-the-art CD methods, including four recent attention-based methods, in both efficiency and accuracy.
Research Object
Bitemporal high-resolution remote sensing image pairs used for change detection
Research Subject
Modeling and detecting semantic changes via a token-based bitemporal image transformer (BIT) that encodes compact spatial–temporal contexts, refines pixel features through a transformer decoder, and improves accuracy and efficiency of deep feature differencing-based change detection
Publication Details
Publication Date
2021-07-20
Journal
Publisher
ISSN
Cited by
1150
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest