Attention mechanisms in computer vision: A survey
Механизмы внимания в компьютерном зрении: обзор
2022-03-15
SCID: 54.1/x26sq9wq
Discuss with AI
attention mechanismschannel attentioncomputer visionspatial attentiontemporal attention
Figures from the paper
Abstract (AI)
Humans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multimodal tasks, and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention, and branch attention; a related repository https://github.com/MenghaoGuo/Awesome-Vision-Attentions is dedicated to collecting related work. We also suggest future directions for attention mechanism research.
Key Findings
1
Attention mechanisms dynamically adjust feature weights according to input-image characteristics, modeling the human visual system’s focus on salient regions.
2
Attention mechanisms have achieved substantial success across image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multimodal tasks, and self-supervised learning.
3
The paper provides a dedicated repository collecting related vision-attention research and identifies future directions for advancing attention mechanisms.
4
The survey comprehensively reviews attention mechanisms developed for computer vision and organizes them into channel, spatial, temporal, and branch attention categories.
Research Object
Attention mechanisms in computer vision
Research Subject
Their approaches, categories, and applications across visual tasks, including dynamic feature-weight adjustment and future research directions
Publication Details
Publication Date
2022-03-15
Journal
Publisher
ISSN
Cited by
2511
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai17
ImageNet: A large-scale hierarchical image database2009
Adam: A Method for Stochastic Optimization2015
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Aggregated Residual Transformations for Deep Neural Networks2017
Proceedings of the 24th international conference on Machine learning2007
Non-local Neural Networks2018
ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks2020
Emerging Properties in Self-Supervised Vision Transformers2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Point Transformer2021
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Pre-Trained Image Processing Transformer2021
An Empirical Study of Training Self-Supervised Vision Transformers2021
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
Self-Attention with Relative Position Representations2018