COCOCOCO datasetCharadesCharades datasetKineticsKinetics datasetlong-range dependenciesnon-local meansnon-local neural networksnon-local operationnon-local operationsobject detectionpose estimationsegmentationvideo classification
Figures from the paper
Abstract (AI)
Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time. In this paper, we present non-local operations as a generic family of building blocks for capturing long-range dependencies. Inspired by the classical non-local means method in computer vision, our non-local operation computes the response at a position as a weighted sum of the features at all positions. This building block can be plugged into many computer vision architectures. On the task of video classification, even without any bells and whistles, our non-local models can compete or outperform current competition winners on both Kinetics and Charades datasets. In static image recognition, our non-local models improve object detection/segmentation and pose estimation on the COCO suite of tasks. Code is available at https://github.com/facebookresearch/video-nonlocal-net .
Key Findings
1
For static image tasks, non-local models improve object detection and segmentation performance on the COCO benchmark.
2
For static image tasks, non-local models improve pose estimation performance on the COCO benchmark.
3
Non-local operations are introduced as a generic building block that captures long-range dependencies by computing each position's response as a weighted sum over all positions.
4
On static image recognition tasks, non-local models improve object detection, segmentation, and pose estimation on the COCO suite of tasks.
5
On video classification, non-local models can compete with or outperform current competition winners on the Charades dataset without extra bells and whistles.
6
On video classification, non-local models can compete with or outperform current competition winners on the Kinetics and Charades datasets, even without extra bells and whistles.
7
On video classification, non-local models can compete with or outperform current competition winners on the Kinetics dataset without extra bells and whistles.
8
The authors provide code implementation publicly at the specified GitHub repository.
9
The authors provide code implementing non-local neural networks at the referenced GitHub repository.
10
The non-local building block is inspired by non-local means in computer vision and can be plugged into many computer vision architectures.
11
The non-local building block is inspired by the classical non-local means method and can be plugged into many computer vision architectures.
Research Object
Non-local operations (non-local neural network building block) applied within computer vision architectures for video and image tasks
Research Subject
Ability of the non-local operation to capture long-range dependencies and improve performance (video classification, object detection/segmentation, pose estimation) when integrated into standard convolutional/recurrent architectures
Publication Details
Publication Date
2018-06-01
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai7
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Aggregated Residual Transformations for Deep Neural Networks2017
Interaction Networks for Learning about Objects, Relations and Physics2016
Deep Residual Learning for Image Recognition2016
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
Improving neural networks by preventing co-adaptation of feature detectors2012
The Graph Neural Network Model2008
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Restormer: Efficient Transformer for High-Resolution Image Restoration2022
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation2021
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
UNETR: Transformers for 3D Medical Image Segmentation2022
Attention mechanisms in computer vision: A survey2022
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Uformer: A General U-Shaped Transformer for Image Restoration2022
Pre-Trained Image Processing Transformer2021
Video Swin Transformer2022
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
An Empirical Study of Training Self-Supervised Vision Transformers2021
Remote Sensing Image Change Detection With Transformers2021
Multimodal Learning With Transformers: A Survey2023
CMT: Convolutional Neural Networks Meet Vision Transformers2022
A Small-Sized Object Detection Oriented Multi-Scale Feature Fusion Approach With Application to Defect Detection2022
Spectral–Spatial Transformer Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020