MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection
MS-TCT: Многомасштабный временной ConvTransformer для обнаружения действий
2022-06-01
SCID: 54.1/qrt4rm8a
Discuss with AI
MS-TCTMulti-Scale Temporal ConvTransformerTemporal EncoderTemporal Scale Mixerframe-level classification
Figures from the paper
Abstract (AI)
Action detection is a significant and challenging task, especially in densely-labelled datasets of untrimmed videos. Such data consist of complex temporal relations including composite or co-occurring actions. To detect actions in these complex settings, it is critical to capture both shortterm and long-term temporal information efficiently. To this end, we propose a novel ‘ConvTransformer’ network for action detection: MS-TCT <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code/Models: https://github.com/dairui01/MS-TCT. This network comprises of three main components: (1) a Temporal Encoder module which explores global and local temporal relations at multiple temporal resolutions, (2) a Temporal Scale Mixer module which effectively fuses multi-scale features, creating a unified feature representation, and (3) a Classification module which learns a center-relative position of each action instance in time, and predicts frame-level classification scores. Our experimental results on multiple challenging datasets such as Charades, TSU and MultiTHUMOS, validate the effectiveness of the proposed method, which outperforms the state-of-the-art methods on all three datasets.
Key Findings
1
A Classification module learns center-relative temporal positions for each action instance and predicts frame-level classification scores.
2
A Temporal Scale Mixer module fuses multi-scale temporal features into a unified feature representation.
3
MS-TCT is a novel ConvTransformer network designed for action detection in densely-labelled untrimmed videos.
4
MS-TCT outperforms state-of-the-art methods on Charades, TSU, and MultiTHUMOS datasets.
5
The network contains a Temporal Encoder that captures both global and local temporal relations at multiple temporal resolutions.
Research Object
MS-TCT ConvTransformer network for action detection in untrimmed videos
Research Subject
Capturing and fusing multi-scale short-term and long-term temporal relations to predict frame-level action instances and center-relative action positions for dense action detection
Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
106
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai10
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers2021
PVT v2: Improved baselines with pyramid vision transformer2022
Video Swin Transformer2022
CMT: Convolutional Neural Networks Meet Vision Transformers2022
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Self-Attention Temporal Convolutional Network for Long-Term Daily Living Activity Detection2019
Adam: A Method for Stochastic Optimization2014