MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection

MS-TCT: Многомасштабный временной ConvTransformer для обнаружения действий
Michael S. Ryoo, Rui Dai, Srijan Das, Kumara Kahatapitiya, François Brémond
2022-06-01

MS-TCTMulti-Scale Temporal ConvTransformerTemporal EncoderTemporal Scale Mixerframe-level classification
Action detection is a significant and challenging task, especially in densely-labelled datasets of untrimmed videos. Such data consist of complex temporal relations including composite or co-occurring actions. To detect actions in these complex settings, it is critical to capture both shortterm and long-term temporal information efficiently. To this end, we propose a novel ‘ConvTransformer’ network for action detection: MS-TCT <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code/Models: https://github.com/dairui01/MS-TCT. This network comprises of three main components: (1) a Temporal Encoder module which explores global and local temporal relations at multiple temporal resolutions, (2) a Temporal Scale Mixer module which effectively fuses multi-scale features, creating a unified feature representation, and (3) a Classification module which learns a center-relative position of each action instance in time, and predicts frame-level classification scores. Our experimental results on multiple challenging datasets such as Charades, TSU and MultiTHUMOS, validate the effectiveness of the proposed method, which outperforms the state-of-the-art methods on all three datasets.
1
A Classification module learns center-relative temporal positions for each action instance and predicts frame-level classification scores.
2
A Temporal Scale Mixer module fuses multi-scale temporal features into a unified feature representation.
3
MS-TCT is a novel ConvTransformer network designed for action detection in densely-labelled untrimmed videos.
4
MS-TCT outperforms state-of-the-art methods on Charades, TSU, and MultiTHUMOS datasets.
5
The network contains a Temporal Encoder that captures both global and local temporal relations at multiple temporal resolutions.

MS-TCT ConvTransformer network for action detection in untrimmed videos

Capturing and fusing multi-scale short-term and long-term temporal relations to predict frame-level action instances and center-relative action positions for dense action detection

Publication Details
Publication Date
2022-06-01
Journal
Publisher
ISSN
Cited by
106
Access Type
Author Information
Authors
Michael S. Ryoo
Rui Dai
Srijan Das
Kumara Kahatapitiya
François Brémond
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%