MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Yucheng Zhang, Jenq–Neng Hwang, Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Gaoang Wang
2024-06-16

SCID:  54.1/yhzgw4bz
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat.
Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Yucheng Zhang
Jenq–Neng Hwang
Enxin Song
Wenhao Chai
Guanhong Wang
Haoyang Zhou
Feiyang Wu
Haozhe Chi
Xun Guo
Tian Ye
Yanting Zhang
Yan Lu
Gaoang Wang
Explore More Research
Use the citation graph to discover related papers and expand your research horizons.
Click any node to explore
Download PDF
100%