E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

E-VRAG: Повышение качества понимания длинных видео с помощью ресурсоэффективной генерации с поддержкой поиска
Zeyu Xu, J Zhang, Qiang Wang, Yi Liu
2025-08-03

E-VRAGcomputational cost reductionframe pre-filteringframe scoringglobal statistical distribution of inter-frame scoreshierarchical query decompositionlightweight VLMlong video understandingmulti-view question answeringretrieval-augmented generationvideo RAG
Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by the restricted context window and the high computational cost required to process long videos with thousands of frames. Retrieval-augmented generation (RAG) addresses this challenge by selecting only the most relevant frames as input, thereby reducing the computational burden. Nevertheless, existing video RAG methods struggle to balance retrieval efficiency and accuracy, particularly when handling diverse and complex video content. To address these limitations, we propose E-VRAG, a novel and efficient video RAG framework for video understanding. We first apply a frame pre-filtering method based on hierarchical query decomposition to eliminate irrelevant frames, reducing computational costs at the data level. We then employ a lightweight VLM for frame scoring, further reducing computational costs at the model level. Additionally, we propose a frame retrieval strategy that leverages the global statistical distribution of inter-frame scores to mitigate the potential performance degradation from using a lightweight VLM. Finally, we introduce a multi-view question answering scheme for the retrieved frames, enhancing the VLM's capability to extract and comprehend information from long video contexts. Experiments on four public benchmarks show that E-VRAG achieves about 70% reduction in computational cost and higher accuracy compared to baseline methods, all without additional training. These results demonstrate the effectiveness of E-VRAG in improving both efficiency and accuracy for video RAG tasks.
1
A frame retrieval strategy leveraging the global statistical distribution of inter-frame scores mitigates performance degradation from using a lightweight VLM.
2
A hierarchical query decomposition pre-filtering step removes irrelevant frames, reducing computational costs at the data level.
3
A lightweight vision-language model is used for frame scoring to further cut computational costs at the model level.
4
A multi-view question answering scheme over retrieved frames enhances the VLM's ability to extract and comprehend information from long video contexts.
5
E-VRAG is a novel retrieval-augmented generation (RAG) framework designed to improve long video understanding while being resource-efficient.
6
On four public benchmarks, E-VRAG achieves about 70% reduction in computational cost and higher accuracy than baseline methods without additional training.

Long videos processed for video understanding using retrieval-augmented generation (video RAG)

Resource-efficient retrieval-augmented generation methods (E-VRAG) to select and score relevant frames and answer multi-view questions so as to improve computational efficiency and accuracy of long-video understanding

Publication Details
Publication Date
2025-08-03
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Zeyu Xu
J Zhang
Qiang Wang
Yi Liu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%