E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
E-VRAG: Повышение качества понимания длинных видео с помощью ресурсоэффективной генерации с поддержкой поиска
2025-08-03
SCID: 54.1/bqnaygtv
Discuss with AI
E-VRAGcomputational cost reductionframe pre-filteringframe scoringglobal statistical distribution of inter-frame scoreshierarchical query decompositionlightweight VLMlong video understandingmulti-view question answeringretrieval-augmented generationvideo RAG
Figures from the paper
Abstract (AI)
Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by the restricted context window and the high computational cost required to process long videos with thousands of frames. Retrieval-augmented generation (RAG) addresses this challenge by selecting only the most relevant frames as input, thereby reducing the computational burden. Nevertheless, existing video RAG methods struggle to balance retrieval efficiency and accuracy, particularly when handling diverse and complex video content. To address these limitations, we propose E-VRAG, a novel and efficient video RAG framework for video understanding. We first apply a frame pre-filtering method based on hierarchical query decomposition to eliminate irrelevant frames, reducing computational costs at the data level. We then employ a lightweight VLM for frame scoring, further reducing computational costs at the model level. Additionally, we propose a frame retrieval strategy that leverages the global statistical distribution of inter-frame scores to mitigate the potential performance degradation from using a lightweight VLM. Finally, we introduce a multi-view question answering scheme for the retrieved frames, enhancing the VLM's capability to extract and comprehend information from long video contexts. Experiments on four public benchmarks show that E-VRAG achieves about 70% reduction in computational cost and higher accuracy compared to baseline methods, all without additional training. These results demonstrate the effectiveness of E-VRAG in improving both efficiency and accuracy for video RAG tasks.
Key Findings
1
A frame retrieval strategy leveraging the global statistical distribution of inter-frame scores mitigates performance degradation from using a lightweight VLM.
2
A hierarchical query decomposition pre-filtering step removes irrelevant frames, reducing computational costs at the data level.
3
A lightweight vision-language model is used for frame scoring to further cut computational costs at the model level.
4
A multi-view question answering scheme over retrieved frames enhances the VLM's ability to extract and comprehend information from long video contexts.
5
E-VRAG is a novel retrieval-augmented generation (RAG) framework designed to improve long video understanding while being resource-efficient.
6
On four public benchmarks, E-VRAG achieves about 70% reduction in computational cost and higher accuracy than baseline methods without additional training.
Research Object
Long videos processed for video understanding using retrieval-augmented generation (video RAG)
Research Subject
Resource-efficient retrieval-augmented generation methods (E-VRAG) to select and score relevant frames and answer multi-view questions so as to improve computational efficiency and accuracy of long-video understanding
Publication Details
Publication Date
2025-08-03
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest