VGSG: Vision-Guided Semantic-Group Network for Text-Based Person Search
VGSG: Сеть с визуально-направленными семантическими группами для поиска людей по тексту
2023-12-05
SCID: 54.1/qn7jhtts
Discuss with AI
Semantic-Group Textual LearningVision-Guided Semantic-Group NetworkVision-guided Knowledge Transfertext-based person searchvision-language similarity transfer
Figures from the paper
Abstract (AI)
Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing methods utilize external tools or heavy cross-modal interaction to achieve explicit alignment of cross-modal fine-grained features, which is inefficient and time-consuming. In this work, we propose a Vision-Guided Semantic-Group Network (VGSG) for text-based person search to extract well-aligned fine-grained visual and textual features. In the proposed VGSG, we develop a Semantic-Group Textual Learning (SGTL) module and a Vision-guided Knowledge Transfer (VGKT) module to extract textual local features under the guidance of visual local clues. In SGTL, in order to obtain the local textual representation, we group textual features from the channel dimension based on the semantic cues of language expression, which encourages similar semantic patterns to be grouped implicitly without external tools. In VGKT, a vision-guided attention is employed to extract visual-related textual features, which are inherently aligned with visual cues and termed vision-guided textual features. Furthermore, we design a relational knowledge transfer, including a vision-language similarity transfer and a class probability transfer, to adaptively propagate information of the vision-guided textual features to semantic-group textual features. With the help of relational knowledge transfer, VGKT is capable of aligning semantic-group textual features with corresponding visual features without external tools and complex pairwise interaction. Experimental results on two challenging benchmarks demonstrate its superiority over state-of-the-art methods.
Key Findings
1
Proposed VGSG (Vision-Guided Semantic-Group Network) extracts well-aligned fine-grained visual and textual features for text-based person search.
2
Relational knowledge transfer (vision-language similarity transfer and class probability transfer) adaptively propagates information from vision-guided textual features to semantic-group textual features, enabling alignment without complex pairwise interactions or external tools.
3
Semantic-Group Textual Learning (SGTL) groups textual features along the channel dimension by semantic cues to obtain local textual representations without external tools.
4
VGSG outperforms state-of-the-art methods on two challenging benchmarks, demonstrating superiority in text-based person search.
5
Vision-guided Knowledge Transfer (VGKT) uses vision-guided attention to produce visual-related textual features inherently aligned with visual cues (vision-guided textual features).
Research Object
Text-based person search system (Vision-Guided Semantic-Group Network applied to cross-modal pedestrian retrieval)
Research Subject
Alignment and extraction of fine-grained visual and textual local features for cross-modal retrieval, specifically vision-guided semantic-group textual feature learning and relational knowledge transfer to align textual features with corresponding visual cues
Publication Details
Publication Date
2023-12-05
Journal
Publisher
ISSN
Cited by
107
Access Type
Author Information
Download PDF
Subscribe to digest