Manifold knowledge-guided feature fusion network for multimodal sentiment analysis

Сеть слияния признаков, управляемая многообразным знанием, для мультимодального анализа сентимента
Yijia Zhang, Xingang Wang, Mengyi Wang, Hai Cui
2025-04-09

knowledge contrastive learningknowledge filtermanifold knowledge-guided feature fusion networkmanifold learningmultimodal sentiment analysis
With the continuous progress of multimedia and information technology, multimodal sentiment analysis (MSA) has become one of the most advanced and challenging research directions in the field of artificial intelligence . Multimodal data, including text, visual and audio information, provides additional perspectives for sentiment analysis . However, extraneous information in non-verbal modalities affects the accuracy of sentiment analysis, as sentiment-related features are mainly concentrated in changes in mouth movements and pitch changes, which poses a challenge for accurate sentiment analysis. To solve this problem, we propose a manifold knowledge-guided feature fusion network (MKGN). MKGN uses manifold knowledge generated by manifold learning algorithms to guide neural networks to extract effective non-verbal features and establish associations between multiple features while reducing dimensionality. In addition, in order to improve the quality of knowledge, we propose two knowledge enhancement methods: knowledge filter (KF) and knowledge contrastive learning (CL). Among them, KF is used to filter out unreliable knowledge, and CL further strengthens retained knowledge by changing the distance between knowledge. Importantly, the proposed MKGN achieves excellent performance on three datasets compared to state-of-the-art models. On the MOSI dataset, the accuracy is improved by 2% and 1%, respectively. On the MOSEI dataset, the accuracy improved by 3.8% and 1.8%, respectively. On the UR-FUNNY dataset, the accuracy improved by 0.4%.
1
Introduced two knowledge enhancement methods: a knowledge filter (KF) to remove unreliable manifold knowledge and knowledge contrastive learning (CL) to strengthen retained knowledge by altering knowledge distances.
2
MKGN establishes associations between multimodal features while mitigating extraneous non-verbal information (e.g., irrelevant visual/audio cues) that harms sentiment analysis.
3
MKGN outperforms state-of-the-art models on three datasets: MOSI accuracy improved by 2% and 1%, MOSEI accuracy improved by 3.8% and 1.8%, and UR-FUNNY accuracy improved by 0.4%.
4
Proposed a manifold knowledge-guided feature fusion network (MKGN) that uses manifold learning-generated knowledge to guide extraction of effective non-verbal features and reduce dimensionality.

Multimodal sentiment analysis system operating on text, visual, and audio data (specifically non-verbal features like mouth movements and pitch) using a manifold knowledge-guided feature fusion network (MKGN)

Extraction and fusion of effective non-verbal features and their associations via manifold-knowledge guidance, with dimensionality reduction and knowledge-quality enhancement (knowledge filter and contrastive learning) to improve multimodal sentiment classification accuracy

Publication Details
Publication Date
2025-04-09
Journal
Publisher
ISSN
Cited by
7
Access Type
Author Information
Authors
Yijia Zhang
Xingang Wang
Mengyi Wang
Hai Cui
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%