FMSA: Few-shot Multimodal Sentiment Analysis for Social Media via Integrated Prompt Learning and Vision-Language Models

FMSA: Мультимодальный анализ тональности социальных сетей в условиях малого числа примеров на основе интегрированного обучения с подсказками и визуально-языковых моделей
Xianxun Zhu, Heyang Feng, Erik Cambria, Xiaohan Yu, José Santamaría, Xuhui Fan, Rui Wang
2026-01-01

Query Transformer (Q-Former)few-shot multimodal sentiment analysisprompt learningsocial media datasetsvision-language models
Multimodal sentiment analysis (MSA) on social media is increasingly critical for understanding complex emotional expressions, yet it faces significant challenges in data-scarce environments where annotated multimodal content is limited. Here, we present FMSA, a novel framework that integrates prompt-based learning with advanced vision-language models to enable robust few-shot MSA. Our approach leverages an instruction-aware Query Transformer (Q-Former) to dynamically extract and align visual features with task-specific textual prompts, enhancing cross-modal fusion. We introduce a distributed consistency sampling strategy to construct representative few-shot datasets, ensuring statistical diversity under constrained conditions. Evaluated across six benchmark social media datasets-including MVSA-S, MVSA-M, and Twitter-Depression-FMSA outperforms state-of-the-art methods, achieving an accuracy of 63.47% and an F1 score of 57.34% on MVSA-S with full data, and 61.25% accuracy with 56.06% F1 in few-shot settings using just 1% of the data. By fine-tuning lightweight components while preserving pretrained model robustness, FMSA mitigates overfitting and delivers generalizable performance. We also release a curated few-shot dataset as a community resource. This framework advances MSA by offering an efficient, scalable solution for interpreting multimodal emotions in low-resource scenarios.
1
Across six social media benchmarks, FMSA achieves 63.47% accuracy and 57.34% F1 on MVSA-S using full data.
2
Distributed consistency sampling constructs statistically diverse and representative few-shot datasets under severe annotation constraints.
3
FMSA integrates instruction-aware Q-Former prompt learning with vision-language models to align visual features and task-specific textual prompts for few-shot multimodal sentiment analysis.
4
Fine-tuning lightweight components while preserving pretrained model robustness reduces overfitting and improves generalization; the authors also release a curated few-shot dataset.
5
Using only 1% of the data in few-shot settings, FMSA reaches 61.25% accuracy and 56.06% F1, outperforming state-of-the-art methods.

multimodal sentiment analysis of social media content in few-shot, data-scarce settings

robust and generalizable interpretation of multimodal emotions through cross-modal feature alignment and fusion under limited annotated data

Publication Details
Publication Date
2026-01-01
Journal
Publisher
ISSN
Cited by
3
Access Type
Author Information
Authors
Xianxun Zhu
Heyang Feng
Erik Cambria
Xiaohan Yu
José Santamaría
Xuhui Fan
Rui Wang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%