Multimodal Learning With Transformers: A Survey
Мультимодальное обучение с трансформерами: обзор
2023-05-11
SCID: 54.1/6x2cf2ed
Discuss with AI
Multimodal TransformerMultimodal learningMultimodal pretrainingTransformerVision Transformer
Figures from the paper
Abstract (AI)
Transformer is a promising neural network learner, and has achieved great success in various machine learning tasks. Thanks to the recent prevalence of multimodal applications and Big Data, Transformer-based multimodal learning has become a hot topic in AI research. This paper presents a comprehensive survey of Transformer techniques oriented at multimodal data. The main contents of this survey include: (1) a background of multimodal learning, Transformer ecosystem, and the multimodal Big Data era, (2) a systematic review of Vanilla Transformer, Vision Transformer, and multimodal Transformers, from a geometrically topological perspective, (3) a review of multimodal Transformer applications, via two important paradigms, i.e., for multimodal pretraining and for specific multimodal tasks, (4) a summary of the common challenges and designs shared by the multimodal Transformer models and applications, and (5) a discussion of open problems and potential research directions for the community.
Key Findings
1
Multimodal Transformer techniques are organized into two application paradigms: multimodal pretraining and specific multimodal tasks.
2
The paper identifies common challenges and shared design patterns across multimodal Transformer models and applications.
3
The survey discusses open problems and potential future research directions for multimodal Transformer community.
4
The survey provides a systematic review of Vanilla Transformer, Vision Transformer, and multimodal Transformers from a geometric-topological perspective.
5
Transformer architectures have become central and promising for multimodal learning due to success across machine learning tasks.
Research Object
Transformer-based multimodal learning models and techniques
Research Subject
Surveying architectures, geometric-topological perspectives, pretraining and task-specific applications, common challenges and design patterns, and open research directions for Transformer-based multimodal learning
Publication Details
Publication Date
2023-05-11
Journal
Publisher
ISSN
Cited by
969
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai16
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
The Elements of Statistical Learning: Data Mining, Inference, and Prediction2010
Non-local Neural Networks2018
Emerging Properties in Self-Supervised Vision Transformers2021
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing2022
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Longformer: The Long-Document Transformer2020
Pre-Trained Image Processing Transformer2021
An Empirical Study of Training Self-Supervised Vision Transformers2021
Flamingo: a Visual Language Model for Few-Shot Learning2022
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips2019
Neural Speech Synthesis with Transformer Network2019
Visual Dialog2017