MM-LLMs: Recent Advances in MultiModal Large Language Models

MM-LLMs: последние достижения в области мультимодальных больших языковых моделей
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, Dong Yu
2024-01-01

LLM augmentationmodel architecturemultimodal large language modelsmultimodal taskstraining pipeline
In the past year, MultiModal Large Language Models (MM-LLMs) have undergone substantial advancements, augmenting off-the-shelf LLMs to support MM inputs or outputs via cost-effective training strategies.The resulting models not only preserve the inherent reasoning and decision-making capabilities of LLMs but also empower a diverse range of MM tasks.In this paper, we provide a comprehensive survey aimed at facilitating further research on MM-LLMs.Initially, we outline general design formulations for model architecture and training pipeline.Subsequently, we introduce a taxonomy encompassing 126 MM-LLMs, each characterized by its specific formulations.Furthermore, we review the performance of selected MM-LLMs on mainstream benchmarks and summarize key training recipes to enhance the potency of MM-LLMs.Finally, we explore promising directions for MM-LLMs while concurrently maintaining a real-time tracking website 1 for the latest developments in the field.We hope that this survey contributes to the ongoing advancement of the MM-LLMs domain.
1
It introduces a taxonomy of 126 MM-LLMs, categorizing each according to its architectural and training formulations.
2
MultiModal Large Language Models increasingly extend off-the-shelf LLMs to process multimodal inputs or generate multimodal outputs through cost-effective training strategies.
3
The survey formalizes general MM-LLM architecture designs and training pipelines, providing a framework for analyzing the field.
4
The survey reviews benchmark performance, summarizes effective training recipes, identifies promising research directions, and maintains a real-time tracking website for field developments.
5
These models retain the reasoning and decision-making capabilities of LLMs while supporting a diverse range of multimodal tasks.

MultiModal Large Language Models (MM-LLMs)

Recent advances, architectural and training formulations, capabilities, benchmark performance, and development directions of MM-LLMs

Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
229
Access Type
Author Information
Authors
Duzhen Zhang
Yahan Yu
Jiahua Dong
Chenxing Li
Dan Su
Chenhui Chu
Dong Yu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%