Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Мультимодальные базовые модели: от специализированных систем к универсальным помощникам
Linjie Li, Zhengyuan Yang, Zhe Gan, Jianfeng Gao, Chunyuan Li, Lijuan Wang, Jianwei Yang
2024-05-06

multimodal LLMsmultimodal foundation modelstext-to-image generationvision backbonesvision-language
This monograph presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to generalpurpose assistants. The research landscape encompasses five core topics, categorized into two classes. (i) We start with a survey of well-established research areas: multimodal foundation models pre-trained for specific purposes, including two topics – methods of learning vision backbones for visual understanding and text-to-image generation. (ii) Then, we present recent advances in exploratory, open research areas: multimodal foundation models that aim to play the role of general-purpose assistants, including three topics – unified vision models inspired by large language models (LLMs), end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs. The target audiences of the monograph are researchers, graduate students, and professionals in computer vision and vision-language multimodal communities who are eager to learn the basics and recent advances in multimodal foundation models.
1
Exploratory research toward general-purpose assistants covers unified vision models inspired by LLMs, end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs.
2
Specialist pre-trained models include methods for learning vision backbones for visual understanding and text-to-image generation.
3
Survey categorizes multimodal foundation model research into five core topics across two classes: specialist pre-trained models and exploratory general-purpose assistant models.
4
Target audience comprises researchers, graduate students, and professionals seeking foundational knowledge and recent advances in vision and vision-language multimodal models.
5
The monograph emphasizes the transition trajectory from specialist multimodal models to general-purpose multimodal assistants.

Multimodal foundation models with vision and vision-language capabilities

The taxonomy, evolution, and transition from specialist multimodal models to general-purpose assistant models, including methods for vision backbones, text-to-image generation, unified vision models inspired by LLMs, end-to-end multimodal LLM training, and chaining multimodal tools with LLMs

Publication Details
Publication Date
2024-05-06
Journal
Publisher
ISSN
Cited by
123
Access Type
Author Information
Authors
Linjie Li
Zhengyuan Yang
Zhe Gan
Jianfeng Gao
Chunyuan Li
Lijuan Wang
Jianwei Yang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%