Benchmark Evaluations, Applications, and Challenges of Large Vision Language Models: A Survey
Бенчмаркинговая оценка, приложения и проблемы больших зрительно-языковых моделей: обзор
2025-01-15
SCID: 54.1/hytx9jqy
Discuss with AI
Vision Language Modelsbenchmarks and evaluation metricshallucinationmultimodalzero-shot classification
Figures from the paper
Abstract (AI)
Multimodal Vision Language Models (vlms) have emerged as a transformative technology at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual modalities. For example, models such as CLIP[1], Claude[2], and GPT-4V[3] demonstrate strong reasoning and understanding abilities on visual and textual data and beat classical single modality vision models on zero-shot classification[4]. Despite their rapid advancements in research and growing popularity in applications, a comprehensive survey of existing studies on vlms is notably lacking, particularly for researchers aiming to leverage vlms in their specific domains. To this end, we provide a systematic overview of vlms in the following aspects: [1] model information of the major vlms developed over the past five years (2019-2024); [2] the main architectures and training methods of these vlms; [3] summary and categorization of the popular benchmarks and evaluation metrics of vlms; [4] the applications of vlms including embodied agents, robotics, and video generation; [5] the challenges and issues faced by current vlms such as hallucination, fairness, and safety. Detailed collections including papers and model repository links are listed in https://github.com/zli12321/Awesome-VLM-Papers-And-Models.git.
Key Findings
1
Applications of VLMs are surveyed across domains such as embodied agents, robotics, and video generation.
2
Current VLMs face challenges including hallucination, fairness, and safety, which the paper highlights for future work.
3
The paper provides a systematic overview of major VLMs developed from 2019–2024, including model information, architectures, and training methods.
4
The survey summarizes and categorizes popular benchmarks and evaluation metrics used for VLMs.
5
Vision-language models (VLMs) like CLIP, Claude, and GPT-4V show strong visual-textual reasoning and outperform single-modality vision models on zero-shot classification.
Research Object
Multimodal Vision Language Models (VLMs)
Research Subject
Benchmark evaluations, architectures, training methods, applications, and challenges (e.g., benchmarks and metrics, embodied agents/robotics/video generation applications, and issues like hallucination, fairness, safety) of VLMs
Publication Details
Publication Date
2025-01-15
Journal
Publisher
ISSN
Cited by
25
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai11
A Survey on Evaluation of Large Language Models2024
Flamingo: a Visual Language Model for Few-Shot Learning2022
GPT-NeoX-20B: An Open-Source Autoregressive Language Model2022
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints2023
Vision-language models for medical report generation and visual question answering: a review2024
Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving2024
NaVILA: Legged Robot Vision-Language-Action Model for Navigation2024
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Affordance-Compiled Intelligence: Observable-Only Cognitive Impedance Matching for No-Meta LLM-Integrated Systems2020
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations2017
ImageNet: A large-scale hierarchical image database2009