Flamingo: a Visual Language Model for Few-Shot Learning
Flamingo: визуальная языковая модель для обучения с небольшим числом примеров
2022-04-29
SCID: 54.1/axxm8y2q
Discuss with AI
FlamingoVisual Language Modelfew-shot learningmultimodal web corporavisual question-answering
Figures from the paper
Abstract (AI)
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer; captioning tasks, which evaluate the ability to describe a scene or an event; and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
Key Findings
1
A single Flamingo model achieves new state-of-the-art few-shot performance across diverse tasks (open-ended VQA, captioning, and multiple-choice VQA) by prompting with task-specific examples.
2
Architectural innovations enable Flamingo to bridge pretrained vision-only and language-only models and handle arbitrarily interleaved visual and textual sequences.
3
Flamingo can seamlessly ingest both images and videos as inputs, allowing training on large-scale multimodal web corpora with interleaved text and images.
4
Flamingo is a family of Visual Language Models designed for in-context few-shot learning on multimodal tasks.
5
Flamingo outperforms models that were fine-tuned on thousands of times more task-specific data on numerous benchmarks.
Research Object
Flamingo family of Visual Language Models (VLMs)
Research Subject
Few-shot in-context learning ability of Flamingo to rapidly adapt to novel multimodal tasks (image and video) via handling interleaved visual and textual inputs and achieving state-of-the-art performance on open-ended and close-ended vision-language tasks
Publication Details
Publication Date
2022-04-29
Journal
Publisher
ISSN
Cited by
1251
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by10
Multimodal Learning With Transformers: A Survey2023
A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges2024
Talking about Large Language Models2024
A Comprehensive Review of Deep Learning: Architectures, Recent Advances, and Applications2024
Vision-language models for medical report generation and visual question answering: a review2024
Review of large vision models and visual prompt engineering2023
Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving2024
Interactive computer-aided diagnosis on medical image using large language models2024
Alpha-CLIP: A CLIP Model Focusing on Wherever you Want2024
Benchmark Evaluations, Applications, and Challenges of Large Vision Language Models: A Survey2025