Flamingo: a Visual Language Model for Few-Shot Learning

Flamingo: визуальная языковая модель для обучения с небольшим числом примеров
Oriol Vinyals, Andrew Zisserman, Jacob Menick, Serkan Cabi, Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bińkowski, Ricardo Barreira, Karen Simonyan, Alayrac, Jean-Baptiste
2022-04-29

FlamingoVisual Language Modelfew-shot learningmultimodal web corporavisual question-answering
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose key architectural innovations to: (i) bridge powerful pretrained vision-only and language-only models, (ii) handle sequences of arbitrarily interleaved visual and textual data, and (iii) seamlessly ingest images or videos as inputs. Thanks to their flexibility, Flamingo models can be trained on large-scale multimodal web corpora containing arbitrarily interleaved text and images, which is key to endow them with in-context few-shot learning capabilities. We perform a thorough evaluation of our models, exploring and measuring their ability to rapidly adapt to a variety of image and video tasks. These include open-ended tasks such as visual question-answering, where the model is prompted with a question which it has to answer; captioning tasks, which evaluate the ability to describe a scene or an event; and close-ended tasks such as multiple-choice visual question-answering. For tasks lying anywhere on this spectrum, a single Flamingo model can achieve a new state of the art with few-shot learning, simply by prompting the model with task-specific examples. On numerous benchmarks, Flamingo outperforms models fine-tuned on thousands of times more task-specific data.
1
A single Flamingo model achieves new state-of-the-art few-shot performance across diverse tasks (open-ended VQA, captioning, and multiple-choice VQA) by prompting with task-specific examples.
2
Architectural innovations enable Flamingo to bridge pretrained vision-only and language-only models and handle arbitrarily interleaved visual and textual sequences.
3
Flamingo can seamlessly ingest both images and videos as inputs, allowing training on large-scale multimodal web corpora with interleaved text and images.
4
Flamingo is a family of Visual Language Models designed for in-context few-shot learning on multimodal tasks.
5
Flamingo outperforms models that were fine-tuned on thousands of times more task-specific data on numerous benchmarks.

Flamingo family of Visual Language Models (VLMs)

Few-shot in-context learning ability of Flamingo to rapidly adapt to novel multimodal tasks (image and video) via handling interleaved visual and textual inputs and achieving state-of-the-art performance on open-ended and close-ended vision-language tasks

Publication Details
Publication Date
2022-04-29
Journal
Publisher
ISSN
Cited by
1251
Access Type
Author Information
Authors
Oriol Vinyals
Andrew Zisserman
Jacob Menick
Serkan Cabi
Jean-Baptiste Alayrac
Jeff Donahue
Pauline Luc
Antoine Miech
Iain Barr
Yana Hasson
Karel Lenc
Arthur Mensch
Katie Millican
Malcolm Reynolds
Roman Ring
Eliza Rutherford
Tengda Han
Zhitao Gong
Sina Samangooei
Marianne Monteiro
Sebastian Borgeaud
Andrew Brock
Aida Nematzadeh
Sahand Sharifzadeh
Mikołaj Bińkowski
Ricardo Barreira
Karen Simonyan
Alayrac, Jean-Baptiste
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%