Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Visual Genome: объединение языка и зрения с использованием плотной разметки изображений, созданной краудсорсингом
Li-Jia Li, Li Fei-Fei, Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, David A. Shamma, Michael S. Bernstein
2017-02-06

Visual Genome datasetdense image annotationsimage descriptionobject relationshipsvisual question answering
Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world. However, models used to tackle the rich content in images for cognitive tasks are still being trained using the same datasets designed for perceptual tasks. To achieve success at cognitive tasks, models need to understand the interactions and relationships between objects in an image. When asked “What vehicle is the person riding?”, computers will need to identify the objects in an image as well as the relationships riding(man, carriage) and pulling(horse, carriage) to answer correctly that “the person is riding a horse-drawn carriage.” In this paper, we present the Visual Genome dataset to enable the modeling of such relationships. We collect dense annotations of objects, attributes, and relationships within each image to learn these models. Specifically, our dataset contains over 108K images where each image has an average of $$35$$ objects, $$26$$ attributes, and $$21$$ pairwise relationships between objects. We canonicalize the objects, attributes, relationships, and noun phrases in region descriptions and questions answer pairs to WordNet synsets. Together, these annotations represent the densest and largest dataset of image descriptions, objects, attributes, relationships, and question answer pairs.
1
Annotations include image descriptions, objects, attributes, relationships, and question–answer pairs, providing supervision beyond conventional perceptual datasets.
2
Objects, attributes, relationships, region-description noun phrases, and question–answer pairs are canonicalized to WordNet synsets.
3
The dataset contains over 108K images with dense annotations averaging 35 objects, 26 attributes, and 21 pairwise object relationships per image.
4
The dataset is presented as the largest and densest resource combining these complementary forms of visual-language annotation.
5
Visual Genome is introduced as a dataset designed to support cognitive vision tasks requiring reasoning about interactions and relationships between image objects.

Visual Genome dataset of densely annotated images (objects, attributes, relationships, region descriptions, and question–answer pairs)

Modeling interactions and relationships between objects in images for visual reasoning, image description, and question answering

Publication Details
Publication Date
2017-02-06
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Li-Jia Li
Li Fei-Fei
Ranjay Krishna
Yuke Zhu
Oliver Groth
Justin Johnson
Kenji Hata
Joshua Kravitz
Stephanie Chen
Yannis Kalantidis
David A. Shamma
Michael S. Bernstein
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%