Visual Dialog

Визуальный диалог
Dhruv Batra, Devi Parikh, Avi Singh, Abhishek Das, Satwik Kottur, Khushi Gupta, Deshraj Jain, José M. F. Moura, Deshraj Yadav
2017-07-01

Hierarchical Recurrent EncoderVisual DialogVisual Dialog dataset (VisDial)image-grounded dialogueretrieval-based evaluation (mean reciprocal rank)
We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the image, the agent has to ground the question in image, infer context from history, and answer the question accurately. Visual Dialog is disentangled enough from a specific downstream task so as to serve as a general test of machine intelligence, while being grounded in vision enough to allow objective evaluation of individual responses and benchmark progress. We develop a novel two-person chat data-collection protocol to curate a large-scale Visual Dialog dataset (VisDial). VisDial contains 1 dialog (10 question-answer pairs) on ~140k images from the COCO dataset, with a total of ~1.4M dialog question-answer pairs. We introduce a family of neural encoder-decoder models for Visual Dialog with 3 encoders (Late Fusion, Hierarchical Recurrent Encoder and Memory Network) and 2 decoders (generative and discriminative), which outperform a number of sophisticated baselines. We propose a retrieval-based evaluation protocol for Visual Dialog where the AI agent is asked to sort a set of candidate answers and evaluated on metrics such as mean-reciprocal-rank of human response. We quantify gap between machine and human performance on the Visual Dialog task via human studies. Our dataset, code, and trained models will be released publicly at https://visualdialog.org. Putting it all together, we demonstrate the first visual chatbot!.
1
Collected VisDial dataset: ~140k COCO images each with 1 dialog of 10 Q-A pairs, totaling ~1.4M question-answer pairs via a two-person chat protocol.
2
Designed a retrieval-based evaluation protocol using candidate-answer ranking and metrics like mean-reciprocal-rank to objectively evaluate Visual Dialog responses.
3
Introduced Visual Dialog, a task requiring an AI to hold multi-turn natural language dialog grounded in an image and dialog history.
4
Proposed neural encoder-decoder models with three encoders (Late Fusion, Hierarchical Recurrent Encoder, Memory Network) and two decoders (generative, discriminative) that outperform several strong baselines.
5
Quantified the gap between machine and human performance on Visual Dialog through human studies and released dataset, code, and trained models publicly.

Visual Dialog task and its VisDial dataset (dialogues of question-answer pairs grounded in images)

An AI agent's ability to understand visual content and dialog context to generate or select accurate, contextually grounded answers to questions about images (evaluation of models and gap to human performance)

Publication Details
Publication Date
2017-07-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Dhruv Batra
Devi Parikh
Avi Singh
Abhishek Das
Satwik Kottur
Khushi Gupta
Deshraj Jain
José M. F. Moura
Deshraj Yadav
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%