REALM: Retrieval-Augmented Language Model Pre-Training

REALM: предварительное обучение языковой модели с дополнением за счет извлечения информации
Ming‐Wei Chang, Kelvin Guu, Zora Tung, Kenton Lee, Panupong Pasupat
2020-02-10

REALMlatent knowledge retrievermasked language modelingopen-domain question answeringretrieval-augmented language model pre-training
Language model pre-training has been shown to capture a surprising amount of world knowledge, crucial for NLP tasks such as question answering. However, this knowledge is stored implicitly in the parameters of a neural network, requiring ever-larger networks to cover more facts. To capture knowledge in a more modular and interpretable way, we augment language model pre-training with a latent knowledge retriever, which allows the model to retrieve and attend over documents from a large corpus such as Wikipedia, used during pre-training, fine-tuning and inference. For the first time, we show how to pre-train such a knowledge retriever in an unsupervised manner, using masked language modeling as the learning signal and backpropagating through a retrieval step that considers millions of documents. We demonstrate the effectiveness of Retrieval-Augmented Language Model pre-training (REALM) by fine-tuning on the challenging task of Open-domain Question Answering (Open-QA). We compare against state-of-the-art models for both explicit and implicit knowledge storage on three popular Open-QA benchmarks, and find that we outperform all previous methods by a significant margin (4-16% absolute accuracy), while also providing qualitative benefits such as interpretability and modularity.
1
Fine-tuning REALM for open-domain question answering outperforms prior explicit- and implicit-knowledge models by 4–16% absolute accuracy across three benchmarks.
2
REALM augments language-model pre-training with a latent knowledge retriever that retrieves and attends over documents from large corpora such as Wikipedia.
3
Retrieval-based knowledge storage provides qualitative benefits including greater interpretability and modularity compared with knowledge encoded solely in model parameters.
4
The retriever can be pre-trained unsupervised using masked language modeling, with backpropagation through retrieval over millions of documents.

Retrieval-Augmented Language Model pre-training (REALM) with a latent knowledge retriever over a large document corpus

The effectiveness, interpretability, modularity, and knowledge-retrieval behavior of REALM for open-domain question answering

Publication Details
Publication Date
2020-02-10
Journal
Publisher
ISSN
Cited by
521
Access Type
Author Information
Authors
Ming‐Wei Chang
Kelvin Guu
Zora Tung
Kenton Lee
Panupong Pasupat
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%