Towards Faithful Model Explanation in NLP: A Survey

На пути к достоверному объяснению моделей в обработке естественного языка: обзор
Chris Callison-Burch, Qing Lyu, Marianna Apidianaki
2024-01-01

counterfactual interventionexplainabilityfaithful model explanationnatural language processingself-explanatory models
Abstract End-to-end neural Natural Language Processing (NLP) models are notoriously difficult to understand. This has given rise to numerous efforts towards model explainability in recent years. One desideratum of model explanation is faithfulness, that is, an explanation should accurately represent the reasoning process behind the model’s prediction. In this survey, we review over 110 model explanation methods in NLP through the lens of faithfulness. We first discuss the definition and evaluation of faithfulness, as well as its significance for explainability. We then introduce recent advances in faithful explanation, grouping existing approaches into five categories: similarity-based methods, analysis of model-internal structures, backpropagation-based methods, counterfactual intervention, and self-explanatory models. For each category, we synthesize its representative studies, strengths, and weaknesses. Finally, we summarize their common virtues and remaining challenges, and reflect on future work directions towards faithful explainability in NLP.
1
Existing faithful explanation approaches are organized into five categories: similarity-based methods, model-internal structure analysis, backpropagation-based methods, counterfactual intervention, and self-explanatory models.
2
It examines how faithfulness is defined and evaluated, emphasizing faithfulness as a central desideratum for reliable NLP explainability.
3
It outlines future research directions for developing more faithful explainability methods in NLP.
4
The survey reviews over 110 NLP model explanation methods using faithfulness—the extent to which explanations accurately reflect models’ prediction reasoning—as its organizing criterion.
5
The survey synthesizes the representative studies, strengths, and weaknesses of each category, identifying shared advantages and unresolved challenges.

model explanation methods in NLP models

faithfulness of explanations, including how accurately they represent the model’s reasoning process behind predictions

Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
97
Access Type
Author Information
Authors
Chris Callison-Burch
Qing Lyu
Marianna Apidianaki
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%