A Unified Approach to Interpreting Model Predictions

Единый подход к интерпретации предсказаний моделей
Scott Lundberg, Su‐In Lee
2017-05-22

SHAPShapley additive explanationsadditive feature attributionfeature importancemodel interpretability
Understanding why a model makes a certain prediction can be as crucial as the prediction's accuracy in many applications. However, the highest accuracy for large modern datasets is often achieved by complex models that even experts struggle to interpret, such as ensemble or deep learning models, creating a tension between accuracy and interpretability. In response, various methods have recently been proposed to help users interpret the predictions of complex models, but it is often unclear how these methods are related and when one method is preferable over another. To address this problem, we present a unified framework for interpreting predictions, SHAP (SHapley Additive exPlanations). SHAP assigns each feature an importance value for a particular prediction. Its novel components include: (1) the identification of a new class of additive feature importance measures, and (2) theoretical results showing there is a unique solution in this class with a set of desirable properties. The new class unifies six existing methods, notable because several recent methods in the class lack the proposed desirable properties. Based on insights from this unification, we present new methods that show improved computational performance and/or better consistency with human intuition than previous approaches.
1
Insights from the unification yield new methods with improved computational performance and/or greater consistency with human intuition.
2
SHAP provides a unified framework for interpreting individual predictions by assigning each feature an importance value.
3
The framework identifies a new class of additive feature-importance measures for explaining model predictions.
4
The framework unifies six existing prediction-interpretation methods, several of which lack the proposed desirable properties.
5
Within this class, SHAP is theoretically shown to be the unique solution satisfying a specified set of desirable properties.

Predictions of complex machine-learning models

Feature-importance attribution for individual predictions, including the unification, axiomatic characterization, and computational improvement of additive explanation methods

Publication Details
Publication Date
2017-05-22
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Scott Lundberg
Su‐In Lee
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%