Machine Learning Interpretability: A Survey on Methods and Metrics

Интерпретируемость машинного обучения: обзор методов и метрик
Jaime S. Cardoso, Diogo V. Carvalho, Eduardo M. Pereira
2019-07-26

black-box systemsexplanation methodsexplanation quality metricsinterpretable modelsmachine learning interpretability
Machine learning systems are becoming increasingly ubiquitous. These systems’s adoption has been expanding, accelerating the shift towards a more algorithmic society, meaning that algorithmically informed decisions have greater potential for significant social impact. However, most of these accurate decision support systems remain complex black boxes, meaning their internal logic and inner workings are hidden to the user and even experts cannot fully understand the rationale behind their predictions. Moreover, new regulations and highly regulated domains have made the audit and verifiability of decisions mandatory, increasing the demand for the ability to question, understand, and trust machine learning systems, for which interpretability is indispensable. The research community has recognized this interpretability problem and focused on developing both interpretable models and explanation methods over the past few years. However, the emergence of these methods shows there is no consensus on how to assess the explanation quality. Which are the most suitable metrics to assess the quality of an explanation? The aim of this article is to provide a review of the current state of the research field on machine learning interpretability while focusing on the societal impact and on the developed methods and metrics. Furthermore, a complete literature review is presented in order to identify future directions of work on this field.
1
Interpretability is becoming indispensable because regulations and highly regulated applications require users to question, understand, and trust algorithmic decisions.
2
Machine learning systems increasingly influence socially significant decisions, creating an urgent need for transparency, auditability, and verifiability.
3
Many high-performing decision-support systems remain black boxes whose prediction rationale is inaccessible to users and incompletely understood even by experts.
4
The article surveys interpretability methods and evaluation metrics, examines their societal impact, and identifies directions for future research through a comprehensive literature review.
5
The interpretability field has developed both inherently interpretable models and post-hoc explanation methods, but lacks consensus on how explanation quality should be assessed.

Machine learning models and explanation methods for decision-support systems

interpretability of machine learning systems, including methods for generating explanations and metrics for assessing explanation quality

Publication Details
Publication Date
2019-07-26
Journal
Publisher
ISSN
Cited by
1831
Access Type
Author Information
Authors
Jaime S. Cardoso
Diogo V. Carvalho
Eduardo M. Pereira
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%