Machine Learning Interpretability: A Survey on Methods and Metrics
Интерпретируемость машинного обучения: обзор методов и метрик
2019-07-26
SCID: 54.1/23kgj24j
Discuss with AI
black-box systemsexplanation methodsexplanation quality metricsinterpretable modelsmachine learning interpretability
Figures from the paper
Abstract (AI)
Machine learning systems are becoming increasingly ubiquitous. These systems’s adoption has been expanding, accelerating the shift towards a more algorithmic society, meaning that algorithmically informed decisions have greater potential for significant social impact. However, most of these accurate decision support systems remain complex black boxes, meaning their internal logic and inner workings are hidden to the user and even experts cannot fully understand the rationale behind their predictions. Moreover, new regulations and highly regulated domains have made the audit and verifiability of decisions mandatory, increasing the demand for the ability to question, understand, and trust machine learning systems, for which interpretability is indispensable. The research community has recognized this interpretability problem and focused on developing both interpretable models and explanation methods over the past few years. However, the emergence of these methods shows there is no consensus on how to assess the explanation quality. Which are the most suitable metrics to assess the quality of an explanation? The aim of this article is to provide a review of the current state of the research field on machine learning interpretability while focusing on the societal impact and on the developed methods and metrics. Furthermore, a complete literature review is presented in order to identify future directions of work on this field.
Key Findings
1
Interpretability is becoming indispensable because regulations and highly regulated applications require users to question, understand, and trust algorithmic decisions.
2
Machine learning systems increasingly influence socially significant decisions, creating an urgent need for transparency, auditability, and verifiability.
3
Many high-performing decision-support systems remain black boxes whose prediction rationale is inaccessible to users and incompletely understood even by experts.
4
The article surveys interpretability methods and evaluation metrics, examines their societal impact, and identifies directions for future research through a comprehensive literature review.
5
The interpretability field has developed both inherently interpretable models and post-hoc explanation methods, but lacks consensus on how explanation quality should be assessed.
Research Object
Machine learning models and explanation methods for decision-support systems
Research Subject
interpretability of machine learning systems, including methods for generating explanations and metrics for assessing explanation quality
Publication Details
Publication Date
2019-07-26
Journal
Publisher
ISSN
Cited by
1831
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai7
Greedy function approximation: A gradient boosting machine.2001
"Why Should I Trust You?"2016
Distilling the Knowledge in a Neural Network2015
A Unified Approach to Interpreting Model Predictions2017
Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)2018
Towards A Rigorous Science of Interpretable Machine Learning2017
Causability and explainability of artificial intelligence in medicine2019
Cited by17
Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence2023
The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic review2024
Connecting the dots in trustworthy Artificial Intelligence: From AI principles, ethics, and key requirements to responsible AI systems and regulation2023
Evaluating the Quality of Machine Learning Explanations: A Survey on Methods and Metrics2021
Explainable Artificial Intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions2024
From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI2023
Transparency of deep neural networks for medical image analysis: A review of interpretability methods2021
Trust in AI: progress, challenges, and future directions2024
Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research2022
Interpretability of machine learning‐based prediction models in healthcare2020
Artificial Intelligence for Predictive Maintenance Applications: Key Components, Trustworthiness, and Future Trends2024
Explainable Artificial Intelligence in CyberSecurity: A Survey2022
Recent Applications of Explainable AI (XAI): A Systematic Literature Review2024
Exploring collaborative decision-making: A quasi-experimental study of human and Generative AI interaction2024
A Survey on Explainable Anomaly Detection2023
Explainable Artificial Intelligence (XAI) in Insurance2022
Interpretable machine learning for real estate market analysis2022