From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI

От анекдотических свидетельств к количественным методам оценки: систематический обзор оценки объяснимого искусственного интеллекта
Maurice van Keulen, Meike Nauta, Jan Trienes, Shreyasi Pathak, Elisa Nguyen, Michelle Peters, Yasmin Schmitt, Jörg Schlötterer, Christin Seifert
2023-02-24

Co-12 propertiesXAI evaluationexplainable artificial intelligencequantitative evaluation methodssystematic review
The rising popularity of explainable artificial intelligence (XAI) to understand high-performing black boxes raised the question of how to evaluate explanations of machine learning (ML) models. While interpretability and explainability are often presented as a subjectively validated binary property, we consider it a multi-faceted concept. We identify 12 conceptual properties, such as Compactness and Correctness, that should be evaluated for comprehensively assessing the quality of an explanation. Our so-called Co-12 properties serve as categorization scheme for systematically reviewing the evaluation practices of more than 300 papers published in the past 7 years at major AI and ML conferences that introduce an XAI method. We find that one in three papers evaluate exclusively with anecdotal evidence, and one in five papers evaluate with users. This survey also contributes to the call for objective, quantifiable evaluation methods by presenting an extensive overview of quantitative XAI evaluation methods. Our systematic collection of evaluation methods provides researchers and practitioners with concrete tools to thoroughly validate, benchmark, and compare new and existing XAI methods. The Co-12 categorization scheme and our identified evaluation methods open up opportunities to include quantitative metrics as optimization criteria during model training to optimize for accuracy and interpretability simultaneously.
1
One-third of reviewed papers rely exclusively on anecdotal evidence, while only one-fifth evaluate explanations with users.
2
The Co-12 framework categorizes evaluation practices across more than 300 XAI papers published at major AI and machine learning conferences over seven years.
3
The Co-12 scheme and quantitative metrics could enable explanation properties to become optimization criteria for jointly improving model accuracy and interpretability.
4
The review conceptualizes explanation quality as multifaceted and identifies 12 properties, including Compactness and Correctness, for comprehensive XAI evaluation.
5
The review systematically compiles quantitative XAI evaluation methods that can support validation, benchmarking, and comparison of explanation techniques.

evaluation practices and methods for explanations of machine learning models in explainable artificial intelligence (XAI)

the multifaceted quality of XAI explanations, including their quantitative evaluation across properties such as Compactness and Correctness

Publication Details
Publication Date
2023-02-24
Journal
Publisher
ISSN
Cited by
535
Access Type
Author Information
Authors
Maurice van Keulen
Meike Nauta
Jan Trienes
Shreyasi Pathak
Elisa Nguyen
Michelle Peters
Yasmin Schmitt
Jörg Schlötterer
Christin Seifert
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%