Beware of “Explanations” of AI

Остерегайтесь «объяснений» ИИ
Stefan Feuerriegel, Galit Shmueli, Wouter Verbeke, Niklas Kühl, Theodoros Evgeniou, David Martens, Claudia Perlich, Foster Provost, Christian Janiesch, Mathias Kraus, Patrick Zschech, Kevin Bauer, Sebastian Gabel, Sofie Goethals, Travis Greene, Nadja Klein, Alona Zharova
2026-08-25

AI explanationsexplainable artificial intelligenceexplanation qualitymental model alignmentresponsible AI adoption
Abstract Understanding the decisions made and actions taken by increasingly complex AI systems remains a key challenge. This has led to an expanding field of research in explainable artificial intelligence (XAI), highlighting the potential of explanations to enhance trust, support adoption, and meet regulatory standards. However, the question of what constitutes a “good” explanation is dependent on the goals, stakeholders, and context. At a high level, psychological insights such as the concept of mental model alignment can offer guidance, but success in practice is challenging due to social and technical factors. As a result of this ill-defined nature of the problem, explanations can be of poor quality (e.g., unfaithful, irrelevant, or incoherent), potentially leading to substantial risks. Instead of fostering trust and safety, poorly designed explanations can actually cause harm due to wrong decisions, privacy violations, manipulation, and reduced AI adoption. Therefore, we caution stakeholders to beware of explanations of AI: While they can be vital, they are not automatically a remedy for transparency or responsible AI adoption, and their misuse or limitations can exacerbate harm. Attention to these caveats can help guide future research to improve the quality and impact of AI explanations.
1
AI explanations may be unfaithful, irrelevant, or incoherent because explanation quality is difficult to define and evaluate.
2
Explanations can support transparency and responsible AI adoption, but they are not automatically effective remedies and require careful attention to limitations.
3
Mental model alignment offers psychological guidance for explanation design, but social and technical factors make practical success difficult.
4
Poorly designed explanations can cause wrong decisions, privacy violations, manipulation, and reduced AI adoption rather than improving trust and safety.
5
What constitutes a good AI explanation depends on its goals, stakeholders, and contextual use.

Explanations of decisions and actions produced by complex AI systems

The quality, effectiveness, limitations, and potential harms of AI explanations in relation to trust, transparency, safety, adoption, regulation, and stakeholder goals

Publication Details
Publication Date
2026-08-25
Journal
Publisher
ISSN
Cited by
1
Access Type
Author Information
Authors
Stefan Feuerriegel
Galit Shmueli
Wouter Verbeke
Niklas Kühl
Theodoros Evgeniou
David Martens
Claudia Perlich
Foster Provost
Christian Janiesch
Mathias Kraus
Patrick Zschech
Kevin Bauer
Sebastian Gabel
Sofie Goethals
Travis Greene
Nadja Klein
Alona Zharova
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%