Assessing Interaction Quality in Human–AI Dialogue: An Integrative Review and Multi-Layer Framework for Conversational Agents

Оценка качества взаимодействия в диалоге человека и ИИ: интегративный обзор и многоуровневая модель разговорных агентов
Luca Marconi, Luca Longo, Federico Cabitza
2026-01-26

conversational agentshuman–AI dialogueinteraction qualitylarge language modelsuser-centred evaluation
Conversational agents are transforming digital interactions across various domains, including healthcare, education, and customer service, thanks to advances in large language models (LLMs). As these systems become more autonomous and ubiquitous, understanding what constitutes high-quality interaction from a user perspective is increasingly critical. Despite growing empirical research, the field lacks a unified framework for defining, measuring, and designing user-perceived interaction quality in human–artificial intelligence (AI) dialogue. Here, we present an integrative review of 125 empirical studies published between 2017 and 2025, spanning text-, voice-, and LLM-powered systems. Our synthesis identifies three consistent layers of user judgment: a pragmatic core (usability, task effectiveness, and conversational competence), a social–affective layer (social presence, warmth, and synchronicity), and an accountability and inclusion layer (transparency, accessibility, and fairness). These insights are formalised into a four-layer interpretive framework—Capacity, Alignment, Levers, and Outcomes—operationalised via a Capacity × Alignment matrix that maps distinct success and failure regimes. It also identifies design levers such as anthropomorphism, role framing, and onboarding strategies. The framework consolidates constructs, positions inclusion and accountability as central to quality, and offers actionable guidance for evaluation and design. This research redefines interaction quality as a dialogic construct, shifting the focus from system performance to co-orchestrated, user-centred dialogue quality.
1
An integrative review of 125 empirical studies from 2017–2025 synthesizes interaction-quality evidence across text-, voice-, and LLM-powered conversational systems.
2
Anthropomorphism, role framing, and onboarding are identified as design levers that can shape perceived conversational quality.
3
The framework redefines interaction quality as a dialogic, co-orchestrated, user-centred construct rather than solely a measure of system performance.
4
The proposed four-layer framework—Capacity, Alignment, Levers, and Outcomes—organizes interaction-quality constructs and is operationalized through a Capacity × Alignment matrix of success and failure regimes.
5
User judgments consistently comprise pragmatic, social–affective, and accountability-and-inclusion layers, covering effectiveness, warmth, transparency, accessibility, and fairness.

human–AI dialogue involving conversational agents, including text-, voice-, and LLM-powered systems

user-perceived interaction quality, including pragmatic effectiveness, social–affective experience, transparency, accessibility, and fairness

Publication Details
Publication Date
2026-01-26
Journal
Publisher
ISSN
Cited by
13
Access Type
Author Information
Authors
Luca Marconi
Luca Longo
Federico Cabitza
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%