Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI

Базовые метрики для оценки эффективности медицинских диалогов на основе генеративного искусственного интеллекта
Li-Jia Li, Olivier Gevaert, Yanshan Wang, Zhongqi Yang, Elahe Khatibi, Mahyar Abbasian, Iman Azimi, Ramesh Jain, David Oniani, Zahra Shakeri Hossein Abad, Alexander Thieme, Ram D. Sriram, Bryant Lin, Amir M. Rahmani
2024-03-29

conversational model evaluationgenerative AIhealthcare chatbotslarge language model metricspatient-centered evaluation
Generative Artificial Intelligence is set to revolutionize healthcare delivery by transforming traditional patient care into a more personalized, efficient, and proactive process. Chatbots, serving as interactive conversational models, will probably drive this patient-centered transformation in healthcare. Through the provision of various services, including diagnosis, personalized lifestyle recommendations, dynamic scheduling of follow-ups, and mental health support, the objective is to substantially augment patient health outcomes, all the while mitigating the workload burden on healthcare providers. The life-critical nature of healthcare applications necessitates establishing a unified and comprehensive set of evaluation metrics for conversational models. Existing evaluation metrics proposed for various generic large language models (LLMs) demonstrate a lack of comprehension regarding medical and health concepts and their significance in promoting patients' well-being. Moreover, these metrics neglect pivotal user-centered aspects, including trust-building, ethics, personalization, empathy, user comprehension, and emotional support. The purpose of this paper is to explore state-of-the-art LLM-based evaluation metrics that are specifically applicable to the assessment of interactive conversational models in healthcare. Subsequently, we present a comprehensive set of evaluation metrics designed to thoroughly assess the performance of healthcare chatbots from an end-user perspective. These metrics encompass an evaluation of language processing abilities, impact on real-world clinical tasks, and effectiveness in user-interactive conversations. Finally, we engage in a discussion concerning the challenges associated with defining and implementing these metrics, with particular emphasis on confounding factors such as the target audience, evaluation methods, and prompt techniques involved in the evaluation process.
1
Existing metrics overlook user-centered requirements crucial for healthcare chatbots, including trust, ethics, personalization, empathy, comprehension, and emotional support.
2
Generic LLM evaluation metrics inadequately capture medical concepts and their relevance to patient well-being in healthcare conversations.
3
Implementing healthcare conversation metrics remains challenging because results are influenced by target audiences, evaluation methods, and prompt techniques.
4
The framework is intended to support unified evaluation of healthcare chatbots across services such as diagnosis, lifestyle recommendations, follow-up scheduling, and mental health support.
5
The paper proposes a comprehensive end-user-oriented metric framework covering language processing, real-world clinical task impact, and interactive conversational effectiveness.

LLM-based healthcare chatbots and their interactive conversational models

Comprehensive, end-user-oriented metrics for evaluating their language-processing abilities, clinical-task impact, and effectiveness in healthcare conversations, including trust, ethics, personalization, empathy, comprehension, and emotional support

Publication Details
Publication Date
2024-03-29
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Li-Jia Li
Olivier Gevaert
Yanshan Wang
Zhongqi Yang
Elahe Khatibi
Mahyar Abbasian
Iman Azimi
Ramesh Jain
David Oniani
Zahra Shakeri Hossein Abad
Alexander Thieme
Ram D. Sriram
Bryant Lin
Amir M. Rahmani
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%