balanced accuracyinstruction fine-tuninglarge language modelsmental health predictionprompting
Figures from the paper
Abstract (AI)
Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In this work, we present a comprehensive evaluation of multiple LLMs on various mental health prediction tasks via online text data, including Alpaca, Alpaca-LoRA, FLAN-T5, GPT-3.5, and GPT-4. We conduct a broad range of experiments, covering zero-shot prompting, few-shot prompting, and instruction fine-tuning. The results indicate a promising yet limited performance of LLMs with zero-shot and few-shot prompt designs for mental health tasks. More importantly, our experiments show that instruction finetuning can significantly boost the performance of LLMs for all tasks simultaneously. Our best-finetuned models, Mental-Alpaca and Mental-FLAN-T5, outperform the best prompt design of GPT-3.5 (25 and 15 times bigger) by 10.9% on balanced accuracy and the best of GPT-4 (250 and 150 times bigger) by 4.8%. They further perform on par with the state-of-the-art task-specific language model. We also conduct an exploratory case study on LLMs' capability on mental health reasoning tasks, illustrating the promising capability of certain models such as GPT-4. We summarize our findings into a set of action guidelines for potential methods to enhance LLMs' capability for mental health tasks. Meanwhile, we also emphasize the important limitations before achieving deployability in real-world mental health settings, such as known racial and gender bias. We highlight the important ethical risks accompanying this line of research.
Key Findings
1
Instruction fine-tuning substantially improves performance across all evaluated mental-health tasks simultaneously.
2
Mental-Alpaca and Mental-FLAN-T5 outperform GPT-3.5 prompt designs by 10.9% and GPT-4 prompt designs by 4.8% in balanced accuracy, despite being much smaller.
3
The best fine-tuned models perform on par with state-of-the-art task-specific language models, while GPT-4 shows promising mental-health reasoning capabilities; deployment remains limited by racial and gender bias and ethical risks.
4
The study evaluates Alpaca, Alpaca-LoRA, FLAN-T5, GPT-3.5, and GPT-4 on diverse mental-health prediction tasks using online text data.
5
Zero-shot and few-shot prompting provide promising but limited performance on mental-health tasks across evaluated large language models.
Research Object
large language models applied to mental health prediction using online text data
Research Subject
their performance, enhancement through prompting and instruction fine-tuning, reasoning capability, and limitations including bias and ethical risks
Publication Details
Publication Date
2024-03-06
Journal
Publisher
ISSN
Cited by
214
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest