Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs

Промпт-инжиниринг для обеспечения согласованности и надежности ответов больших языковых моделей в соответствии с доказательными рекомендациями
Jian Li, Xi Chen, Xiangwen Deng, Wang Li, Hao Wen, Mingke You, Weizhi Liu, Qi Li
2024-02-20

Consistency and reliabilityFleiss kappaLarge language modelsOsteoarthritis guidelinesPrompt engineering
The use of large language models (LLMs) in clinical medicine is currently thriving. Effectively transferring LLMs' pertinent theoretical knowledge from computer science to their application in clinical medicine is crucial. Prompt engineering has shown potential as an effective method in this regard. To explore the application of prompt engineering in LLMs and to examine the reliability of LLMs, different styles of prompts were designed and used to ask different LLMs about their agreement with the American Academy of Orthopedic Surgeons (AAOS) osteoarthritis (OA) evidence-based guidelines. Each question was asked 5 times. We compared the consistency of the findings with guidelines across different evidence levels for different prompts and assessed the reliability of different prompts by asking the same question 5 times. gpt-4-Web with ROT prompting had the highest overall consistency (62.9%) and a significant performance for strong recommendations, with a total consistency of 77.5%. The reliability of the different LLMs for different prompts was not stable (Fleiss kappa ranged from -0.002 to 0.984). This study revealed that different prompts had variable effects across various models, and the gpt-4-Web with ROT prompt was the most consistent. An appropriate prompt could improve the accuracy of responses to professional medical questions.
1
Each clinical question was repeated five times to assess both guideline consistency and response reliability across models and prompts.
2
GPT-4-Web using ROT prompting achieved the highest overall guideline consistency at 62.9%.
3
GPT-4-Web with ROT prompting showed 77.5% consistency for strong recommendations.
4
LLM reliability varied substantially by model and prompt, with Fleiss kappa values ranging from -0.002 to 0.984; prompt effects were not stable across models.
5
The study evaluated how prompt styles affect LLM agreement with AAOS evidence-based osteoarthritis guidelines across different recommendation-strength levels.

Prompt engineering applied to large language models (LLMs) for querying AAOS osteoarthritis evidence-based guidelines

the consistency and reliability of LLM responses across prompt styles, models, and repeated queries, including agreement with guideline evidence levels

Publication Details
Publication Date
2024-02-20
Journal
Publisher
ISSN
Cited by
378
Access Type
Author Information
Authors
Jian Li
Xi Chen
Xiangwen Deng
Wang Li
Hao Wen
Mingke You
Weizhi Liu
Qi Li
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%