Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model

Повышение точности результатов GPT-3/4 на биомедицинских данных с использованием языковой модели с дополнительным поиском (RAG)
Brandon W. Higgs, David S. Soong, Sriram Sridhar, Han Si, Jan-Samuel Wagner, Ana Caroline C. Sá, Christina Y. Yu, Kübra Karagoz, Meijian Guan, Sanyam Kumar, Hisham K. Hamadeh
2024-08-21

GPT-3.5GPT-4Prometheus (Microsoft)accuracydiffuse large B-cell lymphoma (DLBCL)domain-specific LLMshallucinationsoncology research-focused RAGreadabilityrelevanceretrieval-augmented generation (RAG)
Large language models (LLMs) have made a significant impact on the fields of general artificial intelligence. General purpose LLMs exhibit strong logic and reasoning skills and general world knowledge but can sometimes generate misleading results when prompted on specific subject areas. LLMs trained with domain-specific knowledge can reduce the generation of misleading information (i.e. hallucinations) and enhance the precision of LLMs in specialized contexts. Training new LLMs on specific corpora however can be resource intensive. Here we explored the use of a retrieval-augmented generation (RAG) model which we tested on literature specific to a biomedical research area. OpenAI's GPT-3.5, GPT-4, Microsoft's Prometheus, and a custom RAG model were used to answer 19 questions pertaining to diffuse large B-cell lymphoma (DLBCL) disease biology and treatment. Eight independent reviewers assessed LLM responses based on accuracy, relevance, and readability, rating responses on a 3-point scale for each category. These scores were then used to compare LLM performance. The performance of the LLMs varied across scoring categories. On accuracy and relevance, the RAG model outperformed other models with higher scores on average and the most top scores across questions. GPT-4 was more comparable to the RAG model on relevance versus accuracy. By the same measures, GPT-4 and GPT-3.5 had the highest scores for readability of answers when compared to the other LLMs. GPT-4 and 3.5 also had more answers with hallucinations than the other LLMs, due to non-existent references and inaccurate responses to clinical questions. Our findings suggest that an oncology research-focused RAG model may outperform general-purpose LLMs in accuracy and relevance when answering subject-related questions. This framework can be tailored to Q&A in other subject areas. Further research will help understand the impact of LLM architectures, RAG methodologies, and prompting techniques in answering questions across different subject areas.
1
A retrieval-augmented generation (RAG) model was tested on biomedical literature for DLBCL and outperformed general LLMs on accuracy and relevance.
2
Across 19 DLBCL questions evaluated by eight independent reviewers on a 3-point scale, the RAG model achieved higher average scores and the most top scores across questions.
3
GPT-4 and GPT-3.5 produced the most readable answers but also exhibited more hallucinations, including non-existent references and inaccurate clinical responses.
4
GPT-4 matched the RAG model more closely on relevance than on accuracy, indicating comparative strength in relevance but weaker accuracy.
5
The oncology research–focused RAG framework can be adapted to other subject areas, suggesting domain-specific retrieval augmentation improves LLM precision without retraining full models.

Retrieval-augmented generation (RAG) model evaluated on biomedical literature about diffuse large B-cell lymphoma (DLBCL)

Accuracy, relevance, and readability of answers produced by the RAG model compared to general-purpose LLMs (GPT-3.5, GPT-4, Prometheus) when answering DLBCL biology and treatment questions, including incidence of hallucinations

Publication Details
Publication Date
2024-08-21
Journal
Publisher
ISSN
Cited by
49
Access Type
Author Information
Authors
Brandon W. Higgs
David S. Soong
Sriram Sridhar
Han Si
Jan-Samuel Wagner
Ana Caroline C. Sá
Christina Y. Yu
Kübra Karagoz
Meijian Guan
Sanyam Kumar
Hisham K. Hamadeh
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%