The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review

Возникновение больших языковых моделей как инструментов для подготовки обзоров литературы: систематический обзор с использованием большой языковой модели
Alexander Bakumenko, Nina Hubig, Leslie Lenert, Dmitry Scherbakov, Vinita Jansari
2025-04-12

GPT-based modelsdata extractionlarge language modelsreview automationsystematic review
OBJECTIVES: This study aims to summarize the usage of large language models (LLMs) in the process of creating a scientific review by looking at the methodological papers that describe the use of LLMs in review automation and the review papers that mention they were made with the support of LLMs. MATERIALS AND METHODS: The search was conducted in June 2024 in PubMed, Scopus, Dimensions, and Google Scholar by human reviewers. Screening and extraction process took place in Covidence with the help of LLM add-on based on the OpenAI GPT-4o model. ChatGPT and Scite.ai were used in cleaning the data, generating the code for figures, and drafting the manuscript. RESULTS: Of the 3788 articles retrieved, 172 studies were deemed eligible for the final review. ChatGPT and GPT-based LLM emerged as the most dominant architecture for review automation (n = 126, 73.2%). A significant number of review automation projects were found, but only a limited number of papers (n = 26, 15.1%) were actual reviews that acknowledged LLM usage. Most citations focused on the automation of a particular stage of review, such as Searching for publications (n = 60, 34.9%) and Data extraction (n = 54, 31.4%). When comparing the pooled performance of GPT-based and BERT-based models, the former was better in data extraction with a mean precision of 83.0% (SD = 10.4) and a recall of 86.0% (SD = 9.8). DISCUSSION AND CONCLUSION: Our LLM-assisted systematic review revealed a significant number of research projects related to review automation using LLMs. Despite limitations, such as lower accuracy of extraction for numeric data, we anticipate that LLMs will soon change the way scientific reviews are conducted.
1
A systematic search identified 172 eligible studies addressing LLM use in review automation or LLM-supported scientific reviews from 3,788 retrieved records.
2
ChatGPT and GPT-based LLMs dominated review automation research, appearing in 126 studies (73.2%).
3
GPT-based models outperformed BERT-based models for data extraction, achieving pooled mean precision of 83.0% and recall of 86.0%, although numeric-data extraction remained less accurate.
4
LLM applications most commonly targeted publication searching (60 studies, 34.9%) and data extraction (54 studies, 31.4%).
5
Only 26 studies (15.1%) were actual reviews that explicitly acknowledged using LLMs.

Large language models used in scientific literature-review automation

LLM applications, stage-specific usage, and performance in automating systematic-review processes

Publication Details
Publication Date
2025-04-12
Journal
Publisher
ISSN
Cited by
130
Access Type
Author Information
Authors
Alexander Bakumenko
Nina Hubig
Leslie Lenert
Dmitry Scherbakov
Vinita Jansari
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%