Generalization bias in large language model summarization of scientific research

Смещение обобщения при суммаризации научных исследований большими языковыми моделями
Benjamin Chin‐Yee, Uwe Peters
2025-04-01

generalization accuracygeneralization biaslarge language modelsovergeneralization of conclusionsscientific text summarization
Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet, comparing 4900 LLM-generated summaries to their original scientific texts. Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26–73% of cases. In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations (odds ratio = 4.85, 95% CI [3.06, 7.70], p < 0.001). Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings. We highlight potential mitigation strategies, including lowering LLM temperature settings and benchmarking LLMs for generalization accuracy.
1
DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralized scientific results in 26–73% of evaluated cases.
2
LLM-generated science summaries were nearly five times more likely than human-authored summaries to contain broad generalizations (odds ratio 4.85, 95% CI 3.06–7.70, p < 0.001).
3
Most of 10 evaluated large language models generated scientific summaries that generalized conclusions more broadly than warranted, even when prompted for accuracy.
4
Newer language models tended to perform worse than earlier models in accurately preserving the generalization scope of scientific findings.
5
The observed overgeneralization bias creates a risk of large-scale misinterpretation; suggested mitigations include lowering temperature settings and benchmarking generalization accuracy.

large language model-generated summaries of scientific research texts

overgeneralization bias and generalization accuracy of scientific conclusions in LLM-generated summaries

Publication Details
Publication Date
2025-04-01
Journal
Publisher
ISSN
Cited by
115
Access Type
Author Information
Authors
Benjamin Chin‐Yee
Uwe Peters
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%