A Review of Current Trends, Techniques, and Challenges in Large Language Models (LLMs)
Обзор современных тенденций, методов и проблем больших языковых моделей (LLM)
2024-03-01
SCID: 54.1/kqhr6mqj
Discuss with AI
benchmarksfine-tuningin-context learninglarge language models (LLMs)natural language generation (NLG)natural language understanding (NLU)pretraining objectivesscalability (depth, width, data size)self-supervised pretrainingtransfer learning
Figures from the paper
Abstract (AI)
Natural language processing (NLP) has significantly transformed in the last decade, especially in the field of language modeling. Large language models (LLMs) have achieved SOTA performances on natural language understanding (NLU) and natural language generation (NLG) tasks by learning language representation in self-supervised ways. This paper provides a comprehensive survey to capture the progression of advances in language models. In this paper, we examine the different aspects of language models, which started with a few million parameters but have reached the size of a trillion in a very short time. We also look at how these LLMs transitioned from task-specific to task-independent to task-and-language-independent architectures. This paper extensively discusses different pretraining objectives, benchmarks, and transfer learning methods used in LLMs. It also examines different finetuning and in-context learning techniques used in downstream tasks. Moreover, it explores how LLMs can perform well across many domains and datasets if sufficiently trained on a large and diverse dataset. Next, it discusses how, over time, the availability of cheap computational power and large datasets have improved LLM’s capabilities and raised new challenges. As part of our study, we also inspect LLMs from the perspective of scalability to see how their performance is affected by the model’s depth, width, and data size. Lastly, we provide an empirical comparison of existing trends and techniques and a comprehensive analysis of where the field of LLM currently stands.
Key Findings
1
Different pretraining objectives, benchmarks, and transfer learning methods are central to LLM performance, with finetuning and in-context learning critical for downstream tasks.
2
Increases in affordable computational power and large datasets have improved LLM capabilities but simultaneously introduced new challenges; model performance scales with depth, width, and data size.
3
LLMs have grown rapidly from millions to trillion-parameter scales, becoming SOTA on NLU and NLG tasks through self-supervised representation learning.
4
Language model architectures evolved from task-specific to task-independent and then to task-and-language-independent designs.
5
Sufficient training on large, diverse datasets enables LLMs to perform well across many domains and datasets.
Research Object
Large language models (LLMs)
Research Subject
Current trends, techniques, scalability, pretraining objectives, finetuning and in-context learning methods, benchmarks, transfer learning, domain/generalization performance, and challenges in LLM development and deployment
Publication Details
Publication Date
2024-03-01
Journal
Publisher
ISSN
Cited by
241
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest