PaLM: Scaling Language Modeling with Pathways
PaLM: Масштабирование языкового моделирования с помощью Pathways
2022-04-05
SCID: 54.1/tzn3x73m
Discuss with AI
BIG-benchPathways Language Modelfew-shot learningmulti-step reasoningscaling language models
Figures from the paper
Abstract (AI)
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Key Findings
1
Many BIG-bench tasks exhibited discontinuous scaling, with steep performance improvements appearing only at the largest model scale.
2
PaLM 540B surpassed fine-tuned state-of-the-art systems on multiple multi-step reasoning tasks and exceeded average human performance on BIG-bench.
3
PaLM demonstrated strong multilingual and source-code generation capabilities, while the study also analyzed bias, toxicity, memorization, and ethical mitigation strategies.
4
PaLM is a densely activated 540-billion-parameter Transformer language model trained efficiently across 6144 TPU v4 chips using the Pathways system.
5
Scaling PaLM produced continued few-shot learning gains, achieving state-of-the-art results across hundreds of language understanding and generation benchmarks.
Research Object
Pathways Language Model (PaLM), a 540-billion-parameter densely activated Transformer language model
Research Subject
The effects of model scale on few-shot learning performance, including language understanding and generation, reasoning, multilingual ability, code generation, bias, toxicity, and training-data memorization
Publication Details
Publication Date
2022-04-05
Journal
Publisher
ISSN
Cited by
2136
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest