PaLM: Scaling Language Modeling with Pathways

PaLM: Масштабирование языкового моделирования с помощью Pathways
Noam Shazeer, Slav Petrov, Katherine Lee, Daphne Ippolito, Douglas Eck, Emily Reif, Aakanksha Chowdhery, Gaurav Mishra, James T. Bradbury, Sharan Narang, Henryk Michalewski, Andrew M. Dai, Orhan Fırat, Michael Isard, Érica Rodrigues Moreira, Jacob Austin, Xavier García, Thanumalayan Sankaranarayana Pillai, Jacob Devlin, Hyeontaek Lim, Shivani Agrawal, Joshua Maynez, Kevin Robinson, Sanjay Ghemawat, Marie Pellat, Vedant Misra, Pengcheng Yin, Charles Sutton, Xuezhi Wang, Anselm Levskaya, Mark Omernick, Zongwei Zhou, Brennan Saeta, Denny Zhou, Parker Schuh, Jason Lee, Yi Tay, Barret Zoph, Maarten Bosma, Jeff Dean, Sebastian Gehrmann, Adam Roberts, Paul Barham, Ryan Sepassi, Rewon Child, Hyung Won Chung, Michele Catasta, Ben Hutchinson, Parker Barnes, Vinodkumar Prabhakaran, Mark Díaz, Aitor Lewkowycz, Alexander Spiridonov, Kensen Shi, Sasha Tsvyashchenko, Abhishek S. Rao, Nan Du, Reiner Pope, Guy Gur-Ari, Toju Duke, Sunipa Dev, Liam Fedus, David Luan, D. Dohan, Oleksandr Polozov, Kathy Meier-Hellstern, Noah Fiedel, Wei, Jason, Meier-Hellstern, Kathy
2022-04-05

BIG-benchPathways Language Modelfew-shot learningmulti-step reasoningscaling language models
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
1
Many BIG-bench tasks exhibited discontinuous scaling, with steep performance improvements appearing only at the largest model scale.
2
PaLM 540B surpassed fine-tuned state-of-the-art systems on multiple multi-step reasoning tasks and exceeded average human performance on BIG-bench.
3
PaLM demonstrated strong multilingual and source-code generation capabilities, while the study also analyzed bias, toxicity, memorization, and ethical mitigation strategies.
4
PaLM is a densely activated 540-billion-parameter Transformer language model trained efficiently across 6144 TPU v4 chips using the Pathways system.
5
Scaling PaLM produced continued few-shot learning gains, achieving state-of-the-art results across hundreds of language understanding and generation benchmarks.

Pathways Language Model (PaLM), a 540-billion-parameter densely activated Transformer language model

The effects of model scale on few-shot learning performance, including language understanding and generation, reasoning, multilingual ability, code generation, bias, toxicity, and training-data memorization

Publication Details
Publication Date
2022-04-05
Journal
Publisher
ISSN
Cited by
2136
Access Type
Author Information
Authors
Noam Shazeer
Slav Petrov
Katherine Lee
Daphne Ippolito
Douglas Eck
Emily Reif
Aakanksha Chowdhery
Gaurav Mishra
James T. Bradbury
Sharan Narang
Henryk Michalewski
Andrew M. Dai
Orhan Fırat
Michael Isard
Érica Rodrigues Moreira
Jacob Austin
Xavier García
Thanumalayan Sankaranarayana Pillai
Jacob Devlin
Hyeontaek Lim
Shivani Agrawal
Joshua Maynez
Kevin Robinson
Sanjay Ghemawat
Marie Pellat
Vedant Misra
Pengcheng Yin
Charles Sutton
Xuezhi Wang
Anselm Levskaya
Mark Omernick
Zongwei Zhou
Brennan Saeta
Denny Zhou
Parker Schuh
Jason Lee
Yi Tay
Barret Zoph
Maarten Bosma
Jeff Dean
Sebastian Gehrmann
Adam Roberts
Paul Barham
Ryan Sepassi
Rewon Child
Hyung Won Chung
Michele Catasta
Ben Hutchinson
Parker Barnes
Vinodkumar Prabhakaran
Mark Díaz
Aitor Lewkowycz
Alexander Spiridonov
Kensen Shi
Sasha Tsvyashchenko
Abhishek S. Rao
Nan Du
Reiner Pope
Guy Gur-Ari
Toju Duke
Sunipa Dev
Liam Fedus
David Luan
D. Dohan
Oleksandr Polozov
Kathy Meier-Hellstern
Noah Fiedel
Wei, Jason
Meier-Hellstern, Kathy
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%