Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Чрезвычайно большие нейронные сети: слой разреженно управляемой смеси экспертов
2017-01-23
SCID: 54.1/ttrgmtdv
Discuss with AI
Mixture-of-Experts (MoE)Sparsely-Gated Mixture-of-Expertsconditional computationlarge-scale language modeling and machine translationsparse gating network
Figures from the paper
Abstract (AI)
The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges. In this work, we address these challenges and finally realize the promise of conditional computation, achieving greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters. We introduce a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks. A trainable gating network determines a sparse combination of these experts to use for each example. We apply the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora. We present model architectures in which a MoE with up to 137 billion parameters is applied convolutionally between stacked LSTM layers. On large language modeling and machine translation benchmarks, these models achieve significantly better results than state-of-the-art at lower computational cost.
Key Findings
1
Applied the MoE to language modeling and machine translation, achieving significantly better results than state-of-the-art at lower computational cost.
2
Constructed architectures applying an MoE (up to 137 billion parameters) convolutionally between stacked LSTM layers.
3
Introduced a Sparsely-Gated Mixture-of-Experts (MoE) layer with up to thousands of feed-forward sub-networks and a trainable sparse gating network.
4
Realized conditional computation yielding greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters.
Research Object
Sparsely-Gated Mixture-of-Experts (MoE) layer consisting of up to thousands of feed-forward sub-networks used in neural network architectures for language modeling and machine translation
Research Subject
Scaling model capacity and conditional computation via sparse gating: per-example expert selection, achieving massive parameter counts (up to 137 billion) with minor computational overhead and improved performance on large language modeling and machine translation benchmarks
Publication Details
Publication Date
2017-01-23
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest