Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Чрезвычайно большие нейронные сети: слой разреженно управляемой смеси экспертов
Noam Shazeer, Quoc V. Le, Geoffrey E. Hinton, Jeff Dean, Andy Davis, Azalia Mirhoseini, Krzysztof Maziarz
2017-01-23

Mixture-of-Experts (MoE)Sparsely-Gated Mixture-of-Expertsconditional computationlarge-scale language modeling and machine translationsparse gating network
The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges. In this work, we address these challenges and finally realize the promise of conditional computation, achieving greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters. We introduce a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks. A trainable gating network determines a sparse combination of these experts to use for each example. We apply the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora. We present model architectures in which a MoE with up to 137 billion parameters is applied convolutionally between stacked LSTM layers. On large language modeling and machine translation benchmarks, these models achieve significantly better results than state-of-the-art at lower computational cost.
1
Applied the MoE to language modeling and machine translation, achieving significantly better results than state-of-the-art at lower computational cost.
2
Constructed architectures applying an MoE (up to 137 billion parameters) convolutionally between stacked LSTM layers.
3
Introduced a Sparsely-Gated Mixture-of-Experts (MoE) layer with up to thousands of feed-forward sub-networks and a trainable sparse gating network.
4
Realized conditional computation yielding greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters.

Sparsely-Gated Mixture-of-Experts (MoE) layer consisting of up to thousands of feed-forward sub-networks used in neural network architectures for language modeling and machine translation

Scaling model capacity and conditional computation via sparse gating: per-example expert selection, achieving massive parameter counts (up to 137 billion) with minor computational overhead and improved performance on large language modeling and machine translation benchmarks

Publication Details
Publication Date
2017-01-23
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Noam Shazeer
Quoc V. Le
Geoffrey E. Hinton
Jeff Dean
Andy Davis
Azalia Mirhoseini
Krzysztof Maziarz
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%