Learning Deep Transformer Models for Machine Translation
Обучение глубоких моделей Transformer для машинного перевода
2019-01-01
SCID: 54.1/cs4fhnyv
Discuss with AI
Transformer-Bigdeep Transformerlayer combinationlayer normalizationmachine translation
Figures from the paper
Abstract (AI)
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT’16 English-German and NIST OpenMT’12 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
Key Findings
1
A truly deep Transformer can surpass Transformer-Big by using proper layer normalization and a novel method of passing combined previous layers to the next.
2
On WMT’16 English-German and NIST OpenMT’12 Chinese-English, a 30/25-layer encoder deep system outperforms 6-layer Transformer-Big/Base by 0.4–2.4 BLEU points.
3
The proposed deep model is 1.6× smaller in size than Transformer-Big.
4
The proposed deep model trains about 3× faster than Transformer-Big.
Research Object
Deep Transformer-based neural machine translation model (deep Transformer encoder-decoder)
Research Subject
Techniques for training and architectural modifications enabling very deep Transformer networks (e.g., layer normalization usage and novel layer-combination/passing) and their impact on translation quality, model size, and training speed compared to Transformer-Big/Base
Publication Details
Publication Date
2019-01-01
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest