Learning Deep Transformer Models for Machine Translation

Обучение глубоких моделей Transformer для машинного перевода
Qiang Wang, Tong Xiao, Bei Li, Jingbo Zhu, Changliang Li, Derek F. Wong, Lidia S. Chao
2019-01-01

Transformer-Bigdeep Transformerlayer combinationlayer normalizationmachine translation
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT’16 English-German and NIST OpenMT’12 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
1
A truly deep Transformer can surpass Transformer-Big by using proper layer normalization and a novel method of passing combined previous layers to the next.
2
On WMT’16 English-German and NIST OpenMT’12 Chinese-English, a 30/25-layer encoder deep system outperforms 6-layer Transformer-Big/Base by 0.4–2.4 BLEU points.
3
The proposed deep model is 1.6× smaller in size than Transformer-Big.
4
The proposed deep model trains about 3× faster than Transformer-Big.

Deep Transformer-based neural machine translation model (deep Transformer encoder-decoder)

Techniques for training and architectural modifications enabling very deep Transformer networks (e.g., layer normalization usage and novel layer-combination/passing) and their impact on translation quality, model size, and training speed compared to Transformer-Big/Base

Publication Details
Publication Date
2019-01-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Qiang Wang
Tong Xiao
Bei Li
Jingbo Zhu
Changliang Li
Derek F. Wong
Lidia S. Chao
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%