Neural Machine Translation by Jointly Learning to Align and Translate
Нейронный машинный перевод посредством совместного обучения выравниванию и переводу
2014-09-01
SCID: 54.1/pr2u9bet
Discuss with AI
English-to-French translationencoder-decoderfixed-length vector bottleneckneural machine translationsoft alignment
Figures from the paper
Abstract (AI)
Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.
Key Findings
1
Fixed-length vector encoding in encoder-decoder NMT is a bottleneck limiting translation performance.
2
Proposed extension allows the model to soft-search source sentence parts relevant for each target word (soft alignment).
3
Qualitative analysis shows the learned soft-alignments agree well with human intuition.
4
The approach achieves translation performance comparable to a state-of-the-art phrase-based system on English-to-French translation.
5
The model jointly learns to align and translate without explicit hard segmentation of source parts.
Research Object
Neural machine translation encoder-decoder model with an attention mechanism that jointly learns to align and translate
Research Subject
Learning soft (implicit) alignments between source and target sentence parts and using them to improve translation performance compared to fixed-length-vector encoder–decoder models
Publication Details
Publication Date
2014-09-01
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai3
Cited by20
Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting2021
Emerging Properties in Self-Supervised Vision Transformers2021
Are Transformers Effective for Time Series Forecasting?2023
Deep Learning Enabled Semantic Communication Systems2021
A Survey on the Explainability of Supervised Machine Learning2021
Deep learning for AI2021
Deep Learning for Anomaly Detection in Time-Series Data: Review, Analysis, and Guidelines2021
A Survey on Explainable Artificial Intelligence (XAI): Toward Medical XAI2020
Attention in Natural Language Processing2020
Knowledge Graph Completion: A Review2020
Neural Speech Synthesis with Transformer Network2019
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Learning Deep Transformer Models for Machine Translation2019
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation2018
Tacotron: Towards End-to-End Speech Synthesis2017
Deep learning in bioinformatics2016
Hierarchical Attention Networks for Document Classification2016
Sequence-Level Knowledge Distillation2016
Modeling Coverage for Neural Machine Translation2016
Effective Approaches to Attention-based Neural Machine Translation2015