Self-Attention with Relative Position Representations
Механизм самовнимания с представлениями относительной позиции
2018-01-01
SCID: 54.1/ng8mdjd6
Discuss with AI
Attention MechanismRelative Position RepresentationsSelf-AttentionShaw Uszkoreit VaswaniTransformers
Figures from the paper
Abstract (AI)
Peter Shaw, Jakob Uszkoreit, Ashish Vaswani. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). 2018.
Key Findings
1
Augmenting self-attention with relative position information allows the model to better capture sequence order and distances between tokens.
2
Introducing relative position representations for self-attention improves modeling of token interactions compared to absolute positional encodings.
3
The proposed relative position representations can be integrated into the transformer-style self-attention mechanism without major architectural changes.
4
Using relative position representations leads to improved empirical performance on tasks evaluated in the paper (reported in the 2018 NAACL short paper).
Research Object
Self-attention mechanism in neural sequence models
Research Subject
Incorporation of relative position representations into self-attention to model positional relationships between sequence elements
Publication Details
Publication Date
2018-01-01
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai5
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation2016
Rethinking the Inception Architecture for Computer Vision2016
Effective Approaches to Attention-based Neural Machine Translation2015
Adam: A Method for Stochastic Optimization2014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation2014