ETC: Encoding Long and Structured Inputs in Transformers

ETC: Кодирование длинных и структурированных входных данных в трансформерах
Joshua Ainslie, Chris Alberti, Santiago Ontañón, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, Yang Li
2020-01-01

ETCTransformersencoding long inputslong-context transformerstructured inputs
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, Li Yang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
1
ETC enables scaling to longer contexts with reduced computational and memory cost compared to dense attention.
2
ETC represents structured inputs (e.g., graphs or hierarchical data) via specialized encoding and attention patterns.
3
ETC uses sparse attention mechanisms to handle much longer sequences than standard Transformers.
4
Introduces ETC, a Transformer variant designed to encode long and structured inputs efficiently.

Transformer-based neural models processing long and structured input sequences (ETC)

Methods for encoding and attending over long and structured inputs in transformers, including sparse attention and relative/global position representations to scale and capture structure

Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Joshua Ainslie
Chris Alberti
Santiago Ontañón
Vaclav Cvicek
Zachary Fisher
Philip Pham
Anirudh Ravula
Sumit Sanghai
Qifan Wang
Yang Li
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%