ETC: Encoding Long and Structured Inputs in Transformers
ETC: Кодирование длинных и структурированных входных данных в трансформерах
2020-01-01
SCID: 54.1/y4v9mjsz
Discuss with AI
ETCTransformersencoding long inputslong-context transformerstructured inputs
Figures from the paper
Abstract (AI)
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, Li Yang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Key Findings
1
ETC enables scaling to longer contexts with reduced computational and memory cost compared to dense attention.
2
ETC represents structured inputs (e.g., graphs or hierarchical data) via specialized encoding and attention patterns.
3
ETC uses sparse attention mechanisms to handle much longer sequences than standard Transformers.
4
Introduces ETC, a Transformer variant designed to encode long and structured inputs efficiently.
Research Object
Transformer-based neural models processing long and structured input sequences (ETC)
Research Subject
Methods for encoding and attending over long and structured inputs in transformers, including sparse attention and relative/global position representations to scale and capture structure
Publication Details
Publication Date
2020-01-01
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai9
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Longformer: The Long-Document Transformer2020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case2019
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Representation Learning with Contrastive Predictive Coding2018
Hierarchical Attention Networks for Document Classification2016
The Graph Neural Network Model2008