Efficient Transformers: A Survey
Эффективные трансформеры: обзор
2022-04-22
SCID: 54.1/tg9kqs9n
Discuss with AI
Efficient TransformersLinformerLongformerPerformerReformerTransformer architecturesX-former modelscomputational efficiencylanguage, vision, and reinforcement learning domainsmemory efficiency
Figures from the paper
Abstract (AI)
Transformer model architectures have garnered immense interest lately due to their effectiveness across a range of domains like language, vision, and reinforcement learning. In the field of natural language processing for example, Transformers have become an indispensable staple in the modern deep learning stack. Recently, a dizzying number of “X-former” models have been proposed—Reformer, Linformer, Performer, Longformer, to name a few—which improve upon the original Transformer architecture, many of which make improvements around computational and memory efficiency . With the aim of helping the avid researcher navigate this flurry, this article characterizes a large and thoughtful selection of recent efficiency-flavored “X-former” models, providing an organized and comprehensive overview of existing work and models across multiple domains.
Key Findings
1
Many recent models focus specifically on efficiency improvements of the Transformer architecture.
2
Numerous 'X-former' variants (e.g., Reformer, Linformer, Performer, Longformer) have been proposed to improve computational and memory efficiency over the original Transformer.
3
This article provides an organized, comprehensive overview and characterization of a large selection of efficiency-focused Transformer models across multiple domains to aid researchers.
4
Transformer architectures are highly effective across domains including language, vision, and reinforcement learning.
Research Object
Transformer model architectures (efficiency-flavored variants such as Reformer, Linformer, Performer, Longformer)
Research Subject
Computational and memory efficiency improvements, design characteristics, and comparative overview of efficiency-focused Transformer variants across domains
Publication Details
Publication Date
2022-04-22
Journal
Publisher
ISSN
Cited by
1063
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai10
ImageNet: A large-scale hierarchical image database2009
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Distilling the Knowledge in a Neural Network2015
Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer2019
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context2019
Longformer: The Long-Document Transformer2020
PaLM: Scaling Language Modeling with Pathways2022
ETC: Encoding Long and Structured Inputs in Transformers2020
Cited by7
On the Opportunities and Risks of Foundation Models2021
LoFTR: Detector-Free Local Feature Matching with Transformers2021
Transformers in Time Series: A Survey2023
Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better2023
Self-Supervised Speech Representation Learning: A Review2022
Large language models (LLMs): survey, technical frameworks, and future challenges2024
A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges2025