Are Transformers Effective for Time Series Forecasting?

Эффективны ли трансформеры для прогнозирования временных рядов?
Lei Zhang, Ailing Zeng, Muxi Chen, Qiang Xu
2023-06-26

LTSF-LinearTransformerslong-term time series forecastingpositional encodingself-attention temporal information loss
Recently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful solution to extract the semantic correlations among the elements in a long sequence. However, in time series modeling, we are to extract the temporal relations in an ordered set of continuous points. While employing positional encoding and using tokens to embed sub-series in Transformers facilitate preserving some ordering information, the nature of the permutation-invariant self-attention mechanism inevitably results in temporal information loss. To validate our claim, we introduce a set of embarrassingly simple one-layer linear models named LTSF-Linear for comparison. Experimental results on nine real-life datasets show that LTSF-Linear surprisingly outperforms existing sophisticated Transformer-based LTSF models in all cases, and often by a large margin. Moreover, we conduct comprehensive empirical studies to explore the impacts of various design elements of LTSF models on their temporal relation extraction capability. We hope this surprising finding opens up new research directions for the LTSF task. We also advocate revisiting the validity of Transformer-based solutions for other time series analysis tasks (e.g., anomaly detection) in the future.
1
A set of simple one-layer linear models (LTSF-Linear) was introduced as a baseline for long-term time series forecasting.
2
Comprehensive empirical studies identify that various design elements of LTSF models affect their ability to extract temporal relations.
3
On nine real-life datasets, LTSF-Linear outperforms existing sophisticated Transformer-based LTSF models in all cases, often by a large margin.
4
The results call into question the validity of Transformer-based solutions for LTSF and suggest revisiting their use for other time series tasks (e.g., anomaly detection).
5
Transformer-based LTSF models can lose temporal information due to the permutation-invariant self-attention mechanism, despite positional encodings and token embeddings.

Long-term time series forecasting (LTSF) models and datasets

Effectiveness of Transformer-based models versus simple one-layer linear models (LTSF-Linear) at extracting temporal relations and achieving forecasting performance in long-term time series forecasting, including impacts of model design elements on temporal relation extraction

Publication Details
Publication Date
2023-06-26
Journal
Publisher
ISSN
Cited by
3080
Access Type
Author Information
Authors
Lei Zhang
Ailing Zeng
Muxi Chen
Qiang Xu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%