Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
Tokens-to-Token ViT: обучение визуальных трансформеров с нуля на ImageNet
2021-10-01
SCID: 54.1/pxmpq2k7
Discuss with AI
ImageNet training from scratchT2T-ViTTokens-to-Token Vision Transformerdeep-narrow transformer backbonelayer-wise Tokens-to-Token transformation
Figures from the paper
Abstract (AI)
Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model their global relation for classification. However, ViT achieves inferior performance to CNNs when trained from scratch on a midsize dataset like ImageNet. We find it is because: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines among neighboring pixels, leading to low training sample efficiency; 2) the redundant attention backbone design of ViT leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we propose a new Tokens-To-Token Vision Transformer (T2T-VTT), which incorporates 1) a layer-wise Tokens-to-Token (T2T) transformation to progressively structurize the image to tokens by recursively aggregating neighboring Tokens into one Token (Tokens-to-Token), such that local structure represented by surrounding tokens can be modeled and tokens length can be reduced; 2) an efficient backbone with a deep-narrow structure for vision transformer motivated by CNN architecture design after empirical study. Notably, T2T-ViT reduces the parameter count and MACs of vanilla ViT by half, while achieving more than 3.0% improvement when trained from scratch on ImageNet. It also outperforms ResNets and achieves comparable performance with MobileNets by directly training on ImageNet. For example, T2T-ViT with comparable size to ResNet50 (21.5M parameters) can achieve 83.3% top1 accuracy in image resolution 384x384 on ImageNet.1
Key Findings
1
Introduced an efficient deep-narrow backbone for vision transformers inspired by CNN design, improving feature richness under fixed computation budgets.
2
Proposed Tokens-to-Token (T2T) transformation progressively aggregates neighboring tokens to capture local image structures and reduce token length.
3
T2T-ViT halves parameters and MACs of vanilla ViT while achieving more than 3.0% accuracy improvement when trained from scratch on ImageNet.
4
T2T-ViT outperforms ResNets and matches MobileNets when trained directly on ImageNet; a 21.5M-parameter T2T-ViT achieves 83.3% top-1 accuracy at 384x384 resolution.
5
Vanilla ViT underperforms CNNs when trained from scratch on midsize datasets like ImageNet due to poor local structure modeling and redundant attention backbone design.
Research Object
Tokens-to-Token Vision Transformer (T2T-ViT) model for image classification trained from scratch on ImageNet
Research Subject
Improving training-from-scratch image classification performance by progressively aggregating local image structure into tokens (layer-wise Tokens-to-Token transformation) and using an efficient deep-narrow backbone to reduce parameters/MACs and increase accuracy on ImageNet
Publication Details
Publication Date
2021-10-01
Journal
Publisher
ISSN
Cited by
2324
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai13
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
ImageNet: A large-scale hierarchical image database2009
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Rethinking the Inception Architecture for Computer Vision2016
Squeeze-and-Excitation Networks2018
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Distilling the Knowledge in a Neural Network2015
Aggregated Residual Transformations for Deep Neural Networks2017
Non-local Neural Networks2018
mixup: Beyond Empirical Risk Minimization2017
Pre-Trained Image Processing Transformer2021
Cited by14
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Attention mechanisms in computer vision: A survey2022
PVT v2: Improved baselines with pyramid vision transformer2022
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
Vision Transformers for Single Image Dehazing2023
AST: Audio Spectrogram Transformer2021
Transformer in Transformer2021
CMT: Convolutional Neural Networks Meet Vision Transformers2022
Scaling Vision Transformers2022
ChatGPT and Open-AI Models: A Preliminary Review2023
SSAST: Self-Supervised Audio Spectrogram Transformer2022
RenAIssance: A Survey Into AI Text-to-Image Generation in the Era of Large Model2024