Aggregated Residual Transformations for Deep Neural Networks
Агрегированные остаточные преобразования для глубоких нейронных сетей
2017-07-01
SCID: 54.1/bmgm3cjm
Discuss with AI
COCO object detectionImageNet image classificationResNeXtaggregated residual transformationscardinality
Figures from the paper
Abstract (AI)
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call cardinality (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.
Key Findings
1
Defines cardinality—the number of parallel transformations—as a distinct architectural dimension alongside network depth and width.
2
Introduces ResNeXt, a modular architecture that repeatedly aggregates multiple transformations sharing the same topology.
3
On ImageNet-1K, increasing cardinality improves classification accuracy even when computational complexity is held constant.
4
ResNeXt achieved second place in the ILSVRC 2016 classification task and outperformed corresponding ResNet models on ImageNet-5K and COCO detection.
5
When increasing model capacity, increasing cardinality is more effective than increasing depth or width.
Research Object
ResNeXt deep neural network architecture (aggregated residual transformations multi-branch network) for image classification
Research Subject
The effects of transformation cardinality, depth, and width on image-classification accuracy and model capacity
Publication Details
Publication Date
2017-07-01
Journal
Publisher
ISSN
Cited by
12039
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai6
ImageNet classification with deep convolutional neural networks2017
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
Rethinking the Inception Architecture for Computer Vision2016
Learning Multiple Layers of Features from Tiny Images2024
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation2016
Deep Residual Learning for Image Recognition2016
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Squeeze-and-Excitation Networks2018
Non-local Neural Networks2018
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices2018
ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks2020
Review of deep learning: concepts, CNN architectures, challenges, applications, future directions2021
Object Detection With Deep Learning: A Review2019
A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Deep Learning for Generic Object Detection: A Survey2019
Attention mechanisms in computer vision: A survey2022
PVT v2: Improved baselines with pyramid vision transformer2022
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet2021
Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks2023
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification2021
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows2022
CMT: Convolutional Neural Networks Meet Vision Transformers2022
Photonic matrix multiplication lights up photonic accelerator and beyond2022
A Small-Sized Object Detection Oriented Multi-Scale Feature Fusion Approach With Application to Defect Detection2022
Deep Learning for Unmanned Aerial Vehicle-Based Object Detection and Tracking: A survey2021