Aggregated Residual Transformations for Deep Neural Networks

Агрегированные остаточные преобразования для глубоких нейронных сетей
Kaiming He, Piotr Dollár, Ross Girshick, Saining Xie, Zhuowen Tu
2017-07-01

COCO object detectionImageNet image classificationResNeXtaggregated residual transformationscardinality
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call cardinality (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.
1
Defines cardinality—the number of parallel transformations—as a distinct architectural dimension alongside network depth and width.
2
Introduces ResNeXt, a modular architecture that repeatedly aggregates multiple transformations sharing the same topology.
3
On ImageNet-1K, increasing cardinality improves classification accuracy even when computational complexity is held constant.
4
ResNeXt achieved second place in the ILSVRC 2016 classification task and outperformed corresponding ResNet models on ImageNet-5K and COCO detection.
5
When increasing model capacity, increasing cardinality is more effective than increasing depth or width.

ResNeXt deep neural network architecture (aggregated residual transformations multi-branch network) for image classification

The effects of transformation cardinality, depth, and width on image-classification accuracy and model capacity

Publication Details
Publication Date
2017-07-01
Journal
Publisher
ISSN
Cited by
12039
Access Type
Author Information
Authors
Kaiming He
Piotr Dollár
Ross Girshick
Saining Xie
Zhuowen Tu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%