Very Deep Convolutional Networks for Large-Scale Image Recognition
Очень глубокие сверточные сети для масштабного распознавания изображений
2014-09-04
SCID: 54.1/qv9azsyz
Discuss with AI
16-19 weight layers3x3 convolution filtersImageNet Challenge 2014large-scale image recognitionvery deep convolutional networks
Figures from the paper
Abstract (AI)
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.
Key Findings
1
Increasing convolutional network depth with very small 3x3 filters significantly improves large-scale image recognition accuracy.
2
Pushing network depth to 16–19 weight layers yields substantial improvement over prior-art configurations.
3
The learned deep representations generalise well to other datasets, achieving state-of-the-art results.
4
Their deep models achieved first and second places in ImageNet Challenge 2014 localisation and classification tracks respectively.
5
Two best-performing ConvNet models were made publicly available to facilitate further research.
Research Object
Very deep convolutional networks (ConvNets) with small (3x3) convolution filters trained for large-scale image recognition
Research Subject
Effect of network depth (increasing to 16–19 weight layers) on classification and localisation accuracy and the generalisation of learned visual representations to other image datasets
Publication Details
Publication Date
2014-09-04
Journal
Publisher
ISSN
Cited by
75434
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Rethinking the Inception Architecture for Computer Vision2016
Squeeze-and-Excitation Networks2018
Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks2017
A survey on Image Data Augmentation for Deep Learning2019
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network2017
Aggregated Residual Transformations for Deep Neural Networks2017
Non-local Neural Networks2018
Learning Deep Features for Discriminative Localization2016
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices2018
Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising2017
ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks2020
Review of deep learning: concepts, CNN architectures, challenges, applications, future directions2021
Image Style Transfer Using Convolutional Neural Networks2016
Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning2016
Object Detection With Deep Learning: A Review2019
Least Squares Generative Adversarial Networks2017
A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
Convolutional neural networks: an overview and application in radiology2018