Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks
Беги, не иди: в поисках более высокой FLOPS для более быстрых нейронных сетей
2023-06-01
SCID: 54.1/yjp3329h
Discuss with AI
FLOPs vs FLOPS (floating-point operations per second)FasterNetdepthwise convolutioninference throughput / latency on GPU, CPU, ARMpartial convolution (PConv)
Figures from the paper
Abstract (AI)
To design fast neural networks, many works have been focusing on reducing the number of floating-point operations (FLOPs). We observe that such reduction in FLOPs, however, does not necessarily lead to a similar level of re-duction in latency. This mainly stems from inefficiently low floating-point operations per second (FLOPS). To achieve faster networks, we revisit popular operators and demonstrate that such low FLOPS is mainly due to frequent memory access of the operators, especially the depthwise con-volution. We hence propose a novel partial convolution (PConv) that extracts spatial features more efficiently, by cutting down redundant computation and memory access simultaneously. Building upon our PConv, we further propose FasterNet, a new family of neural networks, which attains substantially higher running speed than others on a wide range of devices, without compromising on accuracy for various vision tasks. For example, on ImageNet-lk, our tiny FasterNet-TO is 2.8×, 3.3×, and 2.4× faster than MobileViT-XXS on GPU, CPU, and ARM processors, respectively, while being 2.9% more accurate. Our large FasterNet-L achieves impressive 83.5% top-1 accuracy, on par with the emerging Swin-B, while having 36% higher inference throughput on GPU, as well as saving 37% compute time on CPU. Code is available at https://github.com/JierunChen/FasterNet.
Key Findings
1
FasterNet, built on PConv, achieves substantially higher running speed across devices without compromising accuracy on vision tasks.
2
FasterNet-L (large) achieves 83.5% top-1 accuracy (comparable to Swin-B) while delivering 36% higher GPU inference throughput and saving 37% CPU compute time.
3
FasterNet-TO (tiny) is 2.8×, 3.3×, and 2.4× faster than MobileViT-XXS on GPU, CPU, and ARM respectively, while being 2.9% more accurate on ImageNet-1k.
4
Reducing FLOPs does not necessarily reduce latency because many networks have inefficiently low FLOPS due to frequent memory access, especially in depthwise convolution.
5
The authors propose partial convolution (PConv) which reduces redundant computation and memory access to extract spatial features more efficiently.
Research Object
FasterNet family of neural network architectures (including the proposed partial convolution PConv and FasterNet-TO/L variants)
Research Subject
Improving real-world inference speed (throughput/latency) by increasing effective FLOPS and reducing memory access via the partial convolution operator to achieve higher running speed without sacrificing accuracy across vision tasks
Publication Details
Publication Date
2023-06-01
Journal
Publisher
ISSN
Cited by
2253
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai11
ImageNet classification with deep convolutional neural networks2017
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Rethinking the Inception Architecture for Computer Vision2016
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Distilling the Knowledge in a Neural Network2015
Aggregated Residual Transformations for Deep Neural Networks2017
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices2018
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
mixup: Beyond Empirical Risk Minimization2017
Swin Transformer V2: Scaling Up Capacity and Resolution2022