Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks

Беги, не иди: в поисках более высокой FLOPS для более быстрых нейронных сетей
Chul‐Ho Lee, S.-H. Gary Chan, Jierun Chen, Shiu-hong Kao, Hao He, Weipeng Zhuo, Wen Song
2023-06-01

FLOPs vs FLOPS (floating-point operations per second)FasterNetdepthwise convolutioninference throughput / latency on GPU, CPU, ARMpartial convolution (PConv)
To design fast neural networks, many works have been focusing on reducing the number of floating-point operations (FLOPs). We observe that such reduction in FLOPs, however, does not necessarily lead to a similar level of re-duction in latency. This mainly stems from inefficiently low floating-point operations per second (FLOPS). To achieve faster networks, we revisit popular operators and demonstrate that such low FLOPS is mainly due to frequent memory access of the operators, especially the depthwise con-volution. We hence propose a novel partial convolution (PConv) that extracts spatial features more efficiently, by cutting down redundant computation and memory access simultaneously. Building upon our PConv, we further propose FasterNet, a new family of neural networks, which attains substantially higher running speed than others on a wide range of devices, without compromising on accuracy for various vision tasks. For example, on ImageNet-lk, our tiny FasterNet-TO is 2.8×, 3.3×, and 2.4× faster than MobileViT-XXS on GPU, CPU, and ARM processors, respectively, while being 2.9% more accurate. Our large FasterNet-L achieves impressive 83.5% top-1 accuracy, on par with the emerging Swin-B, while having 36% higher inference throughput on GPU, as well as saving 37% compute time on CPU. Code is available at https://github.com/JierunChen/FasterNet.
1
FasterNet, built on PConv, achieves substantially higher running speed across devices without compromising accuracy on vision tasks.
2
FasterNet-L (large) achieves 83.5% top-1 accuracy (comparable to Swin-B) while delivering 36% higher GPU inference throughput and saving 37% CPU compute time.
3
FasterNet-TO (tiny) is 2.8×, 3.3×, and 2.4× faster than MobileViT-XXS on GPU, CPU, and ARM respectively, while being 2.9% more accurate on ImageNet-1k.
4
Reducing FLOPs does not necessarily reduce latency because many networks have inefficiently low FLOPS due to frequent memory access, especially in depthwise convolution.
5
The authors propose partial convolution (PConv) which reduces redundant computation and memory access to extract spatial features more efficiently.

FasterNet family of neural network architectures (including the proposed partial convolution PConv and FasterNet-TO/L variants)

Improving real-world inference speed (throughput/latency) by increasing effective FLOPS and reducing memory access via the partial convolution operator to achieve higher running speed without sacrificing accuracy across vision tasks

Publication Details
Publication Date
2023-06-01
Journal
Publisher
ISSN
Cited by
2253
Access Type
Author Information
Authors
Chul‐Ho Lee
S.-H. Gary Chan
Jierun Chen
Shiu-hong Kao
Hao He
Weipeng Zhuo
Wen Song
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%