Efficient Acceleration of Deep Learning Inference on Resource-Constrained Edge Devices: A Review
Эффективное ускорение вывода моделей глубокого обучения на периферийных устройствах с ограниченными ресурсами: обзор
2022-12-14
SCID: 54.1/xttwfbdw
Discuss with AI
algorithm–hardware codesigndeep learning acceleratorsdeep learning inferenceedge computingresource-constrained edge devices
Figures from the paper
Abstract (AI)
Successful integration of deep neural networks (DNNs) or deep learning (DL) has resulted in breakthroughs in many areas. However, deploying these highly accurate models for data-driven, learned, automatic, and practical machine learning (ML) solutions to end-user applications remains challenging. DL algorithms are often computationally expensive, power-hungry, and require large memory to process complex and iterative operations of millions of parameters. Hence, training and inference of DL models are typically performed on high-performance computing (HPC) clusters in the cloud. Data transmission to the cloud results in high latency, round-trip delay, security and privacy concerns, and the inability of real-time decisions. Thus, processing on edge devices can significantly reduce cloud transmission cost. Edge devices are end devices closest to the user, such as mobile phones, cyber–physical systems (CPSs), wearables, the Internet of Things (IoT), embedded and autonomous systems, and intelligent sensors. These devices have limited memory, computing resources, and power-handling capability. Therefore, optimization techniques at both the hardware and software levels have been developed to handle the DL deployment efficiently on the edge. Understanding the existing research, challenges, and opportunities is fundamental to leveraging the next generation of edge devices with artificial intelligence (AI) capability. Mainly, four research directions have been pursued for efficient DL inference on edge devices: 1) novel DL architecture and algorithm design; 2) optimization of existing DL methods; 3) development of algorithm–hardware codesign; and 4) efficient accelerator design for DL deployment. This article focuses on surveying each of the four research directions, providing a comprehensive review of the state-of-the-art tools and techniques for efficient edge inference.
Key Findings
1
Cloud-based inference introduces latency, round-trip delays, privacy and security concerns, and limitations for real-time decision-making.
2
Deploying deep learning on edge devices is challenging because models are computationally expensive, power-hungry, and memory-intensive.
3
Edge processing can reduce cloud transmission costs and enable more responsive, privacy-aware inference near end users.
4
Efficient edge inference research primarily follows four directions: novel architectures and algorithms, optimization of existing methods, algorithm–hardware codesign, and dedicated accelerator design.
5
The review surveys state-of-the-art tools, techniques, challenges, and opportunities across these four directions for resource-constrained edge devices.
Research Object
Deep learning inference on resource-constrained edge devices
Research Subject
Efficiency optimization of edge deep learning inference across architectures, algorithms, hardware–software codesign, and accelerator design
Publication Details
Publication Date
2022-12-14
Journal
Publisher
ISSN
Cited by
403
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai13
Very Deep Convolutional Networks for Large-Scale Image Recognition2014
Gradient-based learning applied to document recognition1998
Squeeze-and-Excitation Networks2018
Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023
Distilling the Knowledge in a Neural Network2015
Representation Learning: A Review and New Perspectives2013
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices2018
Reinforcement Learning: A Survey1996
Transformers: State-of-the-Art Natural Language Processing2020
Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)2018
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation2016
A survey on semi-supervised learning2019
Power to the People: The Role of Humans in Interactive Machine Learning2014