Abstract
Using Heterogeneous Multi-Processor System-on-Chips (HMPSoCs) for Deep Neural Network (DNN) inference has become commonplace in edge devices. However, reducing the DNN inference latency on resource-constrained edge devices remains a first-class constraint. Early-Exit (EE) DNNs offer variable, input-dependent inference latency, aiming to reduce average inference latency by lowering the latency of simple inputs at the cost of increased latency for complex inputs. However, the overheads introduced by exit branches reduce the potential performance gains with EE DNNs.
Furthermore, existing CPU- or GPU-only implementations of EE DNN inference under-utilize the HMPSoC. To address this limitation, we propose a cooperative and parallel CPU-GPU execution approach1 for EE DNN inference that effectively distributes computations across all HMPSoC processors, minimizing latency variations. Our approach allows EE DNNs to achieve reduced average latency on HMPSoCs relative to their static counterparts, significantly reducing average- and worst-case inference latencies and enhancing the speed-up of EE DNN compared to the best single-processor inference. On average, the worst-case inference latency decreased by 24.8% across three commonly used EE DNNs, providing latency comparable to astatic model without compromising accuracy on an RK3399PROHMPSoC.
Furthermore, existing CPU- or GPU-only implementations of EE DNN inference under-utilize the HMPSoC. To address this limitation, we propose a cooperative and parallel CPU-GPU execution approach1 for EE DNN inference that effectively distributes computations across all HMPSoC processors, minimizing latency variations. Our approach allows EE DNNs to achieve reduced average latency on HMPSoCs relative to their static counterparts, significantly reducing average- and worst-case inference latencies and enhancing the speed-up of EE DNN compared to the best single-processor inference. On average, the worst-case inference latency decreased by 24.8% across three commonly used EE DNNs, providing latency comparable to astatic model without compromising accuracy on an RK3399PROHMPSoC.
| Original language | English |
|---|---|
| Title of host publication | 2025 IEEE International Conference on Edge Computing and Communications, IEEE EDGE 2025 |
| Editors | Rong N. Chang, Carl K. Chang, Jingwei Yang, Nimanthi Atukorala, Dan Chen, Sumi Helal, Sasu Tarkoma, Qiang He, Tevfik Kosar, Claudio Ardagna, Feras Awaysheh, Volker Hilt, Yogesh Simmhan |
| Publisher | Institute of Electrical and Electronics Engineers |
| Pages | 75-82 |
| Number of pages | 8 |
| ISBN (Electronic) | 979-8-3315-5559-7 |
| DOIs | |
| Publication status | Published - 18 Aug 2025 |
| Event | 2025 IEEE International Conference On Edge Computing & Communications - Helsinki, Finland Duration: 7 Jul 2025 → 12 Jul 2025 https://services.conferences.computer.org/2025/edge/ |
Conference
| Conference | 2025 IEEE International Conference On Edge Computing & Communications |
|---|---|
| Abbreviated title | IEEE EDGE 2025 |
| Country/Territory | Finland |
| City | Helsinki |
| Period | 7/07/25 → 12/07/25 |
| Internet address |
Keywords
- Edge Computing
- Dynamic Networks
- Early- Exit Networks, Inference Acceleration, Embedded Systems
- , Inference Acceleration
- Embedded Systems
- Inference Acceleration
- Early-Exit Networks
Fingerprint
Dive into the research topics of 'Early-Exit DNN Inference on HMPSoCs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver