Open this publication in new window or tab >>2026 (English)In: IEEE Transactions on Cloud Computing, ISSN 2168-7161Article in journal (Refereed) Epub ahead of print
Abstract [en]
Edge intelligence enables the deployment of deep neural network (DNN) inference services across heterogeneous and resource-constrained edge servers. In such environments, inference requests must satisfy strict Quality-of-Service (QoS) requirements in terms of deadline and accuracy, while operating under limited computation, memory bandwidth, and energy budgets. The availability of multiple DNN model variants enables adaptive model selection but significantly increases the complexity of inference assignment decisions. Memory access rate plays a critical role in inference latency and energy consumption; however, existing model-selection approaches in this domain rarely incorporate memory bandwidth explicitly into both latency and energy models. In this work, we study QoS-aware DNN inference assignment in multi-server edge systems with energy constraints, and jointly model computation and memory bandwidth in both latency and energy. To address the NP-hard problem efficiently, we propose a polynomial-time, greedy, preference-based scheduling framework with two complementary assignment strategies, namely Model-oriented and Request-oriented. To analyze the performance of the proposed approaches, we conduct extensive experiments under different system settings using multiple DNN model families, along with comprehensive ablation studies. The results demonstrate the significant effect of DNN model characteristics, including computation, memory bandwidth, and output accuracy, as well as system configuration, on the performance of the strategies. Specifically, Model-oriented approaches achieve better performance for computation-intensive model variants or when inference service times are substantial. In contrast, Request-oriented approaches perform better for bandwidth-intensive model variants and shorter service times. We also demonstrate the notably lower run-time of the proposed Energy-aware-Model-oriented approach across all experiments, where such a lightweight and low-latency solution enables instantaneous decision making, critical for scalable, time-sensitive applications on resource-constrained edge servers.
Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2026
Keywords
DNN Inference, Edge Intelligence, Memory Bandwidth, Model Selection, QoS-aware scheduling
National Category
Computer Sciences Computer Systems
Identifiers
urn:nbn:se:umu:diva-257287 (URN)10.1109/TCC.2026.3714995 (DOI)2-s2.0-105045750295 (Scopus ID)
Funder
The Kempe Foundations, 3161Wallenberg AI, Autonomous Systems and Software Program (WASP)
2026-08-072026-08-072026-08-07