Skip to main navigation Skip to search Skip to main content

AIQoSer: Building the efficient Inference-QoS for AI Services

  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The AI inspired methods have entirely changed the network QoS landscape and brought better demand-guided experiences for the end-users. However, the increasing demands of satisfactory experiences require larger AI models, whose inference efficiency becomes the non-negligible drawback in the time-sensitive network QoS. In this work, we defined this challenge as the inference-QoS (iQoS) problem of the network QoS itself, which balances inference efficiency and performance for AI services. We design a unified iQoS metric to evaluate the AI-enhanced QoS frameworks with considerations on model performance, inference latency, and input scale. Then, we propose a two-stage pipeline as the exemplar for leveraging the iQoS metric in QoS-aware AI services: (i) enhance reconstruction ability, pretraining masked autoencoder extracts intrinsic data correlations by multi-scale masking; (ii) improve inference efficiency, forecasting masked decoder uses the data scale pruning in terms of spatial and temporal dimension for prediction. Comprehensive experiments on our method demonstrate its superior inference latency and overwhelming traffic matrix prediction performance.

Original languageEnglish
Title of host publication2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9781665468244
DOIs
StatePublished - 2022
Event30th IEEE/ACM International Symposium on Quality of Service, IWQoS 2022 - Oslo, Norway
Duration: 10 Jun 202212 Jun 2022

Publication series

Name2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022

Conference

Conference30th IEEE/ACM International Symposium on Quality of Service, IWQoS 2022
Country/TerritoryNorway
CityOslo
Period10/06/2212/06/22

Keywords

  • AI Service
  • Data Pruning
  • Performance-Efficiency Balance
  • QoS
  • Traffic Matrix Prediction

Fingerprint

Dive into the research topics of 'AIQoSer: Building the efficient Inference-QoS for AI Services'. Together they form a unique fingerprint.

Cite this