跳到主要导航 跳到搜索 跳到主要内容

AIQoSer: Building the efficient Inference-QoS for AI Services

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The AI inspired methods have entirely changed the network QoS landscape and brought better demand-guided experiences for the end-users. However, the increasing demands of satisfactory experiences require larger AI models, whose inference efficiency becomes the non-negligible drawback in the time-sensitive network QoS. In this work, we defined this challenge as the inference-QoS (iQoS) problem of the network QoS itself, which balances inference efficiency and performance for AI services. We design a unified iQoS metric to evaluate the AI-enhanced QoS frameworks with considerations on model performance, inference latency, and input scale. Then, we propose a two-stage pipeline as the exemplar for leveraging the iQoS metric in QoS-aware AI services: (i) enhance reconstruction ability, pretraining masked autoencoder extracts intrinsic data correlations by multi-scale masking; (ii) improve inference efficiency, forecasting masked decoder uses the data scale pruning in terms of spatial and temporal dimension for prediction. Comprehensive experiments on our method demonstrate its superior inference latency and overwhelming traffic matrix prediction performance.

源语言英语
主期刊名2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9781665468244
DOI
出版状态已出版 - 2022
活动30th IEEE/ACM International Symposium on Quality of Service, IWQoS 2022 - Oslo, 挪威
期限: 10 6月 202212 6月 2022

出版系列

姓名2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022

会议

会议30th IEEE/ACM International Symposium on Quality of Service, IWQoS 2022
国家/地区挪威
Oslo
时期10/06/2212/06/22

学术指纹

探究 'AIQoSer: Building the efficient Inference-QoS for AI Services' 的科研主题。它们共同构成独一无二的学术指纹。

引用此