TY - GEN
T1 - AIQoSer
T2 - 30th IEEE/ACM International Symposium on Quality of Service, IWQoS 2022
AU - Li, Jianxin
AU - Zhu, Tianchen
AU - Zhou, Haoyi
AU - Sun, Qingyun
AU - Jiang, Chunyang
AU - Zhang, Shuai
AU - Hu, Chunming
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - The AI inspired methods have entirely changed the network QoS landscape and brought better demand-guided experiences for the end-users. However, the increasing demands of satisfactory experiences require larger AI models, whose inference efficiency becomes the non-negligible drawback in the time-sensitive network QoS. In this work, we defined this challenge as the inference-QoS (iQoS) problem of the network QoS itself, which balances inference efficiency and performance for AI services. We design a unified iQoS metric to evaluate the AI-enhanced QoS frameworks with considerations on model performance, inference latency, and input scale. Then, we propose a two-stage pipeline as the exemplar for leveraging the iQoS metric in QoS-aware AI services: (i) enhance reconstruction ability, pretraining masked autoencoder extracts intrinsic data correlations by multi-scale masking; (ii) improve inference efficiency, forecasting masked decoder uses the data scale pruning in terms of spatial and temporal dimension for prediction. Comprehensive experiments on our method demonstrate its superior inference latency and overwhelming traffic matrix prediction performance.
AB - The AI inspired methods have entirely changed the network QoS landscape and brought better demand-guided experiences for the end-users. However, the increasing demands of satisfactory experiences require larger AI models, whose inference efficiency becomes the non-negligible drawback in the time-sensitive network QoS. In this work, we defined this challenge as the inference-QoS (iQoS) problem of the network QoS itself, which balances inference efficiency and performance for AI services. We design a unified iQoS metric to evaluate the AI-enhanced QoS frameworks with considerations on model performance, inference latency, and input scale. Then, we propose a two-stage pipeline as the exemplar for leveraging the iQoS metric in QoS-aware AI services: (i) enhance reconstruction ability, pretraining masked autoencoder extracts intrinsic data correlations by multi-scale masking; (ii) improve inference efficiency, forecasting masked decoder uses the data scale pruning in terms of spatial and temporal dimension for prediction. Comprehensive experiments on our method demonstrate its superior inference latency and overwhelming traffic matrix prediction performance.
KW - AI Service
KW - Data Pruning
KW - Performance-Efficiency Balance
KW - QoS
KW - Traffic Matrix Prediction
UR - https://www.scopus.com/pages/publications/85135378281
U2 - 10.1109/IWQoS54832.2022.9812905
DO - 10.1109/IWQoS54832.2022.9812905
M3 - 会议稿件
AN - SCOPUS:85135378281
T3 - 2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022
BT - 2022 IEEE/ACM 30th International Symposium on Quality of Service, IWQoS 2022
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 10 June 2022 through 12 June 2022
ER -