摘要
Large Language Models (LLM) have demonstrated outstanding performance in natural language processing tasks. However, their extremely large parameter scales pose a significant challenge because the limited capacity of GPU memory becomes a performance bottleneck for inference tasks. To address this issue in the context of LLM inference services, this study proposes AdaptiveLLM, which enables the adaptive selection of offloading strategies between tensor swapping and tensor recomputation based on the characteristics of inference task workloads. To evaluate the characteristics of inference task workloads, AdaptiveLLM establishes a black-box Machine Learning (ML) model through an operator-level computational complexity analysis to predict the overhead of tensor recomputation. It also predicts the overhead of tensor swapping by conducting a fine-grained analysis of KV Cache memory usage. For the adaptive selection of offloading strategies, AdaptiveLLM designs a cost-aware memory optimization strategy specifically for the pre-emption scheduling phase. When GPU memory is insufficient, it opts for the offloading approach with a lower overhead. For the initiation scheduling phase, it devises a fairness-based user-request scheduling strategy. When GPU memory is available, it schedules more user requests in accordance with the principle of fairness. Experimental results indicate that, compared with currently widely used LLM inference benchmark frameworks, AdaptiveLLM achieves an overall increase in throughput while reducing the average weighted turnaround time, thereby realizing fair scheduling.
| 投稿的翻译标题 | Inference Optimization for Large Models Based on Adaptive Tensor Swapping and Recomputation |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 27-36 |
| 页数 | 10 |
| 期刊 | Jisuanji Gongcheng/Computer Engineering |
| 卷 | 51 |
| 期 | 10 |
| DOI | |
| 出版状态 | 已出版 - 15 10月 2025 |
关键词
- Fairness
- Inference
- Large Language Models (LLM)
- Tensor recomputation
- Tensor swapping
- Throughput
指纹
探究 '基于自适应张量交换和重算的大模型推理优化' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver