TY - JOUR
T1 - Split Learning Over Wireless Networks
T2 - Parallel Design and Resource Management
AU - Wu, Wen
AU - Li, Mushu
AU - Qu, Kaige
AU - Zhou, Conghao
AU - Shen, Xuemin
AU - Zhuang, Weihua
AU - Li, Xu
AU - Shi, Weisen
N1 - Publisher Copyright:
© 1983-2012 IEEE.
PY - 2023/4/1
Y1 - 2023/4/1
N2 - Split learning (SL) is a collaborative learning framework, which can train an artificial intelligence (AI) model between a device and an edge server by splitting the AI model into a device-side model and a server-side model at a cut layer. The existing SL approach conducts the training process sequentially across devices, which incurs significant training latency especially when the number of devices is large. In this paper, we design a novel SL scheme to reduce the training latency, named Cluster-based Parallel SL (CPSL) which conducts model training in a 'first-parallel-then-sequential' manner. Specifically, the CPSL is to partition devices into several clusters, parallelly train device-side models in each cluster and aggregate them, and then sequentially train the whole AI model across clusters, thereby parallelizing the training process and reducing training latency. Furthermore, we propose a resource management algorithm to minimize the training latency of CPSL considering device heterogeneity and network dynamics in wireless networks. This is achieved by stochastically optimizing the cut layer selection, device clustering, and radio spectrum allocation. The proposed two-timescale algorithm can jointly make the cut layer selection decision in a large timescale and device clustering and radio spectrum allocation decisions in a small timescale. Extensive simulation results on non-independent and identically distributed data demonstrate that the proposed solution can greatly reduce the training latency as compared with the existing SL benchmarks, while adapting to network dynamics.
AB - Split learning (SL) is a collaborative learning framework, which can train an artificial intelligence (AI) model between a device and an edge server by splitting the AI model into a device-side model and a server-side model at a cut layer. The existing SL approach conducts the training process sequentially across devices, which incurs significant training latency especially when the number of devices is large. In this paper, we design a novel SL scheme to reduce the training latency, named Cluster-based Parallel SL (CPSL) which conducts model training in a 'first-parallel-then-sequential' manner. Specifically, the CPSL is to partition devices into several clusters, parallelly train device-side models in each cluster and aggregate them, and then sequentially train the whole AI model across clusters, thereby parallelizing the training process and reducing training latency. Furthermore, we propose a resource management algorithm to minimize the training latency of CPSL considering device heterogeneity and network dynamics in wireless networks. This is achieved by stochastically optimizing the cut layer selection, device clustering, and radio spectrum allocation. The proposed two-timescale algorithm can jointly make the cut layer selection decision in a large timescale and device clustering and radio spectrum allocation decisions in a small timescale. Extensive simulation results on non-independent and identically distributed data demonstrate that the proposed solution can greatly reduce the training latency as compared with the existing SL benchmarks, while adapting to network dynamics.
KW - Split learning
KW - device clustering
KW - parallel model training
KW - resource management
UR - https://www.scopus.com/pages/publications/85148435487
U2 - 10.1109/JSAC.2023.3242704
DO - 10.1109/JSAC.2023.3242704
M3 - 文章
AN - SCOPUS:85148435487
SN - 0733-8716
VL - 41
SP - 1051
EP - 1066
JO - IEEE Journal on Selected Areas in Communications
JF - IEEE Journal on Selected Areas in Communications
IS - 4
ER -