TY - JOUR
T1 - East
T2 - Efficient and Accurate Secure Inference Framework for Transformer
AU - Ding, Yuanchao
AU - Guo, Hua
AU - Guan, Yewei
AU - Liu, Weixin
AU - Huo, Jiarong
AU - Guan, Zhenyu
AU - Zhang, Xiyong
N1 - Publisher Copyright:
© 2008-2012 IEEE.
PY - 2025
Y1 - 2025
N2 - Transformer has been successfully used in practical applications due to its powerful advantages. However, users’ input is leaked to the model provider during the service. With people’s attention to privacy, privacy-preserving Transformer inference is on the demand of such services. Secure protocols for non-linear functions are crucial in privacy-preserving Transformer inference, which are not well studied. Thus, designing practical secure protocols for non-linear functions is hard but significant to model performance. In this work, we propose a framework East to enable efficient and accurate secure Transformer inference. First, we propose a new oblivious piecewise polynomial evaluation algorithm and apply it to the activation functions, which reduces the runtime and communication of GELU by over 1.5× and 2.5×, compared to prior arts. Second, the secure protocols for softmax and layer normalization are carefully designed to faithfully maintain the desired functionality. Third, several optimizations are conducted in detail to enhance the overall efficiency. We applied East to BERT and the results show that the inference accuracy remains consistent with the plaintext inference without fine-tuning. Compared to Iron, we achieve about 1.8× lower communication within 1.2× lower runtime.
AB - Transformer has been successfully used in practical applications due to its powerful advantages. However, users’ input is leaked to the model provider during the service. With people’s attention to privacy, privacy-preserving Transformer inference is on the demand of such services. Secure protocols for non-linear functions are crucial in privacy-preserving Transformer inference, which are not well studied. Thus, designing practical secure protocols for non-linear functions is hard but significant to model performance. In this work, we propose a framework East to enable efficient and accurate secure Transformer inference. First, we propose a new oblivious piecewise polynomial evaluation algorithm and apply it to the activation functions, which reduces the runtime and communication of GELU by over 1.5× and 2.5×, compared to prior arts. Second, the secure protocols for softmax and layer normalization are carefully designed to faithfully maintain the desired functionality. Third, several optimizations are conducted in detail to enhance the overall efficiency. We applied East to BERT and the results show that the inference accuracy remains consistent with the plaintext inference without fine-tuning. Compared to Iron, we achieve about 1.8× lower communication within 1.2× lower runtime.
KW - Privacy-preserving inference
KW - homomorphic encryption
KW - secure multi-party computation
KW - transformer
UR - https://www.scopus.com/pages/publications/105007614589
U2 - 10.1109/TSC.2025.3577491
DO - 10.1109/TSC.2025.3577491
M3 - 文章
AN - SCOPUS:105007614589
SN - 1939-1374
VL - 18
SP - 2038
EP - 2046
JO - IEEE Transactions on Services Computing
JF - IEEE Transactions on Services Computing
IS - 4
ER -