TY - JOUR
T1 - PPMGS
T2 - An efficient and effective solution for distributed privacy-preserving semi-supervised learning
AU - Li, Zhi
AU - Li, Chaozhuo
AU - Li, Zhoujun
AU - Weng, Jian
AU - Huang, Feiran
AU - Zhou, Zhibo
N1 - Publisher Copyright:
© 2024
PY - 2024/9
Y1 - 2024/9
N2 - Recently, distributed semi-supervised learning has attracted increasing research attention due to its tremendous practical value. A promising distributed semi-supervised learning method should not only achieve desirable classification performance but also protect data privacy in distributed scenarios. Existing approaches typically capture the similarities between data instances with privacy-preserving computations. This paradigm introduces extra computation and heuristic changes to the algorithm, resulting in sub-optimal solutions that are time-consuming. In current distributed semi-supervised learning, instance similarities are widely used to capture the underlying manifold or guide label propagation. This paper emphasizes that instance similarities are not necessary because the structure of data connections can be estimated using coarser-grained information. We propose a Privacy-preserving Mixture-distribution based Graph Smoothing (PPMGS) model for distributed privacy-preserving semi-supervised learning. Our motivation is to construct a graph based on a Gaussian mixture distribution instead of individual data instances, which better captures the underlying data distribution and improves model efficiency. PPMGS includes a privacy-preserving expectation-maximization (EM) phase to estimate the Gaussian mixture distribution depicting the input data and a mixture-distribution-based graph smoothing algorithm to learn a distribution-based classifier by fitting a few labeled samples. Experimental results show that the proposed PPMGS achieves 5%-10% higher accuracy and macro-F1 than state-of-the-art privacy-preserving semi-supervised learning methods. In terms of efficiency, it reduces time cost by 97% and communication cost by 96% in the most complex dataset. The numerical results demonstrate that our proposal outperforms state-of-the-art baselines in both efficiency and effectiveness.
AB - Recently, distributed semi-supervised learning has attracted increasing research attention due to its tremendous practical value. A promising distributed semi-supervised learning method should not only achieve desirable classification performance but also protect data privacy in distributed scenarios. Existing approaches typically capture the similarities between data instances with privacy-preserving computations. This paradigm introduces extra computation and heuristic changes to the algorithm, resulting in sub-optimal solutions that are time-consuming. In current distributed semi-supervised learning, instance similarities are widely used to capture the underlying manifold or guide label propagation. This paper emphasizes that instance similarities are not necessary because the structure of data connections can be estimated using coarser-grained information. We propose a Privacy-preserving Mixture-distribution based Graph Smoothing (PPMGS) model for distributed privacy-preserving semi-supervised learning. Our motivation is to construct a graph based on a Gaussian mixture distribution instead of individual data instances, which better captures the underlying data distribution and improves model efficiency. PPMGS includes a privacy-preserving expectation-maximization (EM) phase to estimate the Gaussian mixture distribution depicting the input data and a mixture-distribution-based graph smoothing algorithm to learn a distribution-based classifier by fitting a few labeled samples. Experimental results show that the proposed PPMGS achieves 5%-10% higher accuracy and macro-F1 than state-of-the-art privacy-preserving semi-supervised learning methods. In terms of efficiency, it reduces time cost by 97% and communication cost by 96% in the most complex dataset. The numerical results demonstrate that our proposal outperforms state-of-the-art baselines in both efficiency and effectiveness.
KW - Distributed privacy preserving data mining
KW - Mixture distribution modeling
KW - Semi-supervised learning
UR - https://www.scopus.com/pages/publications/85196169320
U2 - 10.1016/j.ins.2024.120934
DO - 10.1016/j.ins.2024.120934
M3 - 文章
AN - SCOPUS:85196169320
SN - 0020-0255
VL - 678
JO - Information Sciences
JF - Information Sciences
M1 - 120934
ER -