TY - GEN
T1 - Federated Semantic Synergistic Alignment for Food Image Classification under Class Imbalance
AU - Zhang, Ran
AU - Chai, Minkang
AU - Qian, Zheng
AU - Wei, Lu
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2026/3/20
Y1 - 2026/3/20
N2 - Federated Learning (FL) enables distributed model training while preserving data privacy, making it highly applicable to privacy-sensitive tasks such as food image classification. However, challenges such as class imbalance, vacant classes, and high computational and communication costs hinder its performance. This study proposes a Federated Semantic Synergistic Alignment model (FedSSA), which leverages CLIP's multimodal semantic knowledge to enhance food image classification under label-skewed scenarios. FedSSA comprises three core stages: Multimodal Semantic Knowledge Construction (MSKC), which utilizes CLIP to generate multidimensional semantic features encompassing appearance, texture, context, and other attributes, providing prior knowledge for vacant classes; Semantic-Guided Knowledge Transfer (SGKT), which aligns local model outputs with global semantic features via KL divergence to improve classification of vacant classes; and Adaptive Feature Projection Optimization (AFPO), which dynamically adjusts the feature space and reduces client-side computational and communication costs through server-side feature dissemination. Experiments on the Food101, ISIAFood200, and VireoFood172 datasets demonstrate that FedSSA significantly outperforms state-of-the-art methods, including FedAvg, MOON, FedMR, and FedVLS, in severe class imbalance scenarios. Moreover, FedSSA achieves a computational complexity of only 1/16th that of CLIP, reduces inference time to 0.15 seconds, and lowers storage requirements to approximately 200 MB, highlighting its efficiency and robustness.
AB - Federated Learning (FL) enables distributed model training while preserving data privacy, making it highly applicable to privacy-sensitive tasks such as food image classification. However, challenges such as class imbalance, vacant classes, and high computational and communication costs hinder its performance. This study proposes a Federated Semantic Synergistic Alignment model (FedSSA), which leverages CLIP's multimodal semantic knowledge to enhance food image classification under label-skewed scenarios. FedSSA comprises three core stages: Multimodal Semantic Knowledge Construction (MSKC), which utilizes CLIP to generate multidimensional semantic features encompassing appearance, texture, context, and other attributes, providing prior knowledge for vacant classes; Semantic-Guided Knowledge Transfer (SGKT), which aligns local model outputs with global semantic features via KL divergence to improve classification of vacant classes; and Adaptive Feature Projection Optimization (AFPO), which dynamically adjusts the feature space and reduces client-side computational and communication costs through server-side feature dissemination. Experiments on the Food101, ISIAFood200, and VireoFood172 datasets demonstrate that FedSSA significantly outperforms state-of-the-art methods, including FedAvg, MOON, FedMR, and FedVLS, in severe class imbalance scenarios. Moreover, FedSSA achieves a computational complexity of only 1/16th that of CLIP, reduces inference time to 0.15 seconds, and lowers storage requirements to approximately 200 MB, highlighting its efficiency and robustness.
KW - CLIP
KW - Class Imbalance
KW - Federated Learning
KW - Food Image Classification
KW - Multimodal Semantic Knowledge
UR - https://www.scopus.com/pages/publications/105035813039
U2 - 10.1145/3783862.3783886
DO - 10.1145/3783862.3783886
M3 - 会议稿件
AN - SCOPUS:105035813039
T3 - ICCSIT 2025 - Proceedings of the 2025 18th International Conference on Computer Science and Information Technology
SP - 189
EP - 195
BT - ICCSIT 2025 - Proceedings of the 2025 18th International Conference on Computer Science and Information Technology
PB - Association for Computing Machinery, Inc
T2 - 18th International Conference on Computer Science and Information Technology, ICCSIT 2025
Y2 - 27 October 2025 through 29 October 2025
ER -