TY - GEN
T1 - Cross-Modal Proxy Prompt Alignment for Fine-Grained Image Classification
AU - Han, Jin
AU - Hu, Junlin
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Fine-grained image classification remains challenging due to subtle inter-class differences, large intra-class variations, and lack of explicit semantic guidance in purely visual representations. Existing methods often struggle to construct stable and discriminative class prototypes, especially when textual annotations are unavailable. In this paper, we propose a cross-modal proxy prompt alignment (CPPA) method that introduces proxy prompts and obtains learnable class prototypes that function as auxiliary textual information to encode and inject fine-grained semantic cues into the model. The proxy prompts are directly optimized via multiple image-text alignment objectives, enabling them to autonomously acquire class-specific semantic structure without requiring additional annotations. Building on the aligned proxy prompts, we introduce a fusion mechanism that enables deep bidirectional interaction between visual features and textual proxies, effectively integrating local visual details with class-level semantic cues. Through a two-stage training procedure, our CPPA learns an enriched and more discriminative feature space while maintaining training stability. Experiments on CUB-200-2011, Stanford Dogs, and NABirds datasets show that our CPPA consistently enhances fine-grained image classification performance, demonstrating the effectiveness of proxy prompt alignment and cross-modal fusion for fine-grained image classification.
AB - Fine-grained image classification remains challenging due to subtle inter-class differences, large intra-class variations, and lack of explicit semantic guidance in purely visual representations. Existing methods often struggle to construct stable and discriminative class prototypes, especially when textual annotations are unavailable. In this paper, we propose a cross-modal proxy prompt alignment (CPPA) method that introduces proxy prompts and obtains learnable class prototypes that function as auxiliary textual information to encode and inject fine-grained semantic cues into the model. The proxy prompts are directly optimized via multiple image-text alignment objectives, enabling them to autonomously acquire class-specific semantic structure without requiring additional annotations. Building on the aligned proxy prompts, we introduce a fusion mechanism that enables deep bidirectional interaction between visual features and textual proxies, effectively integrating local visual details with class-level semantic cues. Through a two-stage training procedure, our CPPA learns an enriched and more discriminative feature space while maintaining training stability. Experiments on CUB-200-2011, Stanford Dogs, and NABirds datasets show that our CPPA consistently enhances fine-grained image classification performance, demonstrating the effectiveness of proxy prompt alignment and cross-modal fusion for fine-grained image classification.
UR - https://www.scopus.com/pages/publications/105042089701
U2 - 10.1109/CAI68641.2026.11536342
DO - 10.1109/CAI68641.2026.11536342
M3 - 会议稿件
AN - SCOPUS:105042089701
T3 - 2026 IEEE Conference on Artificial Intelligence, CAI 2026
SP - 1616
EP - 1621
BT - 2026 IEEE Conference on Artificial Intelligence, CAI 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 4th IEEE Conference on Artificial Intelligence, CAI 2026
Y2 - 8 May 2026 through 10 May 2026
ER -