Skip to main navigation Skip to search Skip to main content

Cross-Modal Proxy Prompt Alignment for Fine-Grained Image Classification

  • Jin Han
  • , Junlin Hu*
  • *Corresponding author for this work
  • Beihang University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Fine-grained image classification remains challenging due to subtle inter-class differences, large intra-class variations, and lack of explicit semantic guidance in purely visual representations. Existing methods often struggle to construct stable and discriminative class prototypes, especially when textual annotations are unavailable. In this paper, we propose a cross-modal proxy prompt alignment (CPPA) method that introduces proxy prompts and obtains learnable class prototypes that function as auxiliary textual information to encode and inject fine-grained semantic cues into the model. The proxy prompts are directly optimized via multiple image-text alignment objectives, enabling them to autonomously acquire class-specific semantic structure without requiring additional annotations. Building on the aligned proxy prompts, we introduce a fusion mechanism that enables deep bidirectional interaction between visual features and textual proxies, effectively integrating local visual details with class-level semantic cues. Through a two-stage training procedure, our CPPA learns an enriched and more discriminative feature space while maintaining training stability. Experiments on CUB-200-2011, Stanford Dogs, and NABirds datasets show that our CPPA consistently enhances fine-grained image classification performance, demonstrating the effectiveness of proxy prompt alignment and cross-modal fusion for fine-grained image classification.

Original languageEnglish
Title of host publication2026 IEEE Conference on Artificial Intelligence, CAI 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1616-1621
Number of pages6
ISBN (Electronic)9798331560393
DOIs
StatePublished - 2026
Event4th IEEE Conference on Artificial Intelligence, CAI 2026 - Granada, Spain
Duration: 8 May 202610 May 2026

Publication series

Name2026 IEEE Conference on Artificial Intelligence, CAI 2026

Conference

Conference4th IEEE Conference on Artificial Intelligence, CAI 2026
Country/TerritorySpain
CityGranada
Period8/05/2610/05/26

Fingerprint

Dive into the research topics of 'Cross-Modal Proxy Prompt Alignment for Fine-Grained Image Classification'. Together they form a unique fingerprint.

Cite this