Abstract
Prompt tuning has emerged as a pivotal technique for adapting pre-trained vision-language models (VLMs) to a wide range of downstream tasks. Recent developments have introduced multimodal learnable prompts to construct task-specific classifiers. However, these methods often exhibit limited generalization to unseen classes, primarily due to fixed prompt designs that are tightly coupled with seen training data and lack adaptability to novel class distributions. To overcome this limitation, we propose Label-Informed Knowledge Integration (LIKI)—a novel framework that harnesses the robust generalizability of textual label semantics to guide the generation of adaptive visual prompts. Rather than directly mapping textual prompts into the visual domain, LIKI utilizes robust text embeddings as a knowledge source to inform the visual prompt optimization. Central to our method is a simple yet effective Label Semantic Integration (LSI) module, which dynamically incorporates knowledge from both seen and unseen labels into the visual prompts. This label-informed prompting strategy imbues the visual encoder with semantic awareness, thereby enhancing the generalization and discriminative capacity of VLMs across diverse scenarios. Extensive experiments demonstrate that LIKI consistently outperforms state-of-the-art approaches in base-to-novel generalization, cross-dataset transfer, and domain generalization tasks, offering a significant advancement in prompt-based VLM adaptation.
| Original language | English |
|---|---|
| Article number | 104614 |
| Journal | Computer Vision and Image Understanding |
| Volume | 263 |
| DOIs | |
| State | Published - Jan 2026 |
Keywords
- Few-shot learning
- Prompt tuning
- Vision-language models
- Zero-shot learning
Fingerprint
Dive into the research topics of 'Label-informed knowledge integration: Advancing visual prompt for VLMs adaptation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver