跳到主要导航 跳到搜索 跳到主要内容

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

  • Cheng Chen
  • , Yunpeng Zhai
  • , Yifan Zhao*
  • , Jinyang Gao
  • , Bolin Ding
  • , Jia Li*
  • *此作品的通讯作者
  • Beihang University
  • Alibaba Group Holding Ltd.

科研成果: 期刊稿件会议文章同行评审

摘要

In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs.

源语言英语
页(从-至)3826-3835
页数10
期刊Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOI
出版状态已出版 - 2025
活动2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, 美国
期限: 11 6月 202515 6月 2025

指纹

探究 'Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning' 的科研主题。它们共同构成独一无二的指纹。

引用此