Skip to main navigation Skip to search Skip to main content

Guiding reinforcement learning with shaping rewards provided by the vision–language model

  • Beihang University
  • CAS - Institute of Automation

Research output: Contribution to journalArticlepeer-review

Abstract

Enabling robots to learn manipulation tasks is a practical engineering application, and reinforcement learning is one of the key artificial intelligence methods to achieve it, but reinforcement learning always faces the trade-off between the ease of designing reward functions and the ease of learning from rewards. Reward shaping provides a solution and recent works have designed rewards using images and language descriptions, which is a simple and convenient shaping method for non-expert users. However, some of them only adopt the pretrained model without fine-tuning, while others train the reward model with absolute score labels, which all have difficulties in capturing the spatial relationships within the images, so the performance of reward shaping models is limited. In this work, we propose a novel reward shaping method to generate additional rewards from task descriptions and scene images. We utilize the pretrained vision–language model as backbone for efficient cross-modal information fusion and design a downstream task trained with pair-wise comparison for reward shaping. Extensive experiments are conducted to demonstrate the effectiveness of each component in the method. The approach is validated in the Meta-World environment and the results demonstrate that it outperforms standard reinforcement learning and the existing work in terms of the policy learning efficiency.

Original languageEnglish
Article number111004
JournalEngineering Applications of Artificial Intelligence
Volume155
DOIs
StatePublished - 1 Sep 2025

Keywords

  • Guided reinforcement learning
  • Policy learning efficiency
  • Reward shaping
  • Robotic manipulation task
  • Vision–language model

Fingerprint

Dive into the research topics of 'Guiding reinforcement learning with shaping rewards provided by the vision–language model'. Together they form a unique fingerprint.

Cite this