Skip to main navigation Skip to search Skip to main content

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking

  • Wu Kui
  • , Hao Chen
  • , Jinzhu Han
  • , Haijun Liu
  • , Churan Wang
  • , Wang Yizhou
  • , Zhoujun Li
  • , Si Liu
  • , Fangwei Zhong*
  • *Corresponding author for this work
  • Beihang University
  • City University of Macau
  • Beijing Normal University
  • Minzu University of China
  • Peking University

Research output: Contribution to journalArticlepeer-review

Abstract

Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platform to another. Existing approaches typically train separate models for each embodiment, leading to poor scalability and limited generalization. To address this, we propose AdaTracker, an adaptive in-context policy learning framework that robustly tracks targets on diverse robot morphologies. Our key insight is to explicitly model embodiment-specific constraints through an Embodiment Context Encoder, which infers embodiment-specific constraints from history. This contextual representation dynamically modulates a Context-Aware Policy, enabling it to infer optimal control actions for unseen embodiments in a zero-shot manner. To enhance robustness, we introduce two auxiliary objectives to ensure accurate context identification and temporal consistency. Experiments in both simulation and the real world demonstrate that AdaTracker significantly outperforms state-of-the-art methods in cross-embodiment generalization, sample efficiency, and zero-shot adaptation.

Original languageEnglish
Pages (from-to)8308-8314
Number of pages7
JournalIEEE Robotics and Automation Letters
Volume11
Issue number7
DOIs
StatePublished - Jul 2026

Keywords

  • Embodied visual tracking
  • cross-embodiment generalization
  • in-context reinforcement learning

Fingerprint

Dive into the research topics of 'AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking'. Together they form a unique fingerprint.

Cite this