TY - JOUR
T1 - AdaTracker
T2 - Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking
AU - Kui, Wu
AU - Chen, Hao
AU - Han, Jinzhu
AU - Liu, Haijun
AU - Wang, Churan
AU - Yizhou, Wang
AU - Li, Zhoujun
AU - Liu, Si
AU - Zhong, Fangwei
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2026/7
Y1 - 2026/7
N2 - Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platform to another. Existing approaches typically train separate models for each embodiment, leading to poor scalability and limited generalization. To address this, we propose AdaTracker, an adaptive in-context policy learning framework that robustly tracks targets on diverse robot morphologies. Our key insight is to explicitly model embodiment-specific constraints through an Embodiment Context Encoder, which infers embodiment-specific constraints from history. This contextual representation dynamically modulates a Context-Aware Policy, enabling it to infer optimal control actions for unseen embodiments in a zero-shot manner. To enhance robustness, we introduce two auxiliary objectives to ensure accurate context identification and temporal consistency. Experiments in both simulation and the real world demonstrate that AdaTracker significantly outperforms state-of-the-art methods in cross-embodiment generalization, sample efficiency, and zero-shot adaptation.
AB - Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platform to another. Existing approaches typically train separate models for each embodiment, leading to poor scalability and limited generalization. To address this, we propose AdaTracker, an adaptive in-context policy learning framework that robustly tracks targets on diverse robot morphologies. Our key insight is to explicitly model embodiment-specific constraints through an Embodiment Context Encoder, which infers embodiment-specific constraints from history. This contextual representation dynamically modulates a Context-Aware Policy, enabling it to infer optimal control actions for unseen embodiments in a zero-shot manner. To enhance robustness, we introduce two auxiliary objectives to ensure accurate context identification and temporal consistency. Experiments in both simulation and the real world demonstrate that AdaTracker significantly outperforms state-of-the-art methods in cross-embodiment generalization, sample efficiency, and zero-shot adaptation.
KW - Embodied visual tracking
KW - cross-embodiment generalization
KW - in-context reinforcement learning
UR - https://www.scopus.com/pages/publications/105039254236
U2 - 10.1109/LRA.2026.3692075
DO - 10.1109/LRA.2026.3692075
M3 - 文章
AN - SCOPUS:105039254236
SN - 2377-3766
VL - 11
SP - 8308
EP - 8314
JO - IEEE Robotics and Automation Letters
JF - IEEE Robotics and Automation Letters
IS - 7
ER -