跳到主要导航 跳到搜索 跳到主要内容

InstructTrack: Language-Guided Multi-Object Tracking with Semantic-Aware Association

  • Zishun Zhou
  • , Shuai Wang*
  • , Hao Sheng
  • , Dazhi Yang
  • , Sentan Li
  • , Da Yang
  • , Zhenglong Cui
  • *此作品的通讯作者
  • Beihang University
  • Macao Polytechnic University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Locating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision-language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches.

源语言英语
主期刊名Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
编辑Tat-Seng Chua, Lai-Kuan Wong, Chee Seng Chan, Jinhui Tang, Chong-Wah Ngo, Klaus Schoeffmann, Jiaying Liu, Yo-Sung Ho
出版商Association for Computing Machinery, Inc
ISBN(电子版)9798400720055
DOI
出版状态已出版 - 6 12月 2025
活动7th ACM International Conference on Multimedia in Asia, MMAsia 2025 - Kuala Lumpur, 马来西亚
期限: 9 12月 202512 12月 2025

出版系列

姓名Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025

会议

会议7th ACM International Conference on Multimedia in Asia, MMAsia 2025
国家/地区马来西亚
Kuala Lumpur
时期9/12/2512/12/25

学术指纹

探究 'InstructTrack: Language-Guided Multi-Object Tracking with Semantic-Aware Association' 的科研主题。它们共同构成独一无二的学术指纹。

引用此