TY - GEN
T1 - InstructTrack
T2 - 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
AU - Zhou, Zishun
AU - Wang, Shuai
AU - Sheng, Hao
AU - Yang, Dazhi
AU - Li, Sentan
AU - Yang, Da
AU - Cui, Zhenglong
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/12/6
Y1 - 2025/12/6
N2 - Locating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision-language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches.
AB - Locating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision-language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches.
KW - context gating
KW - instruction-driven tracking
KW - multi-object tracking
KW - person re-identification
KW - vision-language models
UR - https://www.scopus.com/pages/publications/105025107337
U2 - 10.1145/3743093.3771047
DO - 10.1145/3743093.3771047
M3 - 会议稿件
AN - SCOPUS:105025107337
T3 - Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
BT - Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025
A2 - Chua, Tat-Seng
A2 - Wong, Lai-Kuan
A2 - Chan, Chee Seng
A2 - Tang, Jinhui
A2 - Ngo, Chong-Wah
A2 - Schoeffmann, Klaus
A2 - Liu, Jiaying
A2 - Ho, Yo-Sung
PB - Association for Computing Machinery, Inc
Y2 - 9 December 2025 through 12 December 2025
ER -