Skip to main navigation Skip to search Skip to main content

STCTracker: Enhancing sequential temporal consistency in referring multi-object tracking

  • Hong Zhang
  • , Jiabi Zhao
  • , Ding Yuan
  • , Hanyang Liu
  • , Yang Han
  • , Yifan Yang*
  • *Corresponding author for this work
  • Key Laboratory of Precision Opto-Mechatronics Technology (Ministry of Education)
  • Beihang University
  • State Key Laboratory of High-Efficiency Reusable Aerospace Transportation Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Referring Multi-Object Tracking (RMOT) focuses on identifying and tracking objects in video sequences based on natural language queries. A prevalent paradigm employs a two-stage pipeline, combining a tracking-by-detection multi-object tracker with a plug-and-play referring module. The former undertakes the comprehensive tracking of all target entities within the video sequence, while the latter is dedicated to selecting those targets that are semantically aligned with the given query. However, existing multi-object trackers are limited by insufficient perception of motion and occlusion states, while referring modules lack explicit temporal semantic consistency constraints and motion state supervision, leading to temporal ID assignment inconsistency and temporal referring consistency disruption, respectively. In this work, we propose STCTracker, an enhanced Sequential Temporal Consistency RMOT tracker that incorporates two core components to address these limitations: (i) DynaSORT, an optimized StrongSORT variant, integrates three enhanced modules: a Movement Vector Constraint association (MVC) module adapted for robust association, alongside a Bifactor Kalman Filter (BKF) and an Occlusion-aware Exponential Moving Average (OEMA) both tailored for dynamic trajectory updating. (ii) Dual Referring Module (DRM) augments the base referring module with dual auxiliary discriminators: a CLIP-Anchored Categorical Discriminator that validates category-level semantic consistency across temporal frames against query semantics, and an Optical-Flow-Driven State Discriminator for ensuring kinematic state tracking consistency, collectively establishing robust referring grounding. Extensive experiments on Refer-KITTI and KITTI validate the superior performance of our method over existing state-of-the-art solutions.

Original languageEnglish
Article number106073
JournalImage and Vision Computing
Volume173
DOIs
StatePublished - Sep 2026

Keywords

  • KITTI
  • RMOT
  • STCTracker
  • Temporal Consistency

Fingerprint

Dive into the research topics of 'STCTracker: Enhancing sequential temporal consistency in referring multi-object tracking'. Together they form a unique fingerprint.

Cite this