ReferTrack: Referring Then Tracking for Embodied Visual Tracking
TL;DR AI
2 min readKey summary
Researchers introduced ReferTrack, a single-camera embodied visual tracking method that first grounds a language-specified target in bounding boxes, then predicts tracking waypoints.
The system uses temporal box history and the Refer-QA dataset to improve target identification and motion tracking.
ReferTrack achieved state-of-the-art results on EVT-Bench and demonstrated strong performance in real robot deployments.
The work shows that supervised, image-grounded tracking can outperform more abstract reasoning approaches and rival multi-camera systems in hard identification tasks.
