SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
TL;DR AI
2 min readKey summary
Researchers introduced SOCO, a benchmark for semantic object correspondence in vision foundation models.
SOCO includes consistent part annotations and keypoint descriptions across 100 categories and over 1 million correspondence pairs.
Results show vision backbones learn useful semantic structure but still struggle to transfer correspondences across categories.
Vision-language models localize parts better from text prompts than from visual-reference matching.
SOCO correspondence scores strongly predict downstream performance on segmentation, tracking, pose estimation, and 3D detection, often better than ImageNet classification.
