Switch language한국어
Back to the list

QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

TL;DR AI

Key summary

2 min read
  1. QueenVIS is a new framework that improves video instance segmentation using image-only training.

  2. It enriches Mask2Former object queries with feature-prediction and center-prediction auxiliary losses to make queries more stable and discriminative.

  3. At inference, query propagation and a memory bank help preserve instance identities across frames.

  4. The method outperforms prior image-only baselines on YouTube-VIS and OVIS without training on video clips.

Read the original