QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

TL;DR AI
2 min readKey summary
QueenVIS is a new framework that improves video instance segmentation using image-only training.
It enriches Mask2Former object queries with feature-prediction and center-prediction auxiliary losses to make queries more stable and discriminative.
At inference, query propagation and a memory bank help preserve instance identities across frames.
The method outperforms prior image-only baselines on YouTube-VIS and OVIS without training on video clips.
