VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

TL;DR AI
2 min readKey summary
Researchers introduced VIP, a training-free framework for open-vocabulary semantic segmentation.
VIP goes beyond CLIP-style methods by using a spatially aware vision-language model for dense predictions.
It refines ambiguous text queries with visual cues through visual-guided prompt evolution and alias expansion.
The approach improves accuracy and generalization while adding only low inference overhead.
