Visual Contrastive Self-Distillation
TL;DR AI
2 min readKey summary
Researchers introduced VCSD, a contrastive self-distillation method for vision-language models.
VCSD compares image-conditioned and image-erased teacher outputs to create a training signal without external teachers or privileged labels.
On ViRL39K, it improved benchmark performance for Qwen3-VL and Qwen3.5 models.
The approach simplifies multimodal training by removing dependence on external supervision while boosting accuracy.
