Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

TL;DR AI
2 min readKey summary
Researchers introduced IC-VCO, a new training method for vision-language models that reduces multimodal hallucinations.
It contrasts images within a shared context, adds a reliability-gated distillation regularizer, and uses sample editing to create harder negatives.
The approach was tested on five benchmarks and showed strong overall performance, improving reliability in image understanding tasks.
