Switch language한국어
Back to the list

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced IC-VCO, a new training method for vision-language models that reduces multimodal hallucinations.

  2. It contrasts images within a shared context, adds a reliability-gated distillation regularizer, and uses sample editing to create harder negatives.

  3. The approach was tested on five benchmarks and showed strong overall performance, improving reliability in image understanding tasks.

Read the original