Switch language한국어
Back to the list

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RefCaptioner, a video captioning system that uses multiple reference images to ground specific visual phrases more faithfully.

  2. The paper defines a new task, multi-reference image-grounded video captioning, focused on factual, phrase-level descriptions rather than generic captions.

  3. RefCaptioner uses a two-stage training pipeline with supervised fine-tuning and reinforcement learning to improve reference selection, grounding accuracy, and distractor rejection.

  4. The authors also release a large training corpus and MRVBench for evaluation, reporting strong results on open-source benchmarks and favorable human judgments.

  5. This pushes video captioning beyond generic descriptions toward captions that are more factual, verifiable, and useful for source-faithful video generation.

Read the original