Switch language한국어
Back to the list

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

TL;DR AI

Key summary

2 min read
  1. Researchers proposed ReToken, a learnable retrieval token that selects relevant visual tokens from a cached long visual context.

  2. Despite being trained on a small image-QA dataset, it improved performance on image and video retrieval benchmarks.

  3. The method reduces distractors and memory use, enabling training and long-video inference on a single H100 GPU.

Read the original