ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
TL;DR AI
2 min readKey summary
Researchers proposed ReToken, a learnable retrieval token that selects relevant visual tokens from a cached long visual context.
Despite being trained on a small image-QA dataset, it improved performance on image and video retrieval benchmarks.
The method reduces distractors and memory use, enabling training and long-video inference on a single H100 GPU.
