EarlyTom: Early Token Compression Completes Fast Video Understanding
TL;DR AI
2 min readKey summary
Researchers introduced EarlyTom, a training-free token compression framework for video LLMs.
Unlike post-encoder pruning, it compresses visual tokens inside the vision encoder with decoupled spatial token selection.
On LLaVA-OneVision-7B, it cuts time-to-first-token by up to 2.65x and FLOPs by up to 61%.
Accuracy stays close to the full-token baseline, making video understanding models faster and cheaper to deploy.
