Switch language한국어
Back to the list

EarlyTom: Early Token Compression Completes Fast Video Understanding

TL;DR AI

Key summary

2 min read
  1. Researchers introduced EarlyTom, a training-free token compression framework for video LLMs.

  2. Unlike post-encoder pruning, it compresses visual tokens inside the vision encoder with decoupled spatial token selection.

  3. On LLaVA-OneVision-7B, it cuts time-to-first-token by up to 2.65x and FLOPs by up to 61%.

  4. Accuracy stays close to the full-token baseline, making video understanding models faster and cheaper to deploy.

Read the original