Switch language한국어
Back to the list

PEEK: Picking Essential frames via Efficient Knowledge Distillation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced PEEK, a lightweight dynamic frame-sampling method for video captioning.

  2. PEEK distills frame-relevance rankings from a stronger teacher into a temporal model using only visual input.

  3. It outperformed prior adaptive sampling methods on ActivityNet Captions and MSR-VTT, especially with just one or two frames.

  4. The approach improves caption quality at low frame budgets while adding far less captioning time than competing methods.

Read the original