PEEK: Picking Essential frames via Efficient Knowledge Distillation
TL;DR AI
2 min readKey summary
Researchers introduced PEEK, a lightweight dynamic frame-sampling method for video captioning.
PEEK distills frame-relevance rankings from a stronger teacher into a temporal model using only visual input.
It outperformed prior adaptive sampling methods on ActivityNet Captions and MSR-VTT, especially with just one or two frames.
The approach improves caption quality at low frame budgets while adding far less captioning time than competing methods.
