Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs

TL;DR AI
2 min readKey summary
Researchers found many Video-LLMs perform near chance on simple left/right/up/down motion questions.
The failure comes from a binding gap: models can detect motion signals but struggle to map them to the correct answer.
They introduced MoDirect and DeltaDirect to improve motion-direction detection and generalization.
The methods boost directional accuracy without hurting broader video understanding, highlighting a basic perception flaw in video AI.
