GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

TL;DR AI
2 min readKey summary
Researchers introduced GesVLA, a gesture-aware vision-language-action model for robot manipulation in cluttered scenes.
GesVLA fuses gesture representations with vision and language, using synthetic gesture-on-image data and a two-stage training pipeline.
The system improves target grounding and interaction performance on block manipulation, product selection, and produce selection tasks.
The results suggest gestures can reduce instruction ambiguity and boost real-world robot task performance.
