GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
TL;DR AI
2 min readKey summary
Researchers introduced GRASP, a 290K-pair dataset built from 46K videos to train models on social reasoning in multi-person non-verbal interactions.
GRASP links high-level social questions to gaze and deictic gesture events, helping models infer who is interacting with whom.
They also propose Social Grounding Reward to improve grounding, plus GRASP-Bench to evaluate performance on the task.
The work targets a major weakness in video AI: understanding subtle social cues for better video QA and multimodal reasoning.
