Advancing Creative Physical Intelligence in Large Multimodal Models
TL;DR AI
2 min readKey summary
Researchers introduced MM-CreativityBench to test whether multimodal models can solve creative tool-use tasks in visually rich, physically constrained scenes.
The benchmark shows current models often overlook relevant objects, inspect scenes too superficially, and hallucinate object attributes.
To address this, the team proposed affordance-grounded alignment, combining preference optimization with affordance supervision.
This training approach improves grounded exploration and helps reduce hallucinations in creative visual problem-solving.
