Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language Models

TL;DR AI
2 min readKey summary
Researchers introduced FAB-G, an attribute-grounded selective reasoning framework for artwork emotion understanding with multimodal large language models.
They also extended EmoArt with 1,400 new salience annotations from art-trained annotators to better link emotions to relevant visual attributes.
The approach outperformed prompting-based baselines on emotion, arousal, valence, and explanation quality.
The work addresses a common multimodal failure mode by making predictions more interpretable and improving cross-dataset transfer.
