Agent Skills Should Go Beyond Text: The Case for Visual Skills
TL;DR AI
2 min readKey summary
The paper proposes a multimodal skill framework for visual-centric agents that combines textual logic with visual support.
It stores reusable knowledge as static priors, dynamic visual working memory, and interleaved visual skills to capture spatial and state-dependent context.
An automatic pipeline converts task trajectories into these multimodal skills.
The authors report that these skills outperform text-only skills on GUI and other visual-centric benchmarks.
