From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
TL;DR AI
2 min readKey summary
Researchers propose staged post-training for vision-language models, separating perception, visual reasoning, and textual reasoning.
The method consistently outperforms unified training across multiple models and benchmarks, including WeMath and RealWorldQA.
A key takeaway is that improving perception first can boost accuracy and shorten reasoning traces.
The recipe combines SFT and RL in a curriculum-like pipeline, offering a stronger option for open-weight VLMs.
