Switch language한국어
Back to the list

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers propose staged post-training for vision-language models, separating perception, visual reasoning, and textual reasoning.

  2. The method consistently outperforms unified training across multiple models and benchmarks, including WeMath and RealWorldQA.

  3. A key takeaway is that improving perception first can boost accuracy and shorten reasoning traces.

  4. The recipe combines SFT and RL in a curriculum-like pipeline, offering a stronger option for open-weight VLMs.

Read the original