Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
TL;DR AI
2 min readKey summary
Researchers introduced PRISM, a three-stage pipeline that adds a distribution-alignment step between supervised fine-tuning and reinforcement learning for multimodal models.
PRISM uses a black-box on-policy distillation stage with a policy-versus-MoE discriminator game to reduce the train-test mismatch caused by SFT.
On Qwen3-VL, the method improved downstream accuracy across multiple RL methods, including GRPO, DAPO, and GSPO.
The approach showed measurable gains on multimodal reasoning benchmarks, suggesting a practical way to stabilize RLVR training.
