CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
TL;DR AI
2 min readKey summary
Researchers introduced CEPO, a contrastive self-distillation method for reinforcement learning with verifiable rewards.
CEPO uses rejected rollouts as a negative teacher to distinguish decisive reasoning steps from filler tokens.
The method preserves training guarantees, improves credit assignment, and avoids extra sampling cost.
It outperformed prior approaches such as GRPO, OPSD, and SDPO on multimodal math benchmarks.
