Switch language한국어
Back to the list

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced CEPO, a contrastive self-distillation method for reinforcement learning with verifiable rewards.

  2. CEPO uses rejected rollouts as a negative teacher to distinguish decisive reasoning steps from filler tokens.

  3. The method preserves training guarantees, improves credit assignment, and avoids extra sampling cost.

  4. It outperformed prior approaches such as GRPO, OPSD, and SDPO on multimodal math benchmarks.

Read the original