Pass the Baton: Trajectory-Relayed On-Policy Distillation
TL;DR AI
2 min readKey summary
Researchers introduced Relay-OPD, a label-free on-policy distillation method for reasoning models.
When a student starts a bad prefix, Relay-OPD briefly hands control to a teacher model to create a relay trajectory, then resumes the student.
On eight math reasoning benchmarks with Qwen3 models, it beat standard OPD and FastOPD, especially for smaller students.
The method also cut training trajectory length by more than 50%, improving efficiency while boosting performance.
