Not Only Where, But When: Temporal Scheduling for RLVR
TL;DR AI
2 min readKey summary
Researchers proposed temporal scheduling for credit allocation in RLVR, shifting from token-specific behaviors early in training to broader optimization later.
Using trajectory percentiles to separate behaviors worked well and improved learning stability.
The approach preserved policy diversity and boosted results on math and general reasoning benchmarks.
