Switch language한국어
Back to the list

Not Only Where, But When: Temporal Scheduling for RLVR

TL;DR AI

Key summary

2 min read
  1. Researchers proposed temporal scheduling for credit allocation in RLVR, shifting from token-specific behaviors early in training to broader optimization later.

  2. Using trajectory percentiles to separate behaviors worked well and improved learning stability.

  3. The approach preserved policy diversity and boosted results on math and general reasoning benchmarks.

Read the original