When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

TL;DR AI
2 min readKey summary
Researchers proposed RLAVR to make reinforcement learning with verifiable rewards more label-efficient and stable.
RLAVR combines a small set of actively acquired ground-truth labels with pseudo-labels instead of relying on pseudo-labels alone.
The method uses the Corrective Advantage Gap (CAG) to find high-value samples and CARE as its practical acquisition policy.
Experiments show stronger stability and performance across domains and model scales, addressing a major bottleneck in training reasoning models.
