Process Rewards with Learned Reliability
TL;DR AI
2 min readKey summary
Researchers introduced BetaPRM, a process reward model that predicts both step success and how reliable each prediction is using a Beta-Binomial formulation.
The learned reliability signal was used for Adaptive Computation Allocation, helping the model spend more compute only where it is needed.
In Best-of-N reasoning, BetaPRM improved selection quality while cutting token use by up to 33.57%.
The work shows that step-level reward models can better balance accuracy and efficiency when they estimate their own confidence.
