Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

TL;DR AI
2 min readKey summary
Researchers introduced ReBel, a process-level reinforcement learning method for long-horizon LLM agents in partially observable tasks.
ReBel models structured beliefs, adds belief-consistency supervision, and groups trajectories by similar belief states to make training signals denser and less noisy.
On ALFWorld and WebShop, it reportedly beats the episode-level GRPO baseline by up to 20.4 percentage points and improves sample efficiency by 2.1x.
The approach addresses delayed-reward credit assignment without requiring step-level labels, making it more practical for long-horizon agents.
