CAST: Game Solvers as Turn-Level Teachers for LLM Agents
TL;DR AI
2 min readKey summary
CAST is a new training method for LLM agents that turns game-solver value changes into dense, turn-level supervision for RL with verifiable rewards.
It improves credit assignment in long-horizon tasks by giving cheaper intermediate signals instead of relying only on sparse final rewards.
The paper reports better performance than trained baselines on Sokoban, Minesweeper, and Rush Hour.
CAST also shows strong zero-shot results on ALFWorld and WebShop.
