Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO
TL;DR AI
2 min readKey summary
Researchers propose Target Decoupling for multi-timescale actor-critic methods to improve training stability.
The approach removes routing aggregation from the actor and moves multi-horizon fitting to the critic.
It addresses two failure modes, including surrogate objective hacking and policy collapse.
On LunarLander-v2, the decoupled PPO design shows more reliable performance on delayed-reward tasks.
