Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning

TL;DR AI
2 min readKey summary
Researchers introduced ZPRL, a robot adaptation method that keeps a pretrained imitation policy frozen and learns only a residual in a compact latent space.
The approach was tested on eight simulation tasks and four real-world manipulation tasks, where it beat strong baselines and improved sample efficiency.
In real-world trials, ZPRL raised average success by 33.7% over the base imitation policies.
The method offers a safer, more structured alternative to action-space residual fine-tuning for online robot adaptation.
