AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
TL;DR AI
2 min readKey summary
Researchers introduced AR-VLA, a standalone autoregressive Action Expert for vision-language-action robots.
It keeps its own long-term history, refreshes vision-language inputs, and uses re-anchoring to reduce the impact of stale perception.
In simulated and real robot tests, it produced smoother actions and matched or improved success rates over reactive VLA baselines.
The approach separates fast control from slower perception and reasoning, offering a more scalable path for robot policy training and deployment.
