Ω-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

TL;DR AI
2 min readKey summary
Researchers unveiled Omega-QVLA, a training-free post-training quantization method for vision-language-action models.
It uses composite rotation and per-step activation scaling to compress both the language backbone and diffusion action head to W4A4 precision.
On LIBERO, it matched or beat FP16 baselines on Pi 0.5 and GR00T N1.5 while cutting static memory by 71.3%.
Real-world robot tests also showed reliable manipulation, pointing to cheaper and more practical on-device deployment.
