Rethinking VLM Representation for VLA Initialization
TL;DR AI
2 min readKey summary
The study examines how to initialize vision-language-action (VLA) models from pretrained vision-language models (VLMs).
It finds that preserving pretrained VLM representations matters most, while embodied VQA supervision and robot-data pretraining further help.
LoRA gives more stable initialization than full finetuning, and staged LoRA-based robot pretraining performs best overall.
The results offer practical guidance for building stronger robot policies from multimodal pretrained models.
