CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

TL;DR AI
2 min readKey summary
Researchers introduced CoTinyVLA, a compact 0.9B vision-language-action model built on Qwen3.5-0.8B.
It uses dual-view temporal inputs, hierarchical chain-of-thought distillation from a 35B teacher, and paraphrase augmentation.
The model improves robustness and outperforms 3B- to 7B-parameter baselines on LIBERO-Plus and related benchmarks.
Despite its strong results, CoTinyVLA keeps GPU memory use low, making it attractive for embedded robotics and closed-loop control.
