TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
TL;DR AI
2 min readKey summary
TurboVLA is a compact vision-language-action model that maps vision and language directly to robot actions, skipping the usual LLM-centered pipeline.
It reached 97.7% average success on LIBERO with just 0.2B parameters, 31.2 ms latency, and 0.9 GB inference VRAM.
The model ran real-time robot control at 32 Hz on an RTX 4090, showing strong manipulation performance with far lower resource use.
Its efficiency could make practical robot deployment on consumer-grade hardware much more feasible.
