Switch language한국어
Back to the list

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

TL;DR AI

Key summary

2 min read
  1. TurboVLA is a compact vision-language-action model that maps vision and language directly to robot actions, skipping the usual LLM-centered pipeline.

  2. It reached 97.7% average success on LIBERO with just 0.2B parameters, 31.2 ms latency, and 0.9 GB inference VRAM.

  3. The model ran real-time robot control at 32 Hz on an RTX 4090, showing strong manipulation performance with far lower resource use.

  4. Its efficiency could make practical robot deployment on consumer-grade hardware much more feasible.

Read the original