Machine Heart
Aug 3, 2026 · Artificial Intelligence
Does VLA Action Prediction Need an LLM? TurboVLA Achieves 32 Hz with 0.2 B Params on RTX 4090
TurboVLA, a real‑time vision‑language‑action model from Huazhong University of Science and Technology and Huawei, bypasses the large language model bottleneck by directly fusing visual and language features, achieving 32 Hz online action prediction on a single RTX 4090 with only 0.2 B parameters and 0.9 GB VRAM, while maintaining high success rates across LIBERO, RoboTwin 2.0, and real‑robot tasks.
LIBEROLLMRTX 4090
0 likes · 11 min read
