VLAct: 16 GPUs, 20% Data Beats GR00T N1.6 in Cross-Embodiment Transfer
VLAct introduces a representation-centric continued pre-training framework for Vision-Language-Action models, achieving 92.5% on RoboTwin 2.0 and surpassing all World Action Models on RoboDojo using only 16 GPUs and open data; with 20% downstream data it outperforms GR00T N1.6 on unseen GR-1 robot.
