How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5
The talk presents Pelican-VLA 0.5, a unified Vision‑Language‑Action model that leverages attention‑level generalization without task‑specific supervision, achieving over 91% success on RoboTwin benchmarks and demonstrating early zero‑shot generalization through a novel Reasoning Slots bottleneck.
Problem
Vision‑Language‑Action (VLA) models require large amounts of task‑specific robot data, limiting their ability to generalize to unseen objects, scenes, or robot bodies.
Emerging Insight
Recent work suggests that VLA generalization may emerge in stages: first the model learns transferable representations of where to attend and what to manipulate, then it learns precise executable actions.
Proposed Architecture
Pelican‑VLA 0.5 integrates vision‑language understanding, future‑frame generation, and action prediction in a single architecture. The key component is the Reasoning Slots module, a learnable bottleneck placed between perception and action. It routes task‑relevant visual information through a compact slot interface, inducing operation‑centric attention patterns that persist across different strategy structures, including MoT‑style architectures.
Attention‑Level Generalization
Without object labels, segmentation masks, explicit attention supervision, or task‑specific fine‑tuning, the action pathway focuses on instruction‑relevant objects and contact regions. This behavior is retained in unseen scenes and on unseen robot bodies, outperforming open‑source VLA baselines.
Empirical Evaluation
Fine‑tuned on the RoboTwin benchmark, Pelican‑VLA 0.5 achieves 91.4 % success on RoboTwin Clean and 91.0 % on RoboTwin Randomized.
In zero‑shot settings (unseen scenes, objects, and robot bodies) the model attains non‑zero success rates on several tasks, indicating early generalization.
Attention maps before and after fine‑tuning remain highly similar, suggesting fine‑tuning mainly strengthens the mapping from pre‑existing operation‑centric attention regions to executable actions rather than creating new attention patterns.
Technical Report
Full details are available at https://arxiv.org/abs/2607.06655.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
