How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5

The talk presents Pelican-VLA 0.5, a unified Vision‑Language‑Action model that leverages attention‑level generalization without task‑specific supervision, achieving over 91% success on RoboTwin benchmarks and demonstrating early zero‑shot generalization through a novel Reasoning Slots bottleneck.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5

Problem

Vision‑Language‑Action (VLA) models require large amounts of task‑specific robot data, limiting their ability to generalize to unseen objects, scenes, or robot bodies.

Emerging Insight

Recent work suggests that VLA generalization may emerge in stages: first the model learns transferable representations of where to attend and what to manipulate, then it learns precise executable actions.

Proposed Architecture

Pelican‑VLA 0.5 integrates vision‑language understanding, future‑frame generation, and action prediction in a single architecture. The key component is the Reasoning Slots module, a learnable bottleneck placed between perception and action. It routes task‑relevant visual information through a compact slot interface, inducing operation‑centric attention patterns that persist across different strategy structures, including MoT‑style architectures.

Attention‑Level Generalization

Without object labels, segmentation masks, explicit attention supervision, or task‑specific fine‑tuning, the action pathway focuses on instruction‑relevant objects and contact regions. This behavior is retained in unseen scenes and on unseen robot bodies, outperforming open‑source VLA baselines.

Empirical Evaluation

Fine‑tuned on the RoboTwin benchmark, Pelican‑VLA 0.5 achieves 91.4 % success on RoboTwin Clean and 91.0 % on RoboTwin Randomized.

In zero‑shot settings (unseen scenes, objects, and robot bodies) the model attains non‑zero success rates on several tasks, indicating early generalization.

Attention maps before and after fine‑tuning remain highly similar, suggesting fine‑tuning mainly strengthens the mapping from pre‑existing operation‑centric attention regions to executable actions rather than creating new attention patterns.

Technical Report

Full details are available at https://arxiv.org/abs/2607.06655.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

embodied AIroboticszero-shot learningVision-Language-ActionAttention GeneralizationPelican-VLA
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.