AdaRoboVLG: Task-Adaptive Vision-Language Grasping for Diverse Robot Hands

AdaRoboVLG introduces a modular framework that decouples task understanding from grasp synthesis, using composable foundation model priors to generate stable, task-appropriate grasps across different robot hands, validated on multiple benchmarks and real-world cluttered and dynamic scenes.

Machine Heart
Machine Heart
Machine Heart
AdaRoboVLG: Task-Adaptive Vision-Language Grasping for Diverse Robot Hands

AdaRoboVLG is a framework for vision-language grasping (VLG) that separates "how the task requires grasping" from "how the hand stably grasps" . It connects these via a unified grasp interface composed of object geometry, Contact Grasp Representation (CGR), and hand-compatible grasp types. CGR encodes contact position, probability, approach and closing directions, and grasp width; grasp types are derived from human grasp taxonomies filtered by the hand's kinematic capabilities.

01 Decoupling Understanding and Grasping for a Shared Grasp Foundation

The base policy takes CGR and grasp type to generate hand-specific candidates, then scores them with a shared Hand-Object Interaction (HOI) representation that encodes local contact geometry (positions, normals) for force-closure stability prediction. Different hands produce candidates via their own kinematic mappings; the shared scorer is reused. Adding a new hand only requires its kinematic mapping and compatible grasp types.

The base policy was trained on 4.4 million annotated grasp trials in Isaac Sim using three hands: DH3, Allegro, and Inspire. Experiments show HOI converges faster and achieves lower training loss than independent hand-object encodings. Joint training yields grasp-quality prediction accuracies of 82.95% (DH3), 88.02% (Allegro), 91.07% (Inspire) .

Humans use common sense and experience to adjust grasps; robots combine vision, segmentation, and language model priors
Humans use common sense and experience to adjust grasps; robots combine vision, segmentation, and language model priors
AdaRoboVLG overall framework: composable priors connect to base grasp policy via structured interface
AdaRoboVLG overall framework: composable priors connect to base grasp policy via structured interface
Base policy simulation data collection pipeline
Base policy simulation data collection pipeline
Base policy training convergence (left) and cross-hand prediction accuracy (right)
Base policy training convergence (left) and cross-hand prediction accuracy (right)

02 Three Priors Answering "Can Grasp", "How to Grasp", "How to Track"

Spatial Prior: Navigating Clutter

Uses multi-view RGB-D, lifts DINOv3 image features to 3D for scene-level semantics, combines with point cloud, and predicts scene-level CGR via Contact-GraspNet (trained on GraspClutter6D). Evaluated on DexGraspNet 2.0 across three clutter levels and three hands, AdaRoboVLG achieves an average success rate of 88.4% , topping the table on Random and Loose splits; Dense split still has a stronger baseline.

DexGraspNet 2.0 multi-finger grasp success rates
DexGraspNet 2.0 multi-finger grasp success rates

Cognitive Prior: Functional Reasoning and Target Grasping

Two complementary experiments:

Task A – Open-set functional reasoning: 125 image-instruction pairs, 57 tasks, 80 object classes. Base LLM achieves 82% part accuracy and 63% grasp-type accuracy; adding RAG and CoT raises them to 94% and 87% . CoT improves part reasoning more; RAG helps grasp-type selection more.

Task B – Language-guided target grasping: Using the same cognitive module for localization and segmentation, AdaRoboVLG reaches 81.3% (GraspClutter6D) and 86.0% (GraspNet-1Billion) average success across three hands, leading all baselines.

Open-set functional reasoning ablation: part accuracy and grasp-type accuracy with/without RAG and CoT
Open-set functional reasoning ablation: part accuracy and grasp-type accuracy with/without RAG and CoT
Language-guided target grasping: two datasets, three hands success rates
Language-guided target grasping: two datasets, three hands success rates

Temporal Prior: Online Grasp Updates for Moving Targets

Combines SAM3 mask tracking with DINOv3 cross-frame feature consistency to estimate rigid-body motion and update target geometry and CGR online, without extra training. On GraspNet-1Billion Seen/Similar/Novel splits, AdaRoboVLG leads in multi-grasp tracking accuracy, translation error, and rotation error. Long-horizon visualizations show stable association during continuous motion, though symmetric objects may suffer rotation drift and heavy occlusion can cause tracking loss.

Dynamic tracking metrics on Seen/Similar/Novel splits
Dynamic tracking metrics on Seen/Similar/Novel splits
Representative long-horizon tracking: 3D semantic feature projections across frames
Representative long-horizon tracking: 3D semantic feature projections across frames

03 Real-Robot Validation: From Static Clutter to Dynamic Disturbances

Experiment 5A – Find the Right Object, Grasp the Right Part

Tested on 102 daily objects (food, household, tools, toys, electronics, textiles) with XMAN-R1 + ROHand and TianJi Marvin + ROHand. Each object: 5 randomized trials, total 510 real grasps . Success criteria: correct target, functional intent satisfied, collision-free, lift 10 cm, hold 5 s. No real-world fine-tuning. Overall success rate 83.3% . Also deployed on a parallel gripper to show cross-end-effector adaptation.

Language-guided functional grasping in real cluttered scenes
Language-guided functional grasping in real cluttered scenes

Experiment 5B – Tracking Moving Targets on a Conveyor

AgileX Piper + ROHand, 16 representative objects, conveyor speed 2–5 cm/s. Spatial and cognitive priors build initial grasp constraints; temporal prior updates online at 5 Hz . Six tasks demonstrate maintained functional grasp during motion. Further tests with parallel gripper under four human disturbances (out-of-view recovery, fast motion, continuous occlusion, simultaneous translation+rotation) yield 89.7% overall success .

Dynamic conveyor: six tasks synchronized (tennis ball, slipper, red mic, green toy car, rose, kettle)
Dynamic conveyor: six tasks synchronized (tennis ball, slipper, red mic, green toy car, rose, kettle)
Four disturbance types: out-of-view recovery, fast motion, continuous occlusion, translation+rotation
Four disturbance types: out-of-view recovery, fast motion, continuous occlusion, translation+rotation

Experiment 6 – Extending Tasks and Perception Without Retraining Base Policy

Adding a task-allocation module enables dual-arm sorting; adding a depth-completion module enables transparent object handling. Both extensions change task organization or perception input while reusing the same grasp foundation, demonstrating scalable capability composition.

Dual-arm collaborative sorting
Dual-arm collaborative sorting

04 Beyond Scaling Models: Making Grasping Capabilities Scalable

AdaRoboVLG's value lies in task-adaptive, policy-reusable, embodiment-adaptable grasping. The structured interface links task constraints to physical grasps, leaving room for stronger perception and reasoning models. From a systems perspective, robot capability scaling depends not only on model size but on synergy among perception, reasoning, skills, and embodiment. AdaRoboVLG connects foundation-model priors, reusable grasp skills, and diverse hands, showing a path to scale robot capabilities through composition.

Limitations: misidentified functional parts lead to stable but functionally wrong grasps; multi-view localization inconsistency, depth reconstruction noise, and lack of tactile feedback limit reliability. Future directions include tactile closed-loop control and pre-grasp actions (pushing, scooping) under severe occlusion.

From diverse hands to static and dynamic real-world scenes, AdaRoboVLG provides a concrete example of scalable robotic grasping. This decoupled design promises to more efficiently translate foundation-model advances into robot capabilities, opening new development space for systems targeting diverse tasks and embodiments.

Paper: https://arxiv.org/abs/2609.04096

Project page: https://adarobovlg.github.io/

Code: https://github.com/AdaRoboVLG/AdaRoboVLG

Simulation playground: https://github.com/AdaRoboVLG/AdaRoboVLG-Playground

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Foundation ModelsComposable PriorsDynamic TrackingGeneralizable Grasp SynthesisMulti-Hand AdaptationReal-World ValidationRobotic GraspingVision-Language Grasping
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.