Show-Harness: A Semantic Action Interface Lets VLMs Control Real Robots Zero-Shot

Show-Harness introduces an Embodied Harness — a semantic action interface enabling vision-language models to control diverse robots through interpretable, composable actions like move, rotate, grasp, and release, achieving 89–100% zero-shot success across tasks, environments, and embodiments while unifying human teleoperation, GUI agents, and VLM control via the GUMI interface.

Machine Heart
Machine Heart
Machine Heart
Show-Harness: A Semantic Action Interface Lets VLMs Control Real Robots Zero-Shot

Problem: Bridging Foundation Models and Physical Robot Control

Today's frontier vision-language models (VLMs) excel at visual understanding, spatial reasoning, and task decomposition, yet they cannot directly operate real robots. Prior work either trains Vision-Language-Action (VLA) models to regress continuous control signals — requiring re-adaptation for each robot and data distribution — or uses large models only for high-level planning while delegating low-level execution to separate controllers, creating a semantic-to-physical gap.

Show-Harness: An Embodied Harness as a Universal Interface

Researchers from NUS Show Lab propose Show-Harness , an Embodied Harness that sits between a VLM and a robot. Instead of forcing the model to output joint torques or end-effector poses, the harness defines a compact set of semantic action units that are both human-readable and physically grounded. The same interface serves closed-source frontier VLMs in zero-shot mode and small open-source VLMs (e.g., Qwen2.5-2B) with lightweight LoRA fine-tuning (rank-64, ~3% parameters updated, vision encoder and projector frozen).

Semantic Action Space

The action vocabulary includes relative movements (forward/back, left/right, up/down), rotations around specific axes, grasp, release, and terminate. Two design principles govern this space:

Semantically understandable: The model reasons in terms of "which direction is the target?" "should I move closer or lower?" "am I aligned?" rather than predicting joint angles.

Physically grounded: Each semantic unit maps to a small, bounded physical displacement. Robot-specific action interpreters translate the shared intent into safe, embodiment-specific control commands.

Changing the robot body only requires swapping the low-level interpreter; the VLM-facing interface remains identical. For zero-shot frontier agents, cross-embodiment transfer is achieved solely by replacing the interpreter.

Perceive–Reason–Act Closed Loop

Show-Harness organizes a continuous loop: multi-view images, proprioceptive state, recent action history, and the task description are fed to the VLM. The model observes, decomposes or updates sub-tasks, emits a semantic action, and then receives fresh visual and state feedback for correction — avoiding one-shot long-horizon plans that execute blindly.

Figure 3: Show-Harness perceive–reason–act loop with shared semantic interface
Figure 3: Show-Harness perceive–reason–act loop with shared semantic interface

GUMI: A Unified GUI for Humans, Computer-Use Agents, and VLMs

The same semantic actions are exposed as a browser-based graphical interface — GUMI (GUI-based Manipulation Interface) — with clickable controls and keyboard shortcuts. This creates a three-way alignment:

Human operators teleoperate robots via keyboard without specialized hardware.

Computer-Use agents (e.g., agents that click/type/scroll on web pages) can drive the robot through the identical GUI.

VLM robot agents predict semantic actions directly, bypassing the GUI.

Every step records an (observation, semantic action) pair. Because the interpreter also logs the corresponding low-level commands and trajectories, a single demonstration simultaneously serves as training data for Show-Harness policies, continuous-control policies, and supports sim/real, single/dual-arm, and human-in-the-loop correction.

Figure 4: GUMI dual-arm control panel and Computer-Use agent operating it
Figure 4: GUMI dual-arm control panel and Computer-Use agent operating it

Data collected via GUMI: 164 real-robot demonstrations (7.8K decision steps) — 101 on Franka (5.0K steps), 63 on AgileX (2.8K steps); 230 simulated demonstrations (13.5K steps) across ManiSkill and RoboLab environments.

Generalization Across Models, Embodiments, and Tasks

The paper evaluates three axes of generalization on real hardware. Results compare zero-shot (ZS) frontier VLMs, fine-tuned (FT) Qwen2.5-2B, and the strongest baseline in each setting:

Cross-task (10 object–container tasks): ZS 89%, FT 86%, Best Baseline 57%

Cross-environment (background, lighting, viewpoint, distractors): ZS 100%, FT 88%, Best Baseline 65%

Cross-embodiment (Franka 7-DoF vs AgileX 6-DoF dual-arm): ZS 93%, FT 87%, Best Baseline 52%

Tasks range from basic pick-and-place to unfamiliar objects, spatial reasoning, and dual-arm drawer opening. Qualitative demos include erasing text, folding clothes, placing a pen in a holder, and cutting cake.

Figure 5: Real-robot generalization — novel objects, environment changes, spatial reasoning, dual-arm
Figure 5: Real-robot generalization — novel objects, environment changes, spatial reasoning, dual-arm

Sim-to-Real Transfer

Fine-tuned on only the 230 simulated demonstrations (collected via the same GUMI interface), the model achieves 13/20 successes on a real Franka. Comparable trainable VLA baselines score 0/20 under the same protocol.

Physical vs. Semantic Adaptation Analysis

Beyond broad generalization, the authors probe two complementary adaptation capabilities:

Physical adaptivity: handling changes in control precision, action composition, workspace geometry, and embodiment without policy retraining .

Semantic adaptivity: solving tasks that require reasoning or in-context learning from new demonstrations.

Figure 7: 2×2 ablation on action naming vs. physical conventions
Figure 7: 2×2 ablation on action naming vs. physical conventions

A key ablation replaces intuitive action names with abstract labels A–F while keeping the physical conventions intact: success rate remains 95% . Keeping semantic names but dropping explicit conventions yields 90% . Removing both collapses performance to 5% . This shows that semantic names activate the model's spatial priors, while stable "symbol → physical effect" conventions are essential for grounding.

Conclusion and Outlook

Show-Harness demonstrates that a well-designed semantic interface — an Embodied Harness — can translate the existing capabilities of foundation models into robust, generalizable robot control without compressing intelligence into continuous vectors or leaving it stranded at the planning level. The same interface also unifies data collection, human teleoperation, GUI-based agents, and direct VLM control. The authors suggest that future Computer-Use and Robot-Use agents may share the same foundation model, differentiated only by the interface layer connecting to digital versus physical worlds.

Related Work: Survey on Multimodal Embodied Agents

Following Show-Harness, the lab released a survey "Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use" (arXiv:2609.10522) structured around a Perceive → Anticipate → Plan → Act → Verify (PAPAV) framework, with an accompanying Awesome Repo tracking the fast-evolving literature.

Resources

Project page: https://showlab.github.io/Show-Harness/ Paper: https://arxiv.org/abs/2609.10522 Code: https://github.com/showlab/Show-Harness Models: https://huggingface.co/showlab/Show-Harness-VLMs Data: https://huggingface.co/datasets/showlab/Show-Harness-Data Awesome Repo: https://github.com/showlab/Awesome-Multimodal-Embodied-Agent Awesome Project page:

https://showlab.github.io/Awesome-Multimodal-Embodied-Agent/
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Embodied AIVision-Language ModelsZero-Shot Generalizationrobot controlGUMISemantic Action InterfaceShow-HarnessSim-to-Real Transfer
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.