Show-Harness: A Semantic Action Interface Lets VLMs Control Real Robots Zero-Shot
Show-Harness introduces an Embodied Harness — a semantic action interface enabling vision-language models to control diverse robots through interpretable, composable actions like move, rotate, grasp, and release, achieving 89–100% zero-shot success across tasks, environments, and embodiments while unifying human teleoperation, GUI agents, and VLM control via the GUMI interface.
Problem: Bridging Foundation Models and Physical Robot Control
Today's frontier vision-language models (VLMs) excel at visual understanding, spatial reasoning, and task decomposition, yet they cannot directly operate real robots. Prior work either trains Vision-Language-Action (VLA) models to regress continuous control signals — requiring re-adaptation for each robot and data distribution — or uses large models only for high-level planning while delegating low-level execution to separate controllers, creating a semantic-to-physical gap.
Show-Harness: An Embodied Harness as a Universal Interface
Researchers from NUS Show Lab propose Show-Harness , an Embodied Harness that sits between a VLM and a robot. Instead of forcing the model to output joint torques or end-effector poses, the harness defines a compact set of semantic action units that are both human-readable and physically grounded. The same interface serves closed-source frontier VLMs in zero-shot mode and small open-source VLMs (e.g., Qwen2.5-2B) with lightweight LoRA fine-tuning (rank-64, ~3% parameters updated, vision encoder and projector frozen).
Semantic Action Space
The action vocabulary includes relative movements (forward/back, left/right, up/down), rotations around specific axes, grasp, release, and terminate. Two design principles govern this space:
Semantically understandable: The model reasons in terms of "which direction is the target?" "should I move closer or lower?" "am I aligned?" rather than predicting joint angles.
Physically grounded: Each semantic unit maps to a small, bounded physical displacement. Robot-specific action interpreters translate the shared intent into safe, embodiment-specific control commands.
Changing the robot body only requires swapping the low-level interpreter; the VLM-facing interface remains identical. For zero-shot frontier agents, cross-embodiment transfer is achieved solely by replacing the interpreter.
Perceive–Reason–Act Closed Loop
Show-Harness organizes a continuous loop: multi-view images, proprioceptive state, recent action history, and the task description are fed to the VLM. The model observes, decomposes or updates sub-tasks, emits a semantic action, and then receives fresh visual and state feedback for correction — avoiding one-shot long-horizon plans that execute blindly.
GUMI: A Unified GUI for Humans, Computer-Use Agents, and VLMs
The same semantic actions are exposed as a browser-based graphical interface — GUMI (GUI-based Manipulation Interface) — with clickable controls and keyboard shortcuts. This creates a three-way alignment:
Human operators teleoperate robots via keyboard without specialized hardware.
Computer-Use agents (e.g., agents that click/type/scroll on web pages) can drive the robot through the identical GUI.
VLM robot agents predict semantic actions directly, bypassing the GUI.
Every step records an (observation, semantic action) pair. Because the interpreter also logs the corresponding low-level commands and trajectories, a single demonstration simultaneously serves as training data for Show-Harness policies, continuous-control policies, and supports sim/real, single/dual-arm, and human-in-the-loop correction.
Data collected via GUMI: 164 real-robot demonstrations (7.8K decision steps) — 101 on Franka (5.0K steps), 63 on AgileX (2.8K steps); 230 simulated demonstrations (13.5K steps) across ManiSkill and RoboLab environments.
Generalization Across Models, Embodiments, and Tasks
The paper evaluates three axes of generalization on real hardware. Results compare zero-shot (ZS) frontier VLMs, fine-tuned (FT) Qwen2.5-2B, and the strongest baseline in each setting:
Cross-task (10 object–container tasks): ZS 89%, FT 86%, Best Baseline 57%
Cross-environment (background, lighting, viewpoint, distractors): ZS 100%, FT 88%, Best Baseline 65%
Cross-embodiment (Franka 7-DoF vs AgileX 6-DoF dual-arm): ZS 93%, FT 87%, Best Baseline 52%
Tasks range from basic pick-and-place to unfamiliar objects, spatial reasoning, and dual-arm drawer opening. Qualitative demos include erasing text, folding clothes, placing a pen in a holder, and cutting cake.
Sim-to-Real Transfer
Fine-tuned on only the 230 simulated demonstrations (collected via the same GUMI interface), the model achieves 13/20 successes on a real Franka. Comparable trainable VLA baselines score 0/20 under the same protocol.
Physical vs. Semantic Adaptation Analysis
Beyond broad generalization, the authors probe two complementary adaptation capabilities:
Physical adaptivity: handling changes in control precision, action composition, workspace geometry, and embodiment without policy retraining .
Semantic adaptivity: solving tasks that require reasoning or in-context learning from new demonstrations.
A key ablation replaces intuitive action names with abstract labels A–F while keeping the physical conventions intact: success rate remains 95% . Keeping semantic names but dropping explicit conventions yields 90% . Removing both collapses performance to 5% . This shows that semantic names activate the model's spatial priors, while stable "symbol → physical effect" conventions are essential for grounding.
Conclusion and Outlook
Show-Harness demonstrates that a well-designed semantic interface — an Embodied Harness — can translate the existing capabilities of foundation models into robust, generalizable robot control without compressing intelligence into continuous vectors or leaving it stranded at the planning level. The same interface also unifies data collection, human teleoperation, GUI-based agents, and direct VLM control. The authors suggest that future Computer-Use and Robot-Use agents may share the same foundation model, differentiated only by the interface layer connecting to digital versus physical worlds.
Related Work: Survey on Multimodal Embodied Agents
Following Show-Harness, the lab released a survey "Survey on Multimodal Embodied Agents: From Computer-Use to Robot-Use" (arXiv:2609.10522) structured around a Perceive → Anticipate → Plan → Act → Verify (PAPAV) framework, with an accompanying Awesome Repo tracking the fast-evolving literature.
Resources
Project page: https://showlab.github.io/Show-Harness/ Paper: https://arxiv.org/abs/2609.10522 Code: https://github.com/showlab/Show-Harness Models: https://huggingface.co/showlab/Show-Harness-VLMs Data: https://huggingface.co/datasets/showlab/Show-Harness-Data Awesome Repo: https://github.com/showlab/Awesome-Multimodal-Embodied-Agent Awesome Project page:
https://showlab.github.io/Awesome-Multimodal-Embodied-Agent/Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
