Spotter: Embodied Model Leads Execution, VLM Supervises in Parallel for Minute-Level Tasks

Researchers from Griffith University and Tsinghua University introduce Spotter, a robot agent framework where an embodied model executes tasks while a vision-language model supervises in parallel, intervening only upon detected errors, achieving minute-level task completion with 70% less time than VLM-led baselines across RoboCasa, RoboTwin 2.0, and real-robot experiments.

PaperAgent
PaperAgent
PaperAgent
Spotter: Embodied Model Leads Execution, VLM Supervises in Parallel for Minute-Level Tasks

Overview

VLM-driven embodied models are becoming mainstream for open-scene robot manipulation, but existing systems are VLM-led: the VLM plans stepwise while the embodied model acts as a tool, placing the VLM on the critical execution path. System latency then grows with task length rather than error count.

Researchers from Griffith University, Tsinghua University's Institute for Interdisciplinary Information Sciences, and collaborators propose Spotter , which inverts this relationship: the embodied model leads execution while the VLM supervises in parallel, intervening only when an error is confirmed and then reflecting on the repair. The method requires no training or modification of the embodied model and yields consistent improvements on RoboCasa, RoboTwin 2.0, and a real Franka Research 3 robot.

Spotter concept illustration
Spotter concept illustration
Spotter overview
Spotter overview

Successful episodes average 0.7–0.8 minutes , only 13–16 seconds more than the embodied model alone; compared to a VLM-led baseline using the same VLM, successful episode time is reduced by ~70% .

Effect Demonstrations

Real Robot (Franka Research 3, judge model GPT)

In a cup-stacking task the gripper misaligns with the cup rim. The screener raises an alarm, the judge confirms, executes a single lateral adjustment, and returns control — task completes. Video is accelerated; real time shown in top-right corner.

Real robot demo: cup stacking
Real robot demo: cup stacking

Real robot demo: cup stacking

Simulation (RoboCasa, local Qwen throughout)

Left: embodied model alone misses a grasp and fails to notice; right: with Spotter the system detects the empty grasp, retracts to the pre-grasp pose, re-grasps, and succeeds.

Simulation demo: left embodied model alone, right with Spotter
Simulation demo: left embodied model alone, right with Spotter

Simulation demo: left embodied model alone, right with Spotter

Structural Latency of VLM-Led Paradigm

In VLM-led collaboration the fast system (embodied model) and slow system (VLM) switch wholesale: the embodied model waits during VLM inference, and runs unsupervised during execution. The VLM must pre-commit to an execution segment length, creating a hard trade-off: short segments waste time on unnecessary checks; long segments risk irreversible errors before detection.

Per-episode time (seconds): overall, successful, failed episodes
Per-episode time (seconds): overall, successful, failed episodes

Per-episode time (seconds): overall, successful, and failed episodes

On RoboCasa, VLM-led baseline with Qwen takes 146–164 seconds on successful episodes, while the embodied model alone needs only 23–34 seconds . The extra time comes almost entirely from unnecessary waiting and checking on episodes the embodied model could have completed independently.

Motivation: Can Embodied Models Self-Correct?

Language models benefit from an error-detection-and-correction reflection loop. The team first tested whether current embodied models possess similar ability.

Same grasp failure: top embodied model alone, bottom with Spotter
Same grasp failure: top embodied model alone, bottom with Spotter

Same grasp failure: top embodied model alone, bottom with Spotter

Embodied models learn only from successful demonstrations, never seeing post-failure state distributions. Their decisions rely solely on the current observation, unable to distinguish "recovering" from "repeating the same error." Making that distinction requires full episode history and telemetry in textual form — precisely the VLM's strength. Hence Spotter's design principle: embodied model handles routine execution; VLM intervenes only at error moments.

Embodied model self-correction evaluation: (a) repair rate; (b) scorer ranking accuracy in normal vs post-failure states
Embodied model self-correction evaluation: (a) repair rate; (b) scorer ranking accuracy in normal vs post-failure states

Embodied model self-correction evaluation: (a) repair rate; (b) scorer ranking accuracy in normal vs post-failure states

Method

3.1 Problem Definition

Embodied policy π outputs an action chunk given language instruction, RGB images, and robot state (end-effector pose + gripper state). Task ends when success condition is met or step budget exhausted. Spotter freezes π and adds a supervisor that sees the same images and state but is never told whether the task succeeded. Depth images are used only by the executor to convert repair targets from pixel to 3D; the VLM never receives depth. Goal: improve π's success rate within the same step budget while overlapping supervision and execution.

3.2 Overall Framework: Parallel Supervision

Three combination modes: (a) embodied model alone; (b) VLM-Led; (c) Spotter
Three combination modes: (a) embodied model alone; (b) VLM-Led; (c) Spotter

Three combination modes: (a) embodied model alone; (b) VLM-Led; (c) Spotter

Spotter comprises four modules:

Bounded Asynchronous Execution. The embodied model executes action chunks sequentially; after each chunk it submits the image and telemetry to supervision and immediately continues. Supervision may lag but lag is strictly bounded, ensuring judgments are never based on stale observations. The embodied model pauses only when an error is confirmed or the lag bound is reached.

Two-Tier Supervision. First tier: a locally deployed screener performs fast risk assessment on every chunk, optimized for high recall (missed errors cost far more than false alarms). Only when the screener raises consecutive alarms is the second-tier judge (GPT or Qwen) invoked. The judge uses the full episode history to decide continue, intervene, or abort. Expensive inference is thus confined to a few suspicious chunks, and the embodied model keeps executing during judge inference.

Action-Primitive Repair. Upon confirmed error, the judge does not instruct the embodied model (Section 2 showed that fails). Instead it directly generates a repair plan using action primitives: move to image target point, relative translation, lift, wrist rotation, open/close gripper, retreat to historical pose. Control returns to the embodied model after repair.

Expectation-Based Reflection. Each intervention carries an observable expectation, e.g., "gripper aperture must stay above threshold during lift." In subsequent supervision windows the judge first verifies whether the expectation holds: if yes, no repeat repair; if no, the next intervention must use a different repair type. Without this, judges tend to "repair and let go" without confirming effectiveness.

Experiments

Success rate comparison with other embodied models: (a) RoboCasa; (b) RoboTwin 2.0 Hard
Success rate comparison with other embodied models: (a) RoboCasa; (b) RoboTwin 2.0 Hard

Success rate comparison with other embodied models: (a) RoboCasa; (b) RoboTwin 2.0 Hard

Success rates (%) under same step budget
Success rates (%) under same step budget

Success rates (%) under same step budget

Key observations:

Gains concentrate on pick-and-place tasks. With GPT, pick and place tasks improve by 10.3 and 9.0 percentage points respectively; missed grasps are most common and most recoverable via retry.

Error correction alone captures most benefit. On Cosmos Policy, Spotter (GPT) reaches 73.5% , matching a VLM-led baseline ( 73.3% ) that has privileged success signals and a looser step budget.

VLM-Led relies on privileged information. With Qwen, VLM-led baseline drops from 71.3%/69.1% (with success signal) to 65.2%/60.5% (without), falling below embodied model alone. Spotter improves steadily without any privileged signal.

4.1 Intervention Quality

Judge detection rate and single-repair success rate
Judge detection rate and single-repair success rate

Judge detection rate and single-repair success rate

Detection: When Qwen serves as both screener and judge, ~ 85% of judge-intervened episodes are cases where the embodied model alone would fail; with GPT as judge the proportion is higher.

Repair: A single repair recovers 10–20% of errors, outperforming the embodied model given an oracle error signal and three retries ( 7.5–9.7% ).

4.2 Real-Robot Experiments

Real-robot experiments: three tasks and success rates
Real-robot experiments: three tasks and success rates

Real-robot experiments: three tasks and success rates

On a Franka Research 3, the team evaluated cup stacking, placing a cube into a cup, and a long-horizon task composed of two sub-tasks (10 trials each). Adding Spotter (GPT) raised average success rate from 53% (16/30) to 83% (25/30) . The long-horizon task improved most dramatically ( 30% → 70% ): longer task chains amplify the probability that a single-step error causes total failure, so timely correction yields larger gains.

References

https://arxiv.org/abs/2609.36808
https://github.com/zqc3117/Spotter
https://zqc3117.github.io/Spotter
Long Li, [email protected]
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Embodied AIVision-Language ModelError RecoveryRoboCasaReal-Robot ExperimentsParallel SupervisionRobot AgentSpotter
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.