RoboICL: Frozen GPT-6 Astra Achieves SOTA Robot Control via In-Context Learning

RoboICL enables frozen GPT-6 Astra to control dual-arm robots by providing demonstrations and the robot's own execution history as embodied context, achieving 50.64 average progress score on 30 RoboDojo tasks—surpassing prior methods—with detailed analysis of memory design, real-robot transfer, and 929 executable trajectories for inspection.

Machine Heart
Machine Heart
Machine Heart
RoboICL: Frozen GPT-6 Astra Achieves SOTA Robot Control via In-Context Learning

Overview

RoboICL introduces an embodied in-context learning approach that uses a frozen GPT-6 Astra multimodal model for direct dual-arm robot control. Instead of training task-specific vision-language-action (VLA) models, RoboICL supplies two types of context at inference time: a fixed demonstration context from an expert trajectory and an evolving interaction memory that records the robot's own actions, execution receipts, and resulting observations. The method is evaluated on the RoboDojo benchmark across 30 dual-arm manipulation tasks spanning four categories: Open (zero-shot), Memory, Precision, and Long-Horizon (each with one demonstration).

30-Task RoboDojo Results

RoboICL achieves an overall average progress score of 50.64 across 30 tasks, outperforming the previous best method PhysicalRSI (39.56) and the official zero-shot GPT-6 Astra baseline RoboProbe by 20–27 progress-score points in each category. The progress score is RoboDojo's native stage-wise metric (0–100), not a simple success rate. Notably, the Open category (zero-shot) scores 54.75 , showing that organizing online experience alone substantially improves control of the same frozen model.

RoboDojo 30-task comparison chart showing RoboICL vs PhysicalRSI, Simate-beta, VPP2-Preview, and RoboProbe across Open, Memory, Precision, Long-Horizon categories
RoboDojo 30-task comparison chart showing RoboICL vs PhysicalRSI, Simate-beta, VPP2-Preview, and RoboProbe across Open, Memory, Precision, Long-Horizon categories

Transparent Execution Records

The project page releases 929 playable trajectories (880 full evaluation + 49 partial) covering all 30 tasks and 50 layouts each. Each record provides synchronized left-wrist, head, and right-wrist video, official score, evaluation protocol, and GPT-6 Astra's step-by-step execution rationale aligned with video playback. This allows per-layout inspection of failure modes, recovery behaviors, and whether zero-score trajectories exhibit any competence.

Execution record viewer UI showing task/layout selector, triptych video, score, and model reasoning
Execution record viewer UI showing task/layout selector, triptych video, score, and model reasoning

RoboICL Architecture: Unified Observation–Action–Receipt–Observation Syntax

The system keeps GPT-6 Astra frozen. An external controller prepares context, validates action format, runs inverse kinematics, and returns execution receipts. Two context streams share the same syntax:

Demonstration context : Non-overlapping action blocks sampled from different phases of an expert trajectory (fixed per episode).

Interaction memory : The robot's own rollout history, updated after each action chunk.

Each entry follows

Observation → Action → Execution Receipt → Result Observation

. Receipts describe how much of the proposed action was actually executed, which suffixes were dropped, whether the controller interrupted, and the new observation. This closes the loop between planned and real actions, preventing the model from reasoning on stale state.

RoboICL system diagram showing frozen GPT-6 Astra, demonstration context, interaction memory, and control loop
RoboICL system diagram showing frozen GPT-6 Astra, demonstration context, interaction memory, and control loop

Observations are 1920×480 triptychs (three 640×480 RGB images concatenated horizontally) plus proprioception. GPT-6 Astra outputs an H × 14 action matrix via a constrained Act interface, representing Cartesian deltas and gripper targets for both arms.

Bounded Anchored Memory for Long Horizons

To fit within multimodal context budgets while preserving early-task information, RoboICL uses bounded anchored memory : it always keeps the first interaction, a set of anchor points spaced across the timeline, and the most recent completed interaction. Omitted segments are marked with an explicit <TRAJECTORY_GAP> token so the model does not mistake disjoint frames for continuous time. Demonstrations are similarly subsampled into J non-overlapping blocks.

Ablation (total budget J + B = 24) on four tasks shows:

(8, 16) → 57.50

(12, 12) → 80.25 (chosen for main experiments)

(16, 8) → 80.00

(22, 2) → 60.75

Overloading demonstrations while starving online memory hurts performance. Anchored memory (80.25) outperforms a first-segment-plus-recent-history baseline (73.50), with the gap largest on Build Tower, Put Bottles into Dustbin, and Insert Tubes—confirming that long-horizon tasks need both early visual evidence and latest feedback.

Case Study: Deposit Coin (Precision Insertion)

On Deposit Coin (50 layouts), RoboICL averages 78.00 (37 perfect scores) vs. PhysicalRSI's 62.27. Layout 0 trajectory shows the model correcting repeated grasp failures: it adjusts grasp position and height, then at step 205 uses the right-wrist view to detect the slot to the right of the coin, applies lateral and yaw corrections, further calibrates forward/backward and left/right offsets, releases at step 225, and confirms coin entry at step 230 for a score of 100. The demonstration provides procedural prior; online perception and memory handle layout-specific grasp slips, object shifts, and insertion errors.

Deposit Coin layout 0 trajectory visualization with step annotations
Deposit Coin layout 0 trajectory visualization with step annotations

Comparison with GPT-as-Policy (10-Task Panel)

A separate 10-task, 5-layout panel compares RoboICL (1-shot) against Direct (zero-shot) and a hybrid π₀.₅ + GPT-6 Astra system. RoboICL beats Direct by 22.10 points and trails the hybrid by only 2.00 points —without using a learned VLA for action proposals. On Build Tower, zero-shot scores 16, 1-shot reaches 100 on the 5-layout panel; scaling to 50 layouts yields 80.60, confirming the gain is not a single-layout artifact.

10-task comparison bar chart: RoboICL 1-shot vs Direct vs π₀.₅+GPT-6 Astra
10-task comparison bar chart: RoboICL 1-shot vs Direct vs π₀.₅+GPT-6 Astra

Real-Robot Experiments (Franka Research 3)

Three tasks—Peg in Hole, Folding Towel, Building Bridge—tested with three human demonstrations collected via GELLO teleoperation. Observations from wrist-mounted D405 and external D515 cameras. Each condition (0-shot, 1-shot, 3-shot) runs five randomized trials.

Average progress score: 0-shot 14.45 → 1-shot 63.33 → 3-shot 78.89 (monotonic improvement).

Peg in Hole: 2/5 successful insertions at 3-shot.

Building Bridge: +26.67 points from 1-shot to 3-shot.

Towel folding generalization: seen towel 100, unseen same-size towel 90, larger unseen towel 40—indicating demonstrations ground the model's folding knowledge to the specific robot/action space but may overfit motion scale.

Real-robot results table showing progress scores for three tasks across 0/1/3-shot conditions
Real-robot results table showing progress scores for three tasks across 0/1/3-shot conditions

Inference Cost and KV-Cache Reuse

On three tasks (Build Tower, Classify Objects, Put Bottles into Dustbin) across 30 episodes (1810 model calls), input tokens totaled 122.51M, of which 113.15M hit cache ( 92.4% token-weighted KV-cache hit rate ). The high reuse stems from stable prefixes: task instruction, demonstration, camera calibration, and fixed anchors. Only the latest interaction and its adjacent gap change per step.

An optional Jev action-reuse gate (dev-set only) predicts whether to continue executing the unexecuted suffix of the previous action chunk. It reduces GPT-6 Astra calls by ~33% (Align Blocks: 181→121) and ~48% (General Pickup: 116→60). Success rate impact is task-dependent: General Pickup improves 4/5→5/5, Align Blocks drops 3/5→2/5. This remains a development experiment, not a general acceleration solution.

Limitations and Conclusions

Main simulation results use seed 0; training data and evaluation protocols differ across compared methods.

Progress score measures stage-wise progress, not full success rate.

Real-robot evaluation limited to three tasks, five trials each.

Current validation only on GPT-6 Astra; cross-model and cross-embodiment generality untested.

Despite these bounds, the work provides a richer evidence chain than a single average: 30-task category coverage, 10-task controlled comparison, real-robot embodiment transfer, and 929 inspectable episodes. RoboICL demonstrates that when demonstrations and execution history are organized as a shared Observation–Action–Receipt–Observation language, a frozen generalist model can adapt its behavior at deployment time without parameter updates—offering a practical path toward continual embodied learning.

Resources

Project page: https://mosi-ai.github.io/RoboICL-GPT6-Astra.github.io/ Paper: https://arxiv.org/abs/2609.34261 Code:

https://github.com/Mosi-AI/RoboICL
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

embodied AIin-context learningdual-arm manipulationrobot controlRoboDojoGPT-6 Astrafrozen modelRoboICL
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.